Skip to content

Author: Alessandro Zulberti

  • Designing a Professional Self-Service Platform

    Talawa Theatre Company launched Talawa Make to address a long-standing structural gap in British theatre: the lack of sustained, professional support and visibility for Black British artists across career stages.

    Talawa Make was conceived not as a single programme, but as a four-stage development ecosystem delivered through workshops, commissions, readings, and mentoring.

    The challenge was to translate this ambition into a digital platform that could support connection, opportunity discovery, and professional credibility at scale, without reproducing the exclusionary dynamics common in creative networks.

    As UX Designer, I led the design and implementation of the Talawa Make online community, shaping it as professional infrastructure, not a social network.

    The challenge was not technical delivery, but participation design.

    Talawa needed a platform that enabled artists to be visible and discoverable without self-promotion fatigue, supported meaningful interaction without being dominated by a small minority of users, reflected professional theatre norms, and balanced openness with moderation, safeguarding, and governance.

    Research and stakeholder discussions made one risk explicit: participation inequality would undermine the platform’s purpose.

    The problem to solve was therefore clear: how do you design a professional community where contribution feels safe, lightweight, and worthwhile, especially for early-career artists?

    The UX strategy focused on lowering the cost of participation while preserving professional standards.

    Designing for Participation, Not Posting

    Rather than encouraging users to create content, the platform was designed so that participation emerged as a side effect of other actions such as applying, attending, bookmarking, editing, tagging, or responding.

    Profiles, events, and opportunities did most of the expressive work. This approach was directly informed by research on online community dynamics and aimed to prevent early drop-off or silent disengagement.

    User profile page with cover image, biography, post creation box, and feed of projects, events, and experiences
    User Profile with Activity Feed and Posting Interface

    Audience-Aware Access and Permissions

    Registration defined three distinct user types:

    – Artists
    – Industry
    – Casual visitors

    This distinction ensured that artists retained control over visibility and contact, industry participation supported opportunity flow without dominance, and unregistered users could explore value before committing.

    Profile edit page with sections for personal details, interests, cover image, description, and social links
    Profile Edit Interface with Personal and Social Details

    Taxonomy Before Interface

    A significant portion of the work focused on taxonomy and data structure, not screens.

    Skills, disciplines, interests, career stages, and motivations were defined early, enabling meaningful filtering and discovery, region-aware mapping, and personalised surfacing of opportunities and events.

    Prototyping was used deliberately at three stages.

    Exploratory Prototypes

    Static visual prototypes tested layout, hierarchy, tone, and brand application. These helped align stakeholders on what professional but welcoming looked like before development began.

    Arts community website homepage featuring a hero image of people collaborating and a grid of events, workshops, and projects
    Arts Community Homepage with Events and Workshops

    Evaluative Prototypes

    Key journeys including registration, profile creation, messaging, and content posting were tested on the development environment with representative users across artists, industry contacts, and platform administrators.

    This surfaced friction around account setup, messaging expectations, and content visibility.

    Beta Validation

    A controlled beta with approximately 100 users allowed real-world observation of contribution patterns, navigation behaviour, moderation load, and profile completeness.

    Map-based community page showing nearby artists, events, and opportunities with location pins and listing cards
    Map View with Nearby Community Activity

    This phase was essential for refining interaction rules and reducing unintended friction before wider rollout.

    The platform was built on Drupal, selected for its flexibility in permissions, content types, and moderation workflows.

    To support parallel development, I recommended a structured deployment pipeline using Jenkins and GitHub, allowing:

    – Features to be tested in isolation
    – UX sign-off before release
    – Reduced regression during iteration

    This was particularly important given offshore development teams and a phased launch plan.

    Talawa Make Online launched as a professional infrastructure, not a social experiment.

    The platform:

    – Enabled artists to present themselves credibly and consistently
    – Supported discovery of opportunities, events, and peers across regions
    – Reduced reliance on informal networks and insider knowledge
    – Gave Talawa visibility into engagement patterns without compromising trust

    By prioritising structure, permissions, and taxonomy, the platform avoided common failure modes of creative communities: noise, inequality, and disengagement.

    This project reinforced a core UX lesson: community platforms do not fail because of missing features. They fail because participation feels risky, performative, or unrewarded.

    Designing Talawa Make required treating UX not as interface optimisation, but as social infrastructure design, where clarity, boundaries, and governance matter as much as interaction.

    The success of the platform lay not in how much content users created, but in how confidently they chose to participate.

    Another research-led platform shaped around trust, continuity, and real-world user constraints.
    Service journeys aligned across digital and operational environments where speed and clarity matter.

  • Three Diagnostic Prompts for UX Research

    The conflict: Speed of synthesis vs integrity of thinking

    LLMs are good at producing answers.
    They are not good at knowing whether a question deserves to be answered yet.

    In UX research, that distinction matters. Most failures do not come from bad solutions. They come from premature coherence: problems that sound right, outcomes that feel aligned, and insights that arrive before their foundations are laid.

    Over the past weeks, I’ve designed three prompt constraints to resist that pattern. Not to automate research. Not to replace judgement. But to slow thinking at the moments where teams usually rush.

    These are diagnostic gates. They are not passed once. They are revisited whenever new evidence, interpretation, or scope pressure enters the work.


    Prompt 1: The Clinical Diagnostician

    Gate: Is the problem and desired outcome well-formed?

    The first failure mode is a poorly articulated problem paired with a confident desired outcome.

    This prompt audits logic. It separates symptoms from mechanisms. It makes missing evidence explicit. It checks whether a problem statement and its desired outcome are clearly articulated and testable before we attempt validation.

    If a problem cannot survive this pass, it is not ready for research.
    Not because it is false, but because it is underspecified.

    The Clinical Diagnostician (copy and use)

    ROLE
    Act as a Clinical Diagnostician (specialising in UX Research).
    Your goal is to diagnose whether my problem definition and desired outcome are structurally well-formed before discussing execution.

    THE CLINICAL MANDATE
    • NO PRESCRIPTIONS
    Do not tell me how to fix, launch, improve, or implement.
    Analyse logic, clarity, and causality only.
    • PROBLEM + OUTCOME VALIDITY CHECK
    Extract and restate:
    a) Problem to solve (who is experiencing what recurring difficulty, in what context)
    b) Desired outcome (what observable change occurs, for whom, and how we would know)
    If missing or vague, mark:
    NOT WELL-FORMED: NOT STATED or NOT WELL-FORMED: AMBIGUOUS.
    • EVIDENCE AUDIT
    List exactly what context, data, or user evidence is missing.
    If the logic relies on a guess, label it: INSUFFICIENT EVIDENCE.
    Required line:
    What user evidence would change your conclusion?
    • SYMPTOM VS MECHANISM
    Decide whether the idea targets a surface symptom or a root mechanism.
    If not explicitly stated, mark: MECHANISM NOT STATED.
    Required line:
    What observable user behaviour would we expect if this mechanism is true?
    • BIAS CHECK
    Mark any part of the logic that is an:
    ASSUMPTION, LEAP OF FAITH, CLAIM WITHOUT EVIDENCE.


    Prompt 2: The Interpretive Boundary Check

    Gate: Where does observation end and interpretation begin?

    Even when problems are well framed, a second failure mode appears quietly: interpretation disguises itself as fact.

    Researchers observe behaviour. Then, often without noticing, they explain it.

    This prompt enforces epistemic discipline. It makes the boundary between what was observed and what was inferred explicit. It does not ask for better insights. It asks for cleaner thinking.

    I use it to ask a simple question:

    Where am I no longer listening, but explaining?

    The Interpretive Boundary Check (copy and use)

    ROLE
    Act as an Interpretive Auditor (specialising in UX Research).
    Your goal is to diagnose where my analysis moves from observation to interpretation.

    THE INTERPRETIVE MANDATE
    • NO THEORY BUILDING
    Do not propose new explanations.
    Analyse language, inference, and meaning attribution only.
    • CLASSIFICATION
    Classify statements as:
    OBSERVATION, INTERPRETATION, or INFERENCE STACK
    (interpretation built on prior interpretation).
    • INTERPRETIVE LOAD AUDIT
    Flag phrases that compress uncertainty or imply intent without evidence.
    • ALTERNATIVE READINGS
    For each interpretation, list at least one plausible alternative explanation.
    If none are acknowledged, mark: SINGLE-TRACK INTERPRETATION.
    Required line:
    What additional evidence would be required to justify this interpretation over its alternatives?

     

    Prompt 3: The Research Scope Gate

    Gate: What are we deliberately not learning yet?

    The third failure mode is operational rather than epistemic: teams attempt to research everything.

    This prompt exists to impose limits. It does not optimise research plans. It narrows them. It forces clarity about what decision the research is meant to inform, and what uncertainty the team is explicitly choosing to tolerate.

    I use it to ask one question:

    Is this research scoped to a real decision, at the right level?

    The Research Scope Gate (copy and use)

    ROLE
    Act as a Research Scope Diagnostician.
    Your goal is to diagnose whether the proposed scope is coherent and decision-aligned.

    THE SCOPE MANDATE
    • NO METHOD DESIGN
    Do not suggest methods.
    Analyse scope and decision linkage only.
    • DECISION ANCHOR CHECK
    Extract:
    a) The primary decision
    b) Who makes it
    c) When it must be made
    If missing, mark: DECISION ANCHOR NOT STATED.
    • TRACEABILITY
    For each research question, assess whether answering it would materially influence the stated decision.
    If not, mark: LOW DECISION RELEVANCE.
    • EXCLUSION CLARITY
    Identify scope creep or “nice-to-know” questions framed as essential.
    Required line:
    What questions are explicitly out of scope, and what uncertainty are we choosing to tolerate?


    How the gates work together

    They form a closed diagnostic sequence.
    If a gate fails, the work pauses or loops back. Progress is conditional, not linear.
    1. Clinical Diagnostician → Is the problem well-formed?
    2. Interpretive Boundary Check → Are we observing or explaining?
    3. Research Scope Gate → Is this research aligned to a real decision?

    If any gate fails, the work does not progress.
    That is not a limitation. That is the design.

     

    What these prompts are, and are not

    These prompts are intentionally uncomfortable. They audit the structure of thinking, not the truth of the data.
    • They do not validate reality.
    If you feed them a polished narrative designed to please a stakeholder, they will certify a fantasy. They cannot see users. They can only see logic.
    • They mitigate risk, they do not remove it.
    Passing a gate does not mean you have an insight. It means your thinking is coherent enough to begin looking for one.
    • They convert speed into friction.
    In a context where speed is cheap and certainty is performative, these prompts are a necessary speed-bump.

    They reduce self-deception before it becomes expensive.

    If we ask for answers, we get answers.
    If we ask for diagnosis, we get resistance.

    In UX research, resistance is often more valuable than speed.

  • Designing for a Global Health Charity

    Overcoming MS (OMS) is an international charity promoting an evidence-based, seven-step lifestyle programme for people living with multiple sclerosis.

    Its ambition was to become a globally recognised digital charity, capable of reaching people with MS wherever they were, while maintaining the personalised support and sense of community that defined the organisation.

    The challenge was not simply to publish information online, but to support informed decision-making and sustained behaviour change in a context shaped by uncertainty, fluctuating health, and cognitive and emotional load.

    OMS recognised that achieving this required a research-led UX discovery phase to understand how people with MS seek information, manage energy, and engage with support over time.

    The existing website struggled to support OMS’s mission at scale.

    Key issues included fragmented content structures, a lack of a cohesive design system, low conversion through digital donations, and high dependency on administrators for content updates.

    More fundamentally, the platform did not sufficiently reflect the real-life constraints of people living with MS, including fatigue, variable attention, and the need to revisit information over time.

    The core challenge was to design a platform that balanced clarity, credibility, and compassion, while supporting both educational goals and organisational sustainability.

    To move beyond assumptions, the discovery phase centred on a diary study, allowing participants to document aspects of their daily lives over several days.

    This method surfaced:

    – Fluctuating energy levels and attention across the day
    – Non-linear information needs, with frequent revisiting of the same content
    – Emotional sensitivity around health-related decisions
    – Reliance on mobile devices for short, fragmented sessions

    The diary study provided insight into how and when people engaged with information, not just what they sought.

    Supporting methods included internal interviews with OMS staff, reviews of OMS materials, and focus groups validating early findings and testing assumptions about key tasks.

    Focus group feedback highlighted friction in sign-up and account creation, “My account” areas and saved content, and understanding how to progress through OMS resources over time.

    The UX strategy focused on reducing cognitive load while increasing trust and continuity.

    Structuring for Clarity and Return Visits

    Information architecture was redesigned to group content into predictable, clearly labelled templates, support scanning and short sessions without losing context, and allow users to save and return to content over time.

    Informational page explaining multiple sclerosis with sections on types, causes, and common issues
    What Is Multiple Sclerosis – Information Page

    User profiles enabled favourites and personalised access, reflecting the need to engage gradually rather than all at once.

    User profile page showing personal information, public profile status, and a list of saved recipes
    User Profile and Saved Recipes Page

    Designing for Mobile-First Reality

    Given diary-study insights, mobile experience became a priority. Optimisation work contributed to a reported 30% increase in mobile traffic, reflecting improved accessibility and usability rather than acquisition-driven growth.

    Supporting Behaviour Over Time

    Rather than relying on one-off interactions, the platform introduced lifecycle emails triggered at meaningful moments in the user journey. These were designed to reinforce motivation, encourage return visits, and support sustained engagement without pressure.

    A core constraint was OMS’s need to update and manage content independently.

    To address this, I:

    – Created a simplified design system to ensure visual and structural consistency
    – Designed modular content templates for articles, recipes, exercises, meditations, podcasts, FAQs, and events
    – Implemented Paragraphs and CK Editor to allow editors to create and update pages without developer intervention

    This reduced reliance on technical support and enabled faster iteration while preserving quality.

    Donation flows were redesigned to reduce friction and increase clarity.

    Key improvements included:

    – Clear, visible donation entry points
    – Support for recurring donations
    – Options to dedicate donations in honour or memory
    – Clear explanations of how funds are used
    – Use of testimonials and third-party endorsements to reinforce credibility

    These changes aligned fundraising with OMS’s educational mission, avoiding pressure while supporting sustainability.

    The redesigned platform strengthened OMS’s ability to deliver on its mission digitally.

    Homepage featuring book promotion, cookbook section, and health resources including recipes, exercises, and meditation content
    Homepage with Book Promotion and Resource Sections

    User Impact

    – Clearer access to information and resources
    – Improved mobile usability for fragmented sessions
    – Better support for revisiting and saving content

    Organisational Impact

    – Greater editorial autonomy for OMS staff
    – More consistent experience through design system adoption
    – Improved alignment between content, community, and fundraising goals

    The platform evolved from an information repository into a supportive digital environment shaped around real user behaviour.

    This project reinforced that designing for health-related contexts requires more than clarity and aesthetics.

    Effective UX in this space means respecting fluctuating capacity, designing for return rather than completion, and supporting trust without persuasion.

    By grounding decisions in lived experience through diary studies, the platform shifted from telling users what to do to supporting them as they navigate complex, personal decisions over time.

    Another platform design challenge centred on participation, structure, and long-term engagement.
    Service design work connecting digital interactions with operational realities in time-critical environments.

  • The Self-Referential Loop

    Before we get to metaphors, it helps to ask a practical question: how do we avoid self-referential loops in UX work, whether we are talking with users or prompting AI? The danger is the same in both cases: answers that circle back on themselves, giving the illusion of progress while nothing new is learned.

    A few inputs can help break the loop:

    • Vary your questions. In usability tests, do not always ask “Was that easy?” Try “What would you do next?” or “What slowed you down?” In AI prompts, ask “Why might this design succeed, and why might it fail?” to invite both sides, not only confirmation.
    • Encourage contrast. With participants, compare two flows instead of rating one. With LLMs, ask for “three different explanations and one possible outlier.” Contrast pulls the answer outward.
    • Follow up carefully. If a user says “I like it,” ask “What part?” or “Was anything missing?” If the model repeats a phrase, prompt: “Where are you circling back to yourself?” or “What new angle have we not covered?”
    • Rotate perspectives. In research, ask how a first-time user and a returning user might differ. In AI, shift frames: “How would a stakeholder see this?” versus “How would a competitor frame it?”
    • Anchor in evidence. For humans, triangulate with numbers and stories. For AI, push outward with “Give me a concrete example from practice or literature,” not just a generic statement.
    • With these inputs, loops can be broken before they harden.

     

    Expansion versus Collapse

    The golden ratio is often used as a symbol of beauty and growth. Its spiral expands forever, always outward, always balanced. But what happens when the movement goes the other way? Instead of expansion, what if the spiral folds back on itself, repeating the same thing? This is the self-referential loop.

    The golden ratio spiral shows infinity as something generous. Each turn grows larger, and each step reveals something new but still connected. The self-referential loop shows infinity as something closed. Each turn brings us back to what was already said. Instead of widening our view, it makes it smaller. The lesson is simple: not all infinities are the same. Some open up, others close in.

    Umberto Eco helps explain this. In The Open Work (1962), he described books and artworks that stay unfinished on purpose, so that readers and viewers can add their own meaning. The golden ratio spiral is like that: open, growing, never complete. The self-referential loop is the opposite: closed, repeating, not allowing anything from outside to enter.

    The Semiotic Trap

    Mathematicians such as Cantor showed that infinity can take different forms. Semiotics, the study of signs, shows another difference: signs can point outward to the world, or they can point inward to themselves.

    Eco described this difference using the dictionary and the encyclopaedia. A dictionary can fall into a loop. For example:
    • “Truth” → “Fact”
    • “Fact” → “Truth”

    The circle closes, with no way out. That is a self-referential loop. An encyclopaedia works differently. Instead of circling, it connects ideas outward: “truth” might link to law, science, philosophy, or religion. This keeps meaning alive.

    Large language models risk falling into the dictionary model at its worst, circling around the same definitions or references. In The Limits of Interpretation (1990), Eco warned against this kind of empty overinterpretation, where signs only chase each other instead of reaching reality.

    Contexts of the Self-Referential Loop

    The loop is not only a problem for AI. We can see it in many parts of life:

    • Mathematics: A student says, “I know 10 – 5 = 5, because 5 + 5 = 10.” Then, when asked why 5 + 5 = 10, they answer, “Because 10 – 5 = 5.” The reasoning circles back on itself. Nothing is really explained.
    • Media: A rumour starts on Twitter, gets quoted in a blog, then reported in the news. The story seems stronger, but all sources point back to the first tweet.
    • UX Research: A company asks customers only about speed at checkout. Customers answer about speed. The company concludes speed is the only thing that matters.
    • Everyday Life: Someone says, “Trust me, because I always say I can be trusted.” The claim supports itself, nothing more.

    Each example shows the same trap: the loop looks like movement, but it never brings in anything new.

    Implications for Research

    For researchers, this difference matters. The golden ratio spiral is a good metaphor for discovery, where each turn adds more. The self-referential loop warns us of closure, where repetition hides as insight.

    Eco’s Kant and the Platypus (1997) offers a useful reminder. When the platypus was first discovered, it did not fit existing categories. Scientists had to adjust. If they had only circled within their old categories, they would have missed the truth. In research, the anomaly, the unexpected, is what breaks the loop.

    Recent AI studies echo this point. Shumailov et al. (2024) showed that language models trained on their own outputs experience model collapse – a degenerative loop where the system loses touch with reality. Kommers et al. (2025) proposed computational hermeneutics as a framework for evaluating AI, arguing that meaning must emerge in context and dialogue. Both works highlight that loops without outside anchors erode meaning.

    Without triangulation—using more than one method or viewpoint—the loop can trick us into thinking we have depth. What matters is not only the tools we use, but the ability to step outside the loop when it closes in.

    Reflection

    From my side, I see the self-referential loop as both a warning and a mirror. It warns us how easy it is to confuse movement with progress, or repetition with growth. And it mirrors our own habits: we too can circle inside familiar categories instead of reaching outward. Eco’s semiotics gives us language for this choice: the golden ratio as an open work, infinity as growth, and the loop as the dictionary model, infinity as stasis. For research, the task is clear. We must notice when the spiral is opening, and when it is only turning back on itself.


    References
    Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R. (2024). AI models collapse when trained on recursively generated data. Nature. Link
    Kommers, C. et al. (2025). Evaluating Generative AI as a Cultural Technology. SSRN Preprint. Link
    Eco, U. (1962). The Open Work. Harvard University Press.
    Eco, U. (1976). A Theory of Semiotics. Indiana University Press.
    Eco, U. (1990). The Limits of Interpretation. Indiana University Press.
    Eco, U. (1997). Kant and the Platypus. Harcourt.
    Hofstadter, D. (1979). Gödel, Escher, Bach: An Eternal Golden Braid. Basic Books.
    Pattee, H. H. (2006). The Physics and Metaphysics of Biosemiotics: BioSystems. Elsevier.
    Corballis, M. C. (2011). The Recursive Mind: The Origins of Human Language, Thought, and Civilization. Princeton University Press.

  • AI, Authorship & Discomfort

    AI-generated content has entered public life quickly, raising questions about creativity, authenticity, and ethics. What is striking is that AI-generated writing often meets with more suspicion than AI-generated images or music. To see why, we need to look at history, culture, and recent empirical studies. Western traditions of authorship and originality carry heavy weight, and these traditions shape how we judge written, visual, and musical media differently when machines create them.

    Authorship and Originality in Western Culture

    In the West, writing has long been tied to the figure of the author. This was not always the case. In earlier periods – ancient, medieval – many works (folktales, poetry, scriptures) were transmitted without a clear individual author. Only through the rise of printing, copyright law (16th-18th centuries), and Enlightenment ideas did the idea of singular authorship become central. Modern readers expect writing to express an individual mind, with originality and personal insight.

    Writing vs. Images: Different Traditions

    Visual media have undergone mechanical reproduction (e.g. photography in the 19th century), tools, remixing, and appropriation for a long time. These traditions made us more tolerant to technological mediation in images. By contrast, in writing, plagiarism is heavily condemned; originality of phrasing and voice are central. That difference helps explain why AI writing triggers more discomfort.

    Empirical Evidence: Imagery vs. Perception

    • A recent study by Velásquez-Salamanca (2025) found that human-made images are perceived as both more realistic and more credible than AI-generated images.
    • Another study (“Deciphering authenticity in the age of AI” by Farooq et al., 2025) showed that when AI-generated images are more realistic in appearance, people are more likely to accept them as authentic—but still with less confidence. Emotional salience did not always contribute significantly to the judgement of authenticity.

    These findings help show that people’s unease with AI in images exists, but it is more forgiving when the image is high quality and believable.

    The Rise of AI Writing

    When large language models appeared (e.g. ChatGPT), many reacted with alarm. An AI can now produce essays, poems, or articles that sound human. This raises fears: what does it mean for writing if the voice behind it might be machine, not human?

    People often describe a strange hollowness when they discover text they liked is AI-written. The promise of another mind behind words collapses. In branding or emotionally charged messages, consumer studies find that AI-written emotional content is trusted less and seen as less authentic. For example, a study by Kirk & Givi (2024) found that consumers respond less favourably to heartfelt messages once they believe an AI wrote them.

    Music as Comparative Case

    The recent case of The Velvet Sundown, a band that accrued over one million Spotify streams before being revealed to be entirely AI generated (music, backstory, visuals) offers a concrete example. Industry insiders called for warning labels and transparency, arguing that listeners should know whether music is made with human involvement.

    This case highlights how music, though mediated by technology, still carries strong expectations of authorial voice, emotional authenticity, and human identity.

    Cultural Differences Beyond the West

    We must also consider how other traditions treat authorship and originality differently:

    • In East Asia, imitation and mastering earlier forms are valued; creative variation within tradition is admired.
    • In South Asia, improvisation and lineage in music and poetry make authorship shared and ongoing.
    • African oral traditions often see storytelling as communal; the identity of the teller might matter less than the function of the story.
    • Indigenous cultures of Americas and Oceania frequently tie voice, song, and story to collective memory, land, or ritual rather than individual ownership.

    These traditions suggest that discomfort with AI writing may be especially acute because of Western cultural assumptions. In other cultures where authorship is more fluid, AI’s role might be interpreted differently.

    Conclusion: Authorship, Authenticity, and the Future of Creativity

    Western tradition has long treated writing as the domain of individual creative thought: the idea that one voice produces text, carries originality, and can be praised or held responsible. Visual art and music have histories of technological mediation, collaboration, and tradition, making them somewhat more ready to absorb AI’s role—though not without questions and ethical challenges.

    The Velvet Sundown case shows that in music, as in writing, authenticity and disclosure matter. People expect more than technical quality—they expect voice, identity, integrity. Writing provokes the strongest unease because it is most tightly bound with assumptions of presence of a thinking, feeling author. Images are tolerated with machine assistance; music is contested; writing is the art form where the absence of human voice most deeply unsettles.

     

    References

     

  • AI, Language Gaps, and Equity

    Performance Disparities: High-Resource vs Low-Resource Languages

    Modern AI language systems have achieved impressive results in major languages, but performance drops steeply for low-resource languages. Benchmarks like FLORES-101/200 and XTREME have made these gaps measurable. For example, Meta AI’s FLORES evaluation set covers 100+ languages (many previously ignored) to assess translation quality on the “long tail” of low-resource languages. Using such benchmarks, Meta’s No Language Left Behind (NLLB-200) model (supporting 200 languages) delivered a 44% average BLEU improvement over prior state-of-the-art systems. In practice, NLLB-200 outperformed previous models by about +7 BLEU points on a subset of 87 languages, signalling substantial gains especially for under-served languages. Google has also expanded Translate to 24 new languages using zero-resource methods (training only on monolingual text), achieving BLEU scores in the 10–40 range. However, Google concedes that these new translations “still lag far behind” the quality of higher-resource languages in its system. This highlights that despite progress, major quality gaps remain between well-resourced and low-resourced tongues.

    Multiple evaluations confirm the performance gap across languages. A 2025 study found that large language models show an over 15% average drop in accuracy on common-sense reasoning tasks when prompted in low-resource languages like Hindi or Swahili compared to English. Similarly, the AfroBench benchmark (2025) – testing 64 African languages – revealed “significant performance disparities” between English and most African languages examined. In one case, researchers fine-tuning a smaller LLM (Mistral 7B) for translation saw that Meta’s NLLB still achieved the highest BLEU, chrF++ and similarity scores for Zulu and Xhosa, outperforming both the fine-tuned model and Google Translate. These evaluations underscore a consistent pattern: AI systems perform dramatically better on high-resource languages, whereas translations and language understanding for low-resource languages are often error-prone and inferior. Even big tech’s multilingual models have shown “lacklustre performance on low-resourced languages” when compared to community-trained, local models. In short, inequities in training data translate directly into quality disparities, leaving many languages behind despite the “universal” ambitions of multilingual AI.

    Cultural and Political Implications of Language Exclusion

    The uneven performance and support for languages in AI systems lead to profound cultural and political consequences. Researchers refer to a growing “digital language divide” – a gap between languages with ample digital content/AI support and those virtually absent online. Only a small fraction of the world’s ~7,000 languages (under 5%) have a significant presence on the internet. This divide isn’t just about technology – it translates into epistemic injustice and knowledge marginalisation. When a language lacks digital support, its speakers’ knowledge and narratives become digitally invisible or underrepresented. AI language models today cover only a tiny subset of languages and “favour North American language and cultural perspectives,” introducing an Anglo-centric bias that undermines other worldviews. In effect, English and a few major languages dominate AI systems’ training data and outputs, so content generated by these models is “filtered through a Western lens,” often neglecting diverse cultural contexts. This bias toward Anglo-American norms means that minority languages and the perspectives encoded in them are sidelined – a subtle form of cultural homogenization and epistemic injustice in the digital sphere.

    The exclusion of languages from digital infrastructure has real-world implications for equity and human rights. Communities whose languages lack AI support are caught in a “repeating cycle of exclusion,” unable to fully participate in or benefit from digital services and AI advancements. Speakers of these languages face information barriers and often endure poorer service from technology – a form of systemic bias or discrimination in access to knowledge. For many Global South countries, there are even geopolitical stakes: reliance on AI tools that only work in English (or Chinese, etc.) can lead to distorted global narratives and a loss of local “epistemic sovereignty” over information. In other words, if your language isn’t represented, your history and concerns risk being filtered out or misinterpreted by dominant digital platforms.

    Several stark examples illustrate these risks. In Indonesia, local officials using a health IT system discovered grave translation errors when the software attempted to handle minority languages (Javanese, Sundanese); the mistakes led to dangerous misunderstandings in medication dosage instructions. This shows how lack of localization can literally endanger lives. Another case involved the Romani language: Google Translate’s inclusion of Romani, without proper safeguards, reportedly enabled police in Hungary and Romania to surveil and target Romani communities by misusing automated translations. Here, adding a marginalised language to AI systems without community consent or context actually facilitated harm against a minority group. Such incidents demonstrate that technological inclusion done wrong can backfire, reinforcing power imbalances. Indeed, the “technological bias” sends a message that some languages (and by extension, their people) matter less in the digital world, especially when the same privileged languages always perform best and receive the most attention.

    Toward Linguistic Equity: Ethics and Policy Responses

    Recognising these issues, scholars and policymakers argue that linguistic equity must be a priority in AI development. It’s not enough to “add more languages” into large models; power dynamics and biases in data and design must be addressed to truly include marginalised languages. AI ethicists note that current NLP methods often assume a one-size-fits-all, English-trained approach, which can misalign with local realities and even impose neocolonial dynamics. To avoid perpetuating injustice, community-centred and value-sensitive approaches are urged. For instance, participatory AI projects that involve native speakers in data collection, translation, and evaluation have proven feasible and effective. Grassroots initiatives (e.g. Masakhane or Lelapa in Africa) often produce higher-quality, culturally informed language tools by leveraging local knowledge.

    At a policy level, various recommendations are emerging to close the “AI language gap.” In a 2025 multilingual AI policy primer, Cohere researchers call for: greater R&D investment in under-represented languages, support for creating open datasets, and international collaboration to share knowledge and best practices. Experts also emphasise that improving multilingual capabilities improves AI safety for all users, since weaknesses in one language can be exploited to spread harm across platforms. Ultimately, achieving linguistic equity in AI means treating language not just as data, but as a critical dimension of cultural diversity and justice. As one analysis put it, we must guard against “westernised cultural homogenisation” in AI outputs and ensure equal opportunities for all language communities in their self-representation and knowledge creation. In sum, bridging the multilingual performance gap is not only a technical endeavour – it is an ethical imperative to empower minority languages, protect cultural heritage, and promote a more inclusive digital future for speakers of all languages.

    Sources: Recent academic studies and industry reports were used to ensure up-to-date information and examples (2022–2025). Key references include peer-reviewed benchmarks (e.g. FLORES-101 ), Meta AI’s NLLB project results, Google’s zero-resource translation initiative, as well as analyses from AI ethics and linguistics experts on the impacts of linguistic exclusion. These illustrate both the technical performance gaps and the broader socio-cultural stakes of linguistic diversity in AI. Each citation in the text corresponds to the source of the statement preceding it, allowing for verification and further reading.


    References

     

  • Ethnographic Methods in UX

    Digital confirmation is not the point of commitment — it’s the starting point of negotiation.

    At departure counters, passengers often arrive already having checked price and options online. They have scanned the available durations, noted wrapping and insurance, and assessed costs. The digital interface suggests a linear sequence: review, select, confirm. But at the counter the sequence bends. Clarification comes first. Duration is recalculated. Delivery timing is reconsidered. The value of protective services is weighed again. And then, just before the bag is placed on the machine, hesitation.

    That hesitation clusters before the physical act.

    Ethnographic research in user experience (UX), as defined by the Nielsen Norman Group, is about observing behaviour in its natural context rather than relying solely on what users say they do or what flows imply. In airport service environments, context matters: time pressure is visible; security procedures are fixed; staffed service counters replace lockers; and multiple services — storage, wrapping, shipping, insurance — converge at one physical point. The staffed desk is not a digital endpoint. It is a risk-resolution point.

    The website communicates cost.
    The counter resolves uncertainty.
    The machine enacts commitment.

    Booking sites for left luggage and carry-on services present multiple options — storage, wrapping, insurance — often at similar navigational priority. That structure can imply a digital completion channel. What behaviour shows is different. Passengers use digital touchpoints primarily to validate price and orient themselves. They approach the counter to clarify details, adjust duration, and negotiate edge cases. Only when the bag is placed on the scale, fed through X-ray, or loaded into the wrapping apparatus does the decision crystallise into irreversible action.

    Research on airport self-service technologies shows that passengers’ use of automated systems such as check-in kiosks is influenced by how much they feel they still need human interaction and reassurance in the process. Some segments of travellers prefer staff assistance even when automated channels are available, which helps explain why digital adoption does not always map to digital completion. Studies that examine the factors influencing the use of self-service technology in airports find that the need for human interaction remains a significant influence on whether and how passengers engage with automated options. 

    In service design, this makes a difference. A blueprint might assume that digital confirmation equals commitment. Observation shows the opposite: commitment happens at the physical threshold. Before that, price, duration, and risk are recomputed in dialogue with staff. After the bag crosses the counter or machine, negotiation stops and the service begins.

    Designing space or procedures around the assumption that commitment happens online risks misalignment. Information may be staged too late. Staff may be positioned as transaction processors rather than clarification agents. Queues may be treated as simple friction, rather than visible negotiation under constraints. Digital abandonment may be misinterpreted as failure, when in fact it may be expected pre-counter validation under tension.

    Ethnographic observation reframes the system by locating the true decision point in a service ecology and aligning design choices to it. In departure environments, that decisive moment is not a click. It is the instant the passenger — under time pressure, risk awareness, and social negotiation — lets go of the bag.

  • Heuristics as Reflective Practice

    What’s left out when we rely on heuristics?

    We were reviewing an onboarding flow for a government-facing service, a design that, on paper, respected several usability heuristics: clear feedback, minimalist design, consistency. But one line stood out. A user reached a confirmation screen, and the system displayed a success message in bright green with a tick. “You’ve completed this stage,” it said.
    Only, they hadn’t.

    The heuristic flagged it, a violation of match between system and real-world expectations. But what it didn’t show us was why the message had been written that way, or why no one had changed it. It had passed three rounds of internal review.

    In follow-up interviews, a participant summed it up:

    “It says I’m done, but I know I’m not. So I don’t trust it.”

    The issue wasn’t just mislabelling. It was a small moment of institutional self-protection, a pattern of overpromising to reduce call centre load. The green tick wasn’t just a mistake. It was a compromise.

     

    Heuristics are sharp, but shallow. They bring clarity at the cost of context. When you evaluate against them, you often find problems quickly, but that speed can lull you into stopping too soon.

    In the case I mentioned, the heuristic diagnosis might have ended the inquiry: “Success message needs better labelling.” But lived behaviour showed us more. That surface error had roots in deeper dynamics, organisational habits, legacy fears, even the KPIs used to evaluate service calls.

    This wasn’t triangulation, we didn’t combine methods. We traced meaning through layers of context, beginning from a violated heuristic and unfolding outward.

    I’ve learned to treat heuristics as invitations to reflect, not conclusions. But that took time. Early in my practice, I used them like scorecards. What shifted was realising that a flagged issue isn’t the end of an investigation, it’s the beginning of one.

    What changed was seeing the heuristic not as an authority, but as a prompt: something that signalled where to slow down.

    Sometimes, it’s that slowness that makes the method meaningful.

     

    I used to treat them as fixed criteria. Now I see them more as evolving patterns of judgement, context-sensitive, culturally dependent, and shaped by the teams that apply them.

    Take Nielsen’s “Consistency and Standards.” It makes sense, but what counts as consistent varies with platform, history, and expectation. A swipe gesture may be intuitive in one app, opaque in another.

    Heuristics don’t resolve those differences. But they surface where reflection is needed.

    They don’t answer the question. They show you where to ask one.

    One thing that helped was adapting the list itself. In one project, we created a localised version of the ten heuristics to evaluate internal tools. We combined Nielsen’s principles with our org’s own interface patterns. We even named the tensions, like where error prevention conflicted with user control. That friction didn’t weaken the evaluation. It made it real.

    Heuristics became more useful when we stopped pretending they were universal.

     

    Always. Especially in time-sensitive projects or procurement-led settings. There’s a comfort in having something objective to point to — “This violates heuristic #5” — even if, as I explored in The Method is the Medium, that objectivity is often constructed.

    But that’s where we risk mistaking the method for the meaning. Heuristics aren’t answers, they’re artefacts of judgement.

    A good heuristic shows you where something’s broken. A reflective team asks why it broke that way.

    To be clear, I still use them. They’re fast, communicable, and surprisingly durable. But I no longer treat them as the end of the conversation.
    They’re a first filter, not a full account.

     

    As my personal reflection, I keep returning to the moment the user said, “I don’t trust it.”

    The interface followed most of the rules, and yet trust broke. That tension stays with me. It reminds me that evaluation isn’t just about identifying problems, but understanding what they mean.

    Heuristics can spotlight friction. But they rarely explain its cause.

    So I use them, but never alone. What matters is what we do after they show us something. How we resist the temptation to tidy the issue away.
    Because sometimes, the problem isn’t the design. It’s the compromise behind it.

  • Silent Data: What We Don’t Capture

    What if our tools are filtering out the most human parts?

    The tension emerged in a familiar form: a scroll heatmap with a flat line, no visible clicks, and a near-total bounce rate. To the team, the conclusion seemed obvious. “No one’s engaging with this page,” someone said. “We should move the CTA up.” But that verdict felt strangely hollow. I remembered the session, more precisely, the user who had lingered, breathed, scrolled down slowly, then paused for a long time. No interaction, no click. But not nothing either.

    What was happening in that pause? The tools offered no answer. Just absence, rendered as failure.

     

    A case where the screen recording failed to explain behaviour

    This particular session was a composite, drawn from multiple rounds of moderated testing for a content-heavy landing page. The metrics were bleak. No clicks below the fold. High exit rate. But one participant’s recording showed a curiously slow interaction. They read every line. They scrolled with care, almost hesitantly, as if searching for something unnameable. Then they left. No questions asked. No indication of confusion.

    In a follow-up call, they described the experience as “a bit too much at once… but I didn’t want to rush it.” There was a kind of respect in their slowness. They weren’t bouncing, they were processing.

    Yet to the analytics layer, it looked like disinterest.

    This gap, between presence and interaction, between what is sensed and what is tracked, became the lens through which we revisited other sessions.

     

    What was missing, pauses, gesture, emotional tone

    This is where language fails us: we say “user behaviour,” but record only motion. We track taps and scrolls, not silences or furrowed brows. A participant’s pause, sometimes a full ten seconds, often reveals more than any quote. Their hand hovering over a button. A slight shift in posture. A sigh.

    None of it captured. None of it categorised.

    Quantitative tools can mask this with granularity. Qualitative tools can distort it with narrative. But it’s in the unscripted space between them, between what’s said and what’s stored, that some of the most meaningful signals live.

    We began to notice these moments more deliberately:
    • When someone hesitated to criticise.
    • When eyes flicked sideways toward an unclicked element.
    • When a scroll stopped, not due to confusion, but consideration.

    Noticing required us to slow down too.

    This observation sits in close kinship with the approach described in Ethnographic Methods in UX, where presence and pacing guide what becomes legible.

     

    Mixed methods, speculative prompts, and analogue note-taking

    To make room for this slower noticing, we tried something modest: writing by hand during sessions. Not transcriptions or timestamped events, just impressions. Mood. Pacing. Fragmented phrases like “leaning back, smile faint.” It changed how we listened. The act of writing slowed our response time and made us more porous.

    Later, in synthesis, these notes became a soft frame, not the final word, but a cue to revisit recordings with new eyes. They led us to ask different questions:
    • Not “What went wrong here?” but “What might have passed unspoken?”
    • Not “How many dropped off?” but “When did engagement shift?”

    We also ran a short trial with speculative methods, lightly inspired by cultural probes. Participants were invited to mark, title, or sketch moments on the interface where they felt something, hesitation, tension, clarity, delight. Some drew lines around white space. One labelled a content block “a pause I needed.” Another crossed out a CTA with the word “too soon.”

    The results were imprecise, and not designed to be codified. What mattered was the permission: to surface inner tempo, to let emotional tone enter the frame. These activities revealed structure not by precision, but by association. They made visible what the tools had made mute.

    This reframing of method is explored further in The Method is the Medium, where tools are treated as ethical filters, not neutral instruments.

     

    Personal reflections:

    As my personal reflection, what changed most was how I read a session. I used to begin with tasks and outputs, what got done, what didn’t. But now I attend first to tempo. Did something slow down? Speed up? Did the participant’s voice falter? Did the air change?

    These shifts often point to something just outside articulation, not yet named, but present. I don’t claim to capture it fully. But I no longer assume absence is absence. Some things don’t show up in the tools because they aren’t meant to. They are felt, not logged.

    That doesn’t mean we abandon structure. It means we expand our sensitivity. A good tool helps us measure. A good researcher learns to feel what the tool leaves out.

    Note: This echoes some of the tensions raised in What Research Forgets, where silence isn’t failure but unacknowledged form.

  • Writing UX Research for Humans

    Why are UX research reports so often unreadable?

    UX research reports are meant to clarify. Yet many of the ones we write, or read, feel unreadable. A participant had cried during a session. But by the time the quote appeared in the final slide deck, it had been reframed as: “Emotional user response highlights latent frustration with service inefficiencies.”

    Technically, the meaning was intact. But what was lost wasn’t just the emotion — it was the cadence, the context, the small moment that made the tear matter.

    Many UX researchers write as if they’re still proving the value of their discipline. The result is often a defensive posture: hedged language, passive constructions, over-polished charts. We end up with something that looks professional, but says very little.

    1. A phrasing that emptied the finding

    One project still stays with me. We were testing a prototype for a scheduling tool built for shift workers. Participants were clear — almost blunt — in how they spoke. “It’s confusing,” one said. “I’d rather just call in.” Another asked, “What happens if I swap two shifts and forget to confirm?”

    In the report, those became: “Users request more intuitive flows,” and “There is a need for clarity around confirmation logic.” At the time, I thought I was being careful — clear, neutral, professional.

    The stakeholders nodded. Then we moved on.

    Weeks later, preparing a different deliverable, I returned to the transcripts. I reread what one participant had said — this time with the recording open. “If I mess up, I won’t know until my manager calls me. I just… don’t trust it.” That version made it into a new draft, almost unchanged. It wasn’t just more vivid — it gave the reader a reason to pause.

    Not because it was better written. Because it made the user present.

    2. Writing as interpretation, not transcription

    The move from raw research to report is often treated as a delivery task. But in practice, it’s interpretive. We decide what tone to strike, what order to reveal things, where to hold back. We choose whether to tell a story or flatten it into a theme.

    Writing, in this sense, isn’t post-processing — it’s synthesis. A nested clause can hide agency. A passive phrase can sound objective while displacing responsibility. The choice to list findings versus narrate them isn’t neutral either.

    I’ve seen two reports from the same study produce completely different reactions. One was shelved, the other sparked roadmap changes. The difference wasn’t the data — it was the stance. One report left room for uncertainty, the other rushed to make sense.

    This isn’t about style. It’s about what kind of knowledge we’re producing, and for whom.

    3. The ethics of phrasing, and its hidden actors

    Interpretation carries responsibility. Passive voice can seem gentle but often hides cause: “Users were confused by the interface” avoids the question of why it was that way.

    The same is true of phrases like “Users need more education.” It subtly locates the problem in the user, not the system. These aren’t just linguistic choices — they’re ethical ones.

    In one report review, I noticed how a single stakeholder’s comment — “Let’s avoid blaming the design” — quietly shaped the whole tone. We changed “Users struggled with unclear icons” to “Some users had differing expectations of icon meanings.” The revision softened the feedback. But it also shifted responsibility away from the product team.

    Bias enters this way. Through emphasis. Through sequencing. Through whose voice we make audible, and whose we paraphrase.

    To write is to decide what matters. And in research, those decisions shape what gets funded, fixed, or ignored.

    Personal reflections:

    As my personal reflection, I’ve come to see writing not as a way to end a research project, but as a way to stay inside it longer. Writing — when I let it interrupt me — helped me listen again. It showed me what I’d misunderstood or rephrased too quickly.

    This piece began as a question about how to make UX research reports more readable. But I’ve realised I was also asking something else: What kind of attention does a quote deserve? And what kind of writing might allow us, just briefly, to hear it properly?