The views expressed are my own and do not represent any organization I am affiliated with.
A clinical decision support system flags a patient for early intervention. The attending physician glances at the alert, dismisses it, and continues with her original plan. Three days later, the patient deteriorates in exactly the way the system predicted. The incident review lands on your desk.
As the usability specialist assigned to evaluate what went wrong, you face a familiar question with an unfamiliar shape. The interface was clear. The alert was visible. The physician understood what the system was recommending. She simply did not trust it enough to act.
Was this a usability failure? A calibration problem? A training gap? The answer depends on how you define the scope of usability work. And increasingly, that definition is expanding beyond what our traditional frameworks were built to address.
Introduction
Usability work has always involved interpretation. Evaluators observe behavior, infer intent, and translate findings into actionable guidance. For most of the field’s history, this focused on how people interacted with interfaces and workflows that behaved in stable, predictable ways.
That context is changing. As systems infer, adapt, and influence decisions, usability specialists are increasingly asked to make sense of not just user behavior, but system behavior as well. The work is no longer limited to identifying interface breakdowns. It now includes understanding how complex systems mediate intent, shape outcomes, and earn or lose trust over time.
The Old Professional Identity
Traditionally, usability specialists were positioned as evaluators of interaction quality. Their mandate was to reduce friction, improve efficiency, and prevent user error. Success was measured by clearer interfaces, smoother workflows, and more predictable outcomes.
This role was well defined. Engineers built systems. Designers shaped interfaces. Usability specialists validated whether those interfaces worked as intended. When systems behaved unexpectedly, the assumption was that something had been implemented incorrectly or communicated poorly. That division of labor made sense when system behavior could be fully specified in advance.
The heuristics that guided this work reflected its scope: visibility of system status, match between system and real world, user control and freedom, consistency and standards. These remain valuable. But they were developed for systems whose behavior was deterministic.
The Emerging Reality
In adaptive systems, unexpected behavior is not always a defect. It can be a byproduct of inference, learning, or contextual variation. Systems may change behavior across sessions, respond differently to similar inputs, or produce outputs that are reasonable but difficult to explain.
In these cases, usability issues do not always originate at the interface. They emerge from the interaction between human intent and system inference. Users struggle not because they cannot operate the system, but because they cannot anticipate how it will behave.
Consider the physician in the opening scenario. She knew how to read the alert. She understood what it was recommending. Her hesitation was not a failure of interface literacy. It was a failure of trust calibration, a mismatch between the system’s confidence and her own assessment of its reliability. Traditional usability methods would have certified this interface as effective. The breakdown occurred at a layer those methods were not designed to reach.
This creates a category of usability risk that cannot be addressed solely through interface refinement.
Why Usability Specialists Are Uniquely Positioned
Usability specialists are already trained to work at the boundary between human behavior and system design. They observe hesitation, confusion, overreliance, and workarounds. They are comfortable reasoning about mental models, expectations, and error recovery.
These skills translate naturally to adaptive systems. Making sense of system behavior through observed user interaction is a continuation of existing practice, not a departure from it. What changes is the object of inquiry.
Instead of asking only whether the interface communicates clearly, usability specialists must also ask whether the system’s behavior is legible, consistent, and correctable from the user’s perspective. This expanded scope involves three distinct directions of inquiry, each with its own questions and failure modes.
The Three Interpretive Roles
In adaptive systems, this translation work becomes the central competency. It operates in three directions.
Role 1: Translating System Behavior for Users
Usability specialists help teams understand why a system feels unreliable even when it functions correctly. This involves explaining why the system made a particular recommendation, identifying patterns that create confusion, and surfacing cases where correct behavior appears erratic.
Key questions: Does the user understand why the system behaved this way? Can they predict when similar behavior will recur? Do they know when the system is operating outside its reliable range?
Warning signs: Users develop superstitions about system behavior. They attribute consistency where none exists. They ignore outputs entirely because they cannot distinguish signal from noise.
Role 2: Translating User Intent for Systems
Usability specialists identify where system inference misaligns with real user goals. This requires understanding not just what users do, but what they intend. Adaptive systems often optimize for observable behavior while missing the context that gives that behavior meaning.
Key questions: Is the system learning the right lessons from user behavior? When users override recommendations, is the system correctly reading why? Are there patterns of intent the system cannot detect?
Warning signs: The system becomes increasingly confident about increasingly wrong predictions. Users find themselves fighting the system’s assumptions. Workarounds proliferate to escape unwanted personalization.
Role 3: Surfacing Risk for Organizations
Usability specialists surface cases where trust is miscalibrated. These findings inform not just design changes, but governance and deployment decisions.
The risks take multiple forms: users who defer to systems when they should override, users who ignore reliable recommendations, and trust patterns that vary unpredictably across populations or contexts.
Key questions: Where does overtrust create safety risks? Where does undertrust lead to missed opportunities? How does trust calibration vary across user populations, experience levels, and workflows?
Warning signs: Incident reports cluster around trust miscalibration rather than interface errors. User populations show bimodal adoption patterns. High-stakes decisions show delegation patterns that surprise domain experts.
Case Studies
Case 1: The Sepsis Alert That Cried Wolf
A hospital deploys a sepsis prediction system with strong validation metrics. Six months later, override rates exceed 80%. Clinical leadership asks the usability team to explain why an accurate system is being ignored.
The investigation reveals a trust calibration failure. The system fires alerts for patients who meet technical criteria but whom experienced nurses recognize as low-risk based on contextual factors the model cannot see. The alerts are not wrong, but they are unhelpful. Nurses learn to dismiss them reflexively, and this habit persists even for alerts that warrant attention.
The recommendation: not interface changes, but a governance discussion about alert thresholds, display contexts, and whether the model’s definition of risk matches clinical judgment well enough to warrant interruptive alerts. This is Role 3 work informing deployment decisions.
Case 2: The Recommendation Engine That Narrowed Options
A clinical ordering system introduces AI-powered suggestions to reduce cognitive load. Initial reception is positive. A year later, the usability team notices concerning patterns in order diversity. Physicians are selecting from an increasingly narrow range of options.
Investigation reveals a feedback loop. The system learned from physician behavior, but physicians were already being influenced by the system’s suggestions. The model became increasingly confident about a shrinking set of recommendations, and physicians became increasingly reliant on those recommendations. Neither party understood they were training each other toward convergence.
The recommendation: transparency mechanisms that help users understand the system’s learning patterns, and organizational monitoring for recommendation diversity over time. This is Role 2 work, identifying where system inference misaligns with user intent.
Case 3: The Documentation Assistant That Confused Confidence
A documentation system uses language generation to draft clinical notes. Physicians edit the drafts before signing. Error rates are acceptably low. But a usability study reveals a troubling pattern: physicians are spending less time reviewing sections where the generated text reads fluently, regardless of factual accuracy.
The fluency of the generated text creates a false signal of reliability. Physicians correctly identify errors in awkward passages but trust smooth prose even when it contains subtle inaccuracies. The system’s strengths have become a source of risk.
The recommendation: explicit uncertainty markers independent of linguistic fluency, and training that helps physicians recognize fluency and accuracy as distinct qualities. This is Role 1 work, helping users understand system behavior more accurately.
What This Role Is Not
This shift does not turn usability specialists into ethicists, data scientists, or system architects by default. It does not require deep technical implementation knowledge or ownership of model development.
It does require the ability to ask better questions. When does the system overgeneralize? Where does adaptation create confusion? How do users know when to intervene? These questions sit squarely within the usability tradition, even if the systems being evaluated are more complex.
Implementation Guidance
Practitioners looking to extend their work into adaptive systems can begin with methods they already know.
Think-aloud protocols transfer directly. Ask users to verbalize not just what they are doing, but what they expect the system to do next and why. Note discrepancies between expectation and outcome. These gaps reveal where system behavior is illegible.
Heuristic evaluation expands in scope. Add questions about behavioral consistency across sessions, transparency of inference, and recoverability when system predictions diverge from user intent. The original heuristics remain relevant; they simply address a narrower slice of the problem.
Longitudinal observation becomes essential. Adaptive systems change over time, and so does user trust. Single-session studies miss the dynamics of calibration. Plan for follow-up observations at intervals that match the system’s learning rate.
Collaborate early with data scientists. Usability specialists do not need to understand model internals, but they do need to know what kinds of behavior are possible. A brief conversation about the system’s learning mechanisms can prevent misattribution of behavior during evaluation.
Risks and Limitations
Extending into this territory creates new ways to get things wrong.
Usability specialists may overclaim expertise in domains that require deeper technical or ethical training. Identifying a trust calibration problem is not the same as knowing how to fix it at the model level. Role clarity matters: surface the issue, inform the decision, but recognize the boundaries of your contribution.
Organizations may also resist this expansion. Some view usability as a tactical function focused on interface polish, not a strategic input to deployment decisions. Practitioners extending into this work should expect to justify the scope, not assume it will be welcomed.
The interpretive frame can become an excuse for scope creep. Not every adaptive system requires this kind of analysis, and not every usability problem in adaptive systems is a trust or inference problem. Sometimes the button is just in the wrong place.
Finally, this work is not neutral. Usability specialists bring assumptions about how systems should behave and how users should relate to them. These assumptions deserve scrutiny, especially when findings inform high-stakes deployment decisions.
What It Does Mean
As systems become more autonomous, usability specialists gain influence but also responsibility. Their findings increasingly affect decisions about deployment readiness, trust boundaries, and acceptable risk. The work moves upstream, informing design choices earlier, and downstream, shaping how systems are monitored over time.
This evolution strengthens the relevance of usability practice. It positions usability specialists not as gatekeepers of interface polish, but as interpreters of human-system coordination.
Conclusion
The future of usability work is not about replacing established methods or adopting entirely new roles. It is about recognizing that systems now behave in ways users must interpret, not just operate.
Usability specialists are already skilled interpreters of human behavior. Extending that skill to system behavior is a natural progression. As systems think back, usability work becomes less about fixing screens and more about making complex behavior understandable, predictable, and safe.
That is not a departure from the field’s foundations. It is an affirmation of them.





Okay, this article comes at the perfect time! It's super interesting how you highlight the need to interpret system behavior now, not just user. How do we even begin to measure that 'earning trust' bit, tho?