<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[The Human Factors Brief]]></title><description><![CDATA[Human Factors Brief is a newsletter exploring usability, human behavior, and AI through the lens of a usability analyst focused on making complex systems work for real people.]]></description><link>https://johnwbrown.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!OCGY!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83fb7846-875f-4674-81fb-5b573ebc4dad_1254x1254.png</url><title>The Human Factors Brief</title><link>https://johnwbrown.substack.com</link></image><generator>Substack</generator><lastBuildDate>Thu, 30 Jul 2026 19:10:59 GMT</lastBuildDate><atom:link href="https://johnwbrown.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[John W Brown]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[johnwbrown@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[johnwbrown@substack.com]]></itunes:email><itunes:name><![CDATA[John W Brown]]></itunes:name></itunes:owner><itunes:author><![CDATA[John W Brown]]></itunes:author><googleplay:owner><![CDATA[johnwbrown@substack.com]]></googleplay:owner><googleplay:email><![CDATA[johnwbrown@substack.com]]></googleplay:email><googleplay:author><![CDATA[John W Brown]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[The Interface Is Quiet, but the Workflow Is Loud]]></title><description><![CDATA[Six dimensions for finding the usability problems screenshots can't show]]></description><link>https://johnwbrown.substack.com/p/the-interface-is-quiet-but-the-workflow</link><guid isPermaLink="false">https://johnwbrown.substack.com/p/the-interface-is-quiet-but-the-workflow</guid><dc:creator><![CDATA[John W Brown]]></dc:creator><pubDate>Tue, 28 Jul 2026 06:00:16 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!ViRC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9947839-f8db-45d5-bfeb-8d34d3179f36_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ViRC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9947839-f8db-45d5-bfeb-8d34d3179f36_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ViRC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9947839-f8db-45d5-bfeb-8d34d3179f36_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!ViRC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9947839-f8db-45d5-bfeb-8d34d3179f36_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!ViRC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9947839-f8db-45d5-bfeb-8d34d3179f36_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!ViRC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9947839-f8db-45d5-bfeb-8d34d3179f36_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ViRC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9947839-f8db-45d5-bfeb-8d34d3179f36_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e9947839-f8db-45d5-bfeb-8d34d3179f36_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3043541,&quot;alt&quot;:&quot;The Interface Is Quiet, but the Workflow Is Loud&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/206485911?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9947839-f8db-45d5-bfeb-8d34d3179f36_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The Interface Is Quiet, but the Workflow Is Loud" title="The Interface Is Quiet, but the Workflow Is Loud" srcset="https://substackcdn.com/image/fetch/$s_!ViRC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9947839-f8db-45d5-bfeb-8d34d3179f36_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!ViRC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9947839-f8db-45d5-bfeb-8d34d3179f36_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!ViRC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9947839-f8db-45d5-bfeb-8d34d3179f36_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!ViRC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9947839-f8db-45d5-bfeb-8d34d3179f36_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Image created with generative AI</figcaption></figure></div><div><hr></div><p>A referral coordinator in a specialty clinic opens the intake screen. It is a clean design by any conventional standard: one form, six fields, restrained color, concise labels, a single action button. A design review would praise it. A screenshot would look like a case study in restraint.</p><p>Then the task begins.</p><p>The first field wants an authorization number that lives in the payer&#8217;s portal. The second wants a date the coordinator must reconstruct from a scanned fax. The category menu cannot be set until a policy document is checked, so a second window fills with a PDF. The patient&#8217;s identifier goes onto a sticky note because it will not survive the next screen. An exception needs sign-off from a supervisor, which means a message in a chat window and a wait of unknown length. Somewhere in the waiting, the session times out. The coordinator logs back in and reconstructs which steps were finished and which values were still provisional.</p><p>The screen is quiet. The workflow is loud.</p><h2>Where the complexity goes</h2><p>Interface simplicity and work simplicity are different things. A visually restrained design can still impose heavy cognitive, temporal, and coordination demands: hidden state to remember, information scattered across systems, interruptions to survive, colleagues to chase, gaps between records to repair. These demands rarely show up in a screenshot, which is why a screen can pass a visual review while the work system around it stays hard to use.</p><p>In many of these products the complexity has not been removed. It has been displaced.</p><p>A design team trims visible fields by requiring users to fetch values from elsewhere. A screen looks uncluttered because advanced options hide three layers down. A workflow looks short because approvals, handoffs, and preparation happen outside the application. The interface becomes simpler by making the user carry more of the work.</p><p>Health information technology is the best-documented example. AHRQ&#8217;s usability research has repeatedly linked <a href="https://digital.ahrq.gov/program-overview/research-stories/improving-electronic-health-record-usability-patient-safety">electronic health record design to disrupted workflow, impeded communication, and added cognitive burden</a> when systems are built around screens instead of around actual patterns of work. ONC&#8217;s <a href="https://www.healthit.gov/resources/strategies-on-reducing-regulatory-and-administrative-burdens-relating-to-the-use-of-health-it-and-ehrs/">burden-reduction strategy</a> treats usability the same way: as a product of software design, local configuration, implementation choices, and training working together, rather than a property of the pixels. The lesson travels well beyond healthcare.</p><p>None of this is a new insight. Nielsen&#8217;s recognition-rather-than-recall heuristic is three decades old, and the distributed-cognition tradition associated with Edwin Hutchins has long argued that the unit of analysis is the whole system of people, artifacts, and tools rather than any single display. What has changed is the incentive structure. Minimal interfaces demo well, screenshot well, and pass design review well. There is no equivalent reward for the quiet workflow, because workflow burden is invisible in every artifact a product team routinely produces.</p><p>Practitioners need a way to make it visible. Here is a structure for doing that.</p><h2>The Workflow-Loudness Review</h2><p>Assess the work surrounding an interface across six dimensions.</p><p><strong>Information gathering.</strong> How many places must users search before they can act, and how do they judge which source is authoritative?</p><p><strong>Memory.</strong> What must users remember because the system does not preserve or display it?</p><p><strong>Switching.</strong> How often must users change applications, records, devices, or formats, and what does each transition cost?</p><p><strong>Coordination.</strong> Which actions depend on other people, and how is that dependency made visible?</p><p><strong>Interruption.</strong> Can users suspend the task and resume it without reconstructing their progress from memory?</p><p><strong>Repair.</strong> What happens when information is missing, inconsistent, late, or wrong, and who does the fixing?</p><p>A visually simple interface can score badly on all six. The sections that follow walk through the dimensions, and an illustrative case afterward shows what a full review turns up.</p><h2>Gathering and switching: the work between the clicks</h2><p>Interaction logs show which controls users select and how long a page stays open. They say little about what happens in between: the search in another system, the policy lookup, the phone call, the comparison of two records, the arithmetic on a notepad. The application records inactivity. The user is working.</p><p>This gap produces a flattering picture of efficiency. A workflow that contains six recorded interactions can contain eleven minutes of unrecorded assembly work, and each switch between applications carries its own tax. The user must reorient: remember why the new window was opened, find the record, extract the value, carry it back, and reattach it to the original task. The tax rises when systems disagree on terminology, when identifiers fail to match, when sessions time out, and when several similar cases are open at once. The user supplies the integration mentally, one clipboard trip at a time.</p><p>Organizations often describe such an environment as integrated because every application is reachable from the same workstation. That is access, not necessarily integration. Integration is measured by how little reconstruction the user performs. NIST&#8217;s usability guidance for health IT points the same direction: <a href="https://www.nist.gov/publications/nistir-7804-technical-evaluation-testing-and-validation-usability-electronic-health">evaluate with representative users performing realistic tasks</a>, because interface inspection alone cannot surface this burden.</p><h2>Memory: what minimalism asks users to carry</h2><p>A minimalist screen often preserves its calm by dropping persistent information: the value from the previous screen, the reason the case was opened, the status of an external request, the last action taken before an interruption. Users respond by building external memory. They write notes, keep tabs open, paste fragments into scratch documents, and photograph their own monitors. These habits are routinely misread as personal inefficiency. They are adaptations to missing state.</p><p>Recognition over recall applies at the workflow level as much as the control level. A user should be spared remembering anything the system already knows and could display at the moment of need. That requires only a concise task header: current case, purpose, last completed action, unresolved issue, next decision. This is cognitive support, and it costs very little screen space.</p><h2>Coordination: what &#8220;pending&#8221; conceals</h2><p>Many tasks are collaborative even when the interface is built for one person. When a colleague must supply a value, approve an exception, or take the next step, and the system offers no way to see or manage that dependency, users move the coordination elsewhere: email, chat, phone, a shared spreadsheet, a note left on a desk.</p><p>Meanwhile the application displays a serene status such as &#8220;pending.&#8221; The word hides a social process. Pending with whom? Was the request received? Does the other person know it is urgent? Who owns follow-up? A quiet status label can conceal a loud coordination problem.</p><p>Coordination burden also travels. One user closes a case quickly by pushing an under-specified request into a queue; a second user later spends twenty minutes discovering what the first one meant. Local efficiency can be purchased through downstream burden. This is the failure evaluation misses most reliably, because usability studies typically observe the initiating user, measure their task time, and declare success while the repair work lands on someone outside the study.</p><h2>Interruption: whether the workflow has memory</h2><p>Much professional work is interrupt-driven, and in some settings interruption is a defining condition of the job rather than a lapse of discipline. AHRQ identifies <a href="https://psnet.ahrq.gov/primers/primer/43">interruptions and distractions as significant safety concerns</a> that most work environments cannot eliminate.</p><p>The design question is what happens afterward. A user returning to a task needs to see what was active, what was completed, what was confirmed, what was tentative, and what comes next. When the system preserves none of that, the user rebuilds it from memory, and reconstruction consumes time while inviting omission, duplication, and wrong-record errors. Resumability belongs on the list of core usability properties: a good design records enough of the user&#8217;s progress to support safe reentry after attention has moved elsewhere.</p><h2>Repair and expertise: who absorbs what the design omits</h2><p>Repair is where quiet interfaces most depend on unacknowledged human skill. Some screens look simple because experienced users have already learned everything the screen declines to explain: which categories actually route correctly, which warnings can be ignored, where the missing value lives, which colleague handles exceptions. The expert carries the manual in their head. A novice sees the same quiet screen and finds no guidance at all.</p><p>This produces deceptive test results whenever evaluation leans on experienced participants. The system appears learnable because the participants already learned how to survive it. A quiet interface should not require a loud informal training network. The practical countermeasure is recruiting for inexperience on purpose: include new hires, floaters, and low-frequency users in testing, and treat the questions they ask as findings rather than noise.</p><p>Alerts deserve a note here, because they are the opposite failure with the same root. Instead of hiding workflow complexity, an alert-heavy system converts it into a stream of small judgments: relevant? urgent? already handled by someone else? safe to defer? AHRQ describes the result as <a href="https://psnet.ahrq.gov/primer/alert-fatigue">alert fatigue</a>, desensitization produced by sheer volume, which makes users likelier to override the warnings that matter. Design teams often count the number of interface elements but not the number of judgments those elements demand.</p><h2>An illustrative case: six clicks, four systems, three roles</h2><p>The case that follows is illustrative: a hypothetical assembled from patterns documented in the health IT usability literature and familiar across enterprise back-office work. No single site is depicted, and the numbers are for illustration, not data.</p><p>An intake team processes specialist referrals. The vendor&#8217;s analytics report that the median referral takes just under three minutes of active screen time: six interactions on one well-designed form. Leadership cites the number as evidence that the modernization worked.</p><p>Sit beside the coordinator, though, and the same task looks different. Completing one referral involves four applications (the EHR, the payer portal, a document viewer, and chat), two logins, a spreadsheet the team maintains privately to track pending authorizations, and a sticky-note relay for the identifier that does not survive screen transitions. Wall-clock time: eleven minutes, of which the recorded three are the smallest part. The team&#8217;s private spreadsheet, an artifact the official design never acknowledged, turns out to be the actual system of record for case state.</p><p>Follow the case downstream and the loudness spreads. Scheduling receives referrals marked complete that lack the context needed to book correctly. Staff there send roughly one clarification message back to intake for every three referrals, and each round trip adds a day. After one interruption, a coordinator resumes on the wrong patient record and catches the error only because the payer portal rejects the mismatched identifier.</p><p>Scored against the six dimensions, the design performs poorly on every one, while every screenshot of it looks exemplary. The fixes that matter are task-state persistence, a visible dependency status showing who owes what and since when, and a handoff payload that carries booking context to the next role. The form itself never changes.</p><p>The outcome measures are worth pausing on. Six weeks after remediation, clarification traffic from scheduling falls from roughly one message per three referrals to one per ten, and the private spreadsheet is retired because the system now holds case state. Median wall-clock time drops from eleven minutes to about seven. Active screen time rises, from three minutes to nearly four, because work that lived in sticky notes and chat now happens inside the application. By the vendor&#8217;s dashboard, the workflow gets slower. By every measure that touches the actual work, it gets quieter. In-application metrics mislead in both directions.</p><h2>Evaluating and designing for the workflow</h2><p>Evaluation has to follow the work off the screen. Contextual observation, in the environment where work actually happens, surfaces the switching, the notes, the side conversations, and the waiting. Workflow mapping should capture information, decisions, handoffs, delays, and exception paths, along with who knows what and when. Cognitive task analysis reaches the judgments behind visible actions: which cues users rely on, what makes a case hard, how exceptions are recognized. Artifact review collects the unofficial tooling, the spreadsheets, printed guides, and message templates that reveal functions the official design lacks. Interruption testing introduces realistic breaks mid-task and measures time to reconstruct state, duplicated actions, omitted steps, and the user&#8217;s confidence about where things stand. Cross-role research follows one case through every role it touches, which is the only way distributed burden becomes visible.</p><p>On the design side, a handful of strategies reduce workflow loudness without adding visual noise. Preserve task context so a returning user can reenter without rereading the entire record. Bring the evidence, policy, or prior action to the decision point instead of requiring users to memorize it in another window. Make dependencies visible: what is waiting, who owns it, when it was requested, and what happens if it stalls. Keep related values side by side so comparison does not depend on memory. Let provisional states stay visibly provisional instead of forcing premature certainty. Keep messages and clarifications attached to the case they concern. And evaluate every simplification for downstream effect, because a change that speeds one screen can quietly increase rework, review burden, and error somewhere else.</p><h2>Boundary conditions</h2><p>Visual simplicity remains the right goal when tasks are frequent and predictable, goals are clear, required information is at hand, and consequences are reversible. More visible structure earns its place as tasks grow complex: when users must compare evidence, when several roles coordinate, when uncertainty matters, when interruptions are routine, and when errors are hard to reverse.</p><p>Density is no virtue either. A cluttered screen creates perceptual burden and buries priority. The working target is managed complexity: important relationships visible, current state understandable, rare detail available on demand. Some complexity belongs to the task itself and cannot be designed away; a consequential decision may legitimately require several sources and careful review. The system&#8217;s job is to help users manage that complexity rather than hide it until an error surfaces it.</p><h2>Look beyond the white space</h2><p>A clean screen inspires confidence. It suggests the product has reduced the problem to its essentials. Sometimes it has. Sometimes the complexity has moved into the user&#8217;s memory, a second application, a side conversation, a private spreadsheet, or the hands of the next person in the chain.</p><p>The user experiences the search, the switching, the waiting, the interruption, the handoff, and the repair, whether or not anyone designed them. Those activities are part of the product. A successful interface makes the surrounding work more coherent, more visible, and more manageable, and it deserves to be evaluated on exactly that.</p><p>The screen is quiet. Listen to the workflow.</p><h2>Quick reference: the Workflow-Loudness Review</h2><p><strong>Six dimensions</strong></p><ul><li><p>Information gathering: sources consulted before action; how authority between sources is judged</p></li><li><p>Memory: state the user must carry; memory aids in use (notes, tabs, screenshots, private spreadsheets)</p></li><li><p>Switching: application and record transitions per task; reorientation cost of each</p></li><li><p>Coordination: dependencies on other people; whether ownership and status are visible</p></li><li><p>Interruption: whether tasks can be suspended and resumed; time to reconstruct state</p></li><li><p>Repair: handling of missing, late, or conflicting information; where downstream rework lands</p></li></ul><p><strong>Useful measures</strong></p><ul><li><p>applications and information sources used per task</p></li><li><p>time spent outside the primary interface</p></li><li><p>duplicate entry and memory-aid use</p></li><li><p>interruptions per task and time to resume</p></li><li><p>handoffs, clarification requests, and downstream rework</p></li><li><p>wrong-record errors and recovery after interruption</p></li><li><p>perceived mental demand and confidence in task state</p></li></ul><p><em>Sources: <a href="https://digital.ahrq.gov/program-overview/research-stories/improving-electronic-health-record-usability-patient-safety">AHRQ Digital Healthcare Research on EHR usability</a>, <a href="https://psnet.ahrq.gov/primer/alert-fatigue">AHRQ PSNet primer on alert fatigue</a>, <a href="https://psnet.ahrq.gov/primers/primer/43">AHRQ PSNet primer on electronic health records</a>, <a href="https://www.healthit.gov/resources/strategies-on-reducing-regulatory-and-administrative-burdens-relating-to-the-use-of-health-it-and-ehrs/">ONC Strategy on Reducing Regulatory and Administrative Burden Relating to the Use of Health IT and EHRs</a>, <a href="https://www.nist.gov/publications/nistir-7804-technical-evaluation-testing-and-validation-usability-electronic-health">NISTIR 7804</a>, <a href="https://playbook.healthit.gov/playbook/quality-and-patient-safety/">HealthIT.gov Health IT Playbook</a>.</em></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Human Factors Brief is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/p/the-interface-is-quiet-but-the-workflow?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/p/the-interface-is-quiet-but-the-workflow?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Attention Is a Budget, Not a Spotlight]]></title><description><![CDATA[Why visible is not the same as noticed, and what clinical alert fatigue teaches every designer]]></description><link>https://johnwbrown.substack.com/p/attention-is-a-budget-not-a-spotlight</link><guid isPermaLink="false">https://johnwbrown.substack.com/p/attention-is-a-budget-not-a-spotlight</guid><dc:creator><![CDATA[John W Brown]]></dc:creator><pubDate>Tue, 21 Jul 2026 06:02:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K8BQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda40905a-b159-488e-888b-8663d26a8763_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!K8BQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda40905a-b159-488e-888b-8663d26a8763_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!K8BQ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda40905a-b159-488e-888b-8663d26a8763_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!K8BQ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda40905a-b159-488e-888b-8663d26a8763_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!K8BQ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda40905a-b159-488e-888b-8663d26a8763_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!K8BQ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda40905a-b159-488e-888b-8663d26a8763_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!K8BQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda40905a-b159-488e-888b-8663d26a8763_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/da40905a-b159-488e-888b-8663d26a8763_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1149757,&quot;alt&quot;:&quot;Grid&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/206125374?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda40905a-b159-488e-888b-8663d26a8763_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Grid" title="Grid" srcset="https://substackcdn.com/image/fetch/$s_!K8BQ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda40905a-b159-488e-888b-8663d26a8763_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!K8BQ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda40905a-b159-488e-888b-8663d26a8763_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!K8BQ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda40905a-b159-488e-888b-8663d26a8763_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!K8BQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda40905a-b159-488e-888b-8663d26a8763_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Image created with generative AI</figcaption></figure></div><div><hr></div><p><em>The views expressed are my own and do not represent any organization I&#8217;m affiliated with.</em></p><div><hr></div><p>Consider a scene that repeats in hospitals every day. A physician enters a medication order. The system fires an alert warning of a possible drug interaction. She has seen this alert, or one styled exactly like it, many times already this shift, and almost all of those were noise, flagging combinations she had already judged safe. So she does what the interface has trained her to do. She overrides it and moves on.</p><p>Most of the time, the override is correct. Occasionally it is not, and the one alert that mattered is dismissed with the same reflex as the hundred that did not.</p><p>The clinician is behaving rationally. She is rationing a limited resource under time pressure, and the system has taught her that its alerts rarely repay attention. This is the problem in miniature. Users do not arrive at an interface with unlimited attention. They arrive with a task, a memory of similar systems, a set of expectations, and a limited capacity for deciding what matters next. This is why attention is better understood as a budget than a spotlight.</p><p>A spotlight suggests that attention simply shines on one object at a time. A budget suggests something more accurate: users allocate attention under constraint. Every interface element asks for some portion of that budget. Good design helps users spend it wisely. Poor design spends it for them.</p><p>This matters because many usability problems are failures of priority, not visibility. A button can be visible and still go unnoticed. A help message can sit directly on the page and still be functionally invisible if the user&#8217;s attention is already committed elsewhere.</p><p>Good UX lowers the cost of knowing where to look next.</p><h2>The Brain Does Not Process Everything Equally</h2><p>The human brain is powerful, but it is not an unlimited processor. Raichle and Gusnard note that the average adult brain represents about 2 percent of body weight, yet accounts for about 20 percent of the body&#8217;s oxygen and calorie consumption. They also note that this high rate of metabolism remains relatively constant across different mental and motor activities. [1]</p><p>That does not mean every difficult interface &#8220;burns up&#8221; the brain in a literal sense. The better point is more practical: cognition operates under constraint. People do not build a complete, neutral, high-resolution model of everything in front of them. They filter, predict, group, ignore, and prioritize.</p><p>This is the cognitive background behind many familiar UX problems. Users miss things that designers assumed were obvious. They scan instead of read. They follow patterns from previous systems. They rely on labels, location, grouping, and visual hierarchy to infer what matters. They do not inspect every object on the screen with equal care.</p><p>That reflex is efficient. The brain is rationing limited capacity, exactly as it should.</p><p>An interface that expects full inspection is already misaligned with human behavior. The better question is not, &#8220;Is the information present?&#8221; The better question is, &#8220;Can the user identify what matters at the moment they need it?&#8221;</p><h2>Attention Has Costs</h2><p>Nielsen Norman Group defines cognitive load in UX as the amount of mental processing power required to use a system. Their guidance is straightforward: unnecessary complexity, clutter, unfamiliar interactions, and poorly structured information increase the effort required to complete a task. [2]</p><p>Cognitive load is often discussed as if it were a general feeling of difficulty. In practice, it is more specific. It is the cost of holding information in mind, interpreting options, remembering steps, resolving ambiguity, and deciding what to ignore.</p><p>Every poorly labeled button spends attention. Every unnecessary alert spends attention. Every inconsistent pattern spends attention. Every visual distinction that does not communicate a meaningful distinction asks the user to evaluate whether the difference matters.</p><p>This is why attention should be treated as a design budget. Some interface elements deserve attention because they support action, safety, orientation, or recovery. Others consume it without returning value. They may be interesting or convenient, but they still cost the user.</p><p>In complex systems, that cost accumulates.</p><p>A user who must interpret ten competing visual cues before completing a routine action is doing more than coping with a busy screen. They are paying an attentional tax. The interface is asking them to spend limited cognitive resources on the system&#8217;s structure rather than the user&#8217;s goal.</p><h2>Working Memory Is Narrower Than We Like to Admit</h2><p>One reason this matters is that working memory is limited. The older &#8220;seven plus or minus two&#8221; idea is widely known, but later work has argued for a smaller practical capacity in many conditions. Nelson Cowan&#8217;s review reconsidered short-term memory capacity and argued that the central limit is often closer to about four chunks. [3]</p><p>For UX, the exact number is less important than the design implication: users cannot keep unlimited interface states, options, exceptions, and instructions active in mind. They need the interface to reduce unnecessary holding costs.</p><p>This is especially important in enterprise, healthcare, government, and other high-stakes systems where the user is often managing more than the screen itself. The interface is only one part of the task environment. The user may also be managing a patient conversation, a policy requirement, a deadline, an interruption, or a sequence of dependent decisions.</p><p>When an interface requires users to remember what should be visible, infer what should be explicit, or compare options that should have been prioritized, it consumes working memory that may be needed elsewhere.</p><p>Good design does not eliminate thought. It protects thought for the work that matters.</p><h2>Users Allocate Attention by Goal, Expectation, Risk, Novelty, and Habit</h2><p>Attention is not governed by visual salience alone. Connor, Egeth, and Yantis describe the relationship between bottom-up attention, where salient stimuli can attract notice, and top-down attention, where current goals shape what people attend to. [4]</p><p>This distinction is critical for UX.</p><p>A red badge, bold label, animation, or high-contrast button may attract attention, but only in context. If it does not align with the user&#8217;s goal, users may ignore it. If it resembles advertising, users may filter it out. If every other element is also visually emphasized, the intended signal may disappear into the environment.</p><p>Users allocate attention through several filters at once. They weigh what they are trying to do, where a control or piece of information ought to be, and what could go wrong if they miss it. They notice what has changed, and they lean on what similar systems have trained them to ignore.</p><p>This is why &#8220;make it stand out&#8221; is only half a strategy. Standing out helps only when the distinction maps to user need. Otherwise, the interface may succeed at attracting the eye while failing to support the task.</p><h2>Visible Is Not the Same as Noticed</h2><p>One of the most important lessons from attention research is that people can miss visible information. Simons and Chabris demonstrated sustained inattentional blindness, showing that people can fail to notice unexpected but visible events when their attention is engaged in another task. [5]</p><p>UX teams encounter this pattern constantly. A form instruction is visible, but users miss it. A warning is visible, but users continue. A field requirement is visible, but users discover it only after an error.</p><p>The problem is not always placement or contrast. Sometimes the element is visible in a screenshot but unavailable to attention during the actual task.</p><p>This is why usability testing matters. Static review can tell us whether something is present. Testing can tell us whether users notice it, understand it, and act on it while pursuing a real goal.</p><p>The distinction is practical:</p><p>Visibility is a property of the interface.</p><p>Noticeability is a property of the interaction between the interface, the task, and the user&#8217;s attention.</p><p>A design can satisfy the first while failing the second.</p><h2>Overemphasis Creates Attentional Inflation</h2><p>There is another problem: emphasis loses value when it becomes common.</p><p>If every message is urgent, urgency becomes background. If every card is highlighted, highlighting loses meaning. If every button is primary, users must inspect rather than recognize. If every alert is red, red no longer means &#8220;pay attention.&#8221; It means &#8220;this system is noisy.&#8221;</p><p>This is attentional inflation.</p><p>Like monetary inflation, attentional inflation occurs when a signal is overissued. The more often a system uses visual emphasis, the less purchasing power that emphasis retains. Designers then compensate by making the next signal louder: brighter color, larger banner, stronger modal, more intrusive interruption. The interface enters an arms race against its own noise.</p><p>That trains users to discriminate less. Emphasis that means everything means nothing.</p><p>Visual hierarchy depends on restraint. Nielsen Norman Group&#8217;s guidance on visual hierarchy is built on the principle that some elements must carry more emphasis than others, which means emphasis only works when most of the interface stays quiet. [6] If everything is emphasized, nothing is. Important differences need a quiet background.</p><p>Users respond to this noise actively. They learn what to ignore. Nielsen Norman Group&#8217;s work on banner blindness shows that users have learned to disregard content that resembles ads, appears near ads, or occupies locations traditionally associated with advertising. [7] The lesson generalizes well past advertising. If a system repeatedly uses prominent space for low-value messages, users learn that prominent space is not trustworthy. If a dashboard repeatedly highlights items that do not require action, users learn to discount highlights. If notification badges rarely matter, users stop treating them as meaningful.</p><p>This is adaptation. Users protect their attention when systems fail to protect it for them. And once that defensive filter forms, it does not politely exempt the message the designer most needed them to see.</p><h2>The Clearest Case: Clinical Alert Fatigue</h2><p>The medication scene that opened this piece comes straight from the literature. It is one of the most heavily documented failures in health informatics, and it shows attentional inflation operating at full scale, with patient safety as the stake.</p><p>Clinical decision support systems generate alerts to catch drug interactions, dosing errors, allergies, and duplicate therapies. The intent is sound. The execution floods clinicians with warnings, many of low or no clinical relevance. The predictable result is that clinicians override most of them. A 2024 systematic review and meta-analysis found a pooled physician override rate for drug-drug interaction alerts of about 90 percent. [8] Foundational work in the field reported override rates ranging from 49 to 96 percent depending on setting and alert type. [9] More recent reviews continue to report rates reaching as high as 96 percent, and note the direct concern: an override reflex trained on noise will also dismiss the rare alert that carries real risk. [10]</p><p>Look at how precisely this tracks the budget model. The system overissues a signal. The signal loses purchasing power. Clinicians, rationing attention under time pressure, adapt by discounting the entire category. The interface then has no reliable way to make the one critical warning register, because it spent that channel&#8217;s credibility on hundreds of trivial ones.</p><p>The friction points are specific and familiar to any practitioner. Alerts fire on broad screening rules that ignore patient-specific context such as lab values or route of administration. Low-severity and high-severity warnings share the same visual treatment, so severity carries no information. Interruptive modals demand a response for cases that did not warrant interruption. Each of these is a design decision that spends clinician attention without returning proportional value.</p><p>The most instructive finding is a paradox. In evaluations where clinicians override the overwhelming majority of alerts, many of those same clinicians still rate the system as useful. Their complaint is not with decision support itself but with an implementation that has made its own signals unaffordable to read. That is the difference between a tool that has failed and a budget that has been mismanaged. Note what this does and does not require. It keeps every safeguard. It simply spends the attention channel only where the return justifies the cost, which for clinical alerts means tiering by severity, suppressing low-value warnings, and reserving interruption for cases that genuinely warrant it.</p><h2>The Interface Should Lower the Cost of the Next Decision</h2><p>The clinical case is extreme, but the mechanism is universal. A usable interface does more than present information. It reduces the cost of deciding what to do next.</p><p>That does not mean removing all complexity. Some work is complex because the domain is complex. Clinical documentation, benefits administration, aviation maintenance, cybersecurity monitoring, and financial review cannot be made simple by wishful design. But even complex work can be structured so that users are not forced to spend attention unnecessarily.</p><p>Good interfaces answer practical questions quickly:</p><p>Where am I?</p><p>What changed?</p><p>What matters now?</p><p>What can I safely ignore?</p><p>What happens if I do nothing?</p><p>When the interface answers these questions clearly, users can reserve attention for judgment. When it does not, users must spend attention diagnosing the interface before they can complete the task.</p><p>That is the core cost of poor UX in complex systems. It moves cognitive effort away from the domain and toward the tool.</p><h2>Quiet Backgrounds Make Important Differences Possible</h2><p>A quiet background can still be rich and detailed. What makes it quiet is stability: consistent placement, predictable labels, restrained color, familiar interaction patterns, and clear grouping. That stability lets users form expectations, and those expectations are what allow them to notice meaningful violations.</p><p>A medication warning stands out because normal medication information is structured consistently.</p><p>An overdue task stands out because routine tasks are not styled as emergencies.</p><p>A destructive action stands out because ordinary actions do not share the same treatment.</p><p>In each case, the difference works because the surrounding pattern is disciplined. The design has protected the value of attention.</p><p>This is also why design systems matter. A design system is more than a library of components. At its best, it is a cognitive agreement with the user. It teaches what routine interaction looks like so that exceptions can carry meaning.</p><h2>Design Implications</h2><p>Treating attention as a budget changes the questions we ask of a design. For every element, ask what it costs the user to interpret and what it returns in orientation, safety, or clarity. Reserve the expensive signals, urgency and contrast and color and interruption, for the differences that deserve them. And test attention in context. Seeing an element in a screenshot proves little. What matters is whether users notice it, understand it, and act on it while pursuing a real goal.</p><p>The practical rule is simple:</p><p>Spend user attention only when the return justifies the cost.</p><h2>An Attention Budget Audit</h2><p>Principles are easier to affirm than to apply. The following is a practical pass you can run on an existing screen, flow, or dashboard. It does not require special tooling, only the discipline to treat attention as a spend rather than a free resource.</p><p>Start with an inventory. List every element on the screen that asks for attention: alerts, badges, banners, highlights, color-coded states, animations, and primary buttons. For each one, write down the meaningful distinction it is supposed to signal. Elements that cannot be tied to a real distinction are the first candidates for removal.</p><p>Next, count your emphasis. If more than a handful of elements use high-contrast, high-urgency treatment, emphasis is being overissued and is losing value. A screen where everything is loud has no way to signal what is actually loud.</p><p>Then check severity encoding. Where the domain has genuine tiers of importance, confirm that the interface encodes them differently. If a low-severity notice and a critical warning look the same, the interface has thrown away information the user needs.</p><p>Now audit interruption specifically. For every modal, confirmation, or blocking alert, ask whether the case genuinely warranted stopping the user&#8217;s work. Interruption is the most expensive attention spend available. It should be rare and it should be earned.</p><p>Finally, look for trained blindness. If your telemetry or testing shows users routinely dismissing, overriding, or ignoring a particular element, resist the urge to make it louder. That is the arms race. Ask instead why the element lost credibility, and whether it is firing in cases that do not deserve attention.</p><p>Run this audit honestly and most interfaces reveal the same thing: a handful of elements that earn their cost, and a longer list that spends attention the design cannot afford.</p><h2>Conclusion: Attention Without Prioritization Is Waste</h2><p>Attention is one of the most important resources in UX because it is always limited and always contested. Users bring only so much of it to any task. They spend it on goals, uncertainty, risk, memory, interruption, and interpretation.</p><p>Good UX works with that limit instead of against it. It helps users decide what matters next rather than expecting them to inspect everything. It makes routine actions predictable, meaningful changes noticeable, and unnecessary interpretation rare.</p><p>This is why attention is better understood as a budget than a spotlight. A spotlight tells us where the eye might go. A budget reminds us that every demand has a cost.</p><p>Designers do not own user attention. They borrow it.</p><p>The responsibility is to spend it carefully.</p><p>What is one element on a screen you own that spends user attention without earning it, and what would it take to retire it?</p><h2>Sources</h2><p>[1] Marcus E. Raichle and Debra A. Gusnard, &#8220;Appraising the brain&#8217;s energy budget,&#8221; <em>Proceedings of the National Academy of Sciences</em>, 2002.</p><p>[2] Kathryn Whitenton, &#8220;Minimize Cognitive Load to Maximize Usability,&#8221; Nielsen Norman Group, 2013.</p><p>[3] Nelson Cowan, &#8220;The magical number 4 in short-term memory: A reconsideration of mental storage capacity,&#8221; <em>Behavioral and Brain Sciences</em>, 2001.</p><p>[4] Charles E. Connor, Howard E. Egeth, and Steven Yantis, &#8220;Visual attention: Bottom-up versus top-down,&#8221; <em>Current Biology</em>, 2004.</p><p>[5] Daniel J. Simons and Christopher F. Chabris, &#8220;Gorillas in our midst: Sustained inattentional blindness for dynamic events,&#8221; <em>Perception</em>, 1999.</p><p>[6] Page Laubheimer, &#8220;Visual Hierarchy in UX: Definition,&#8221; Nielsen Norman Group.</p><p>[7] Kara Pernice, &#8220;Banner Blindness Revisited: Users Dodge Ads on Mobile and Desktop,&#8221; Nielsen Norman Group, 2018.</p><p>[8] Mariano Felisberto et al., &#8220;Override rate of drug-drug interaction alerts in clinical decision support systems: A brief systematic review and meta-analysis,&#8221; <em>Health Informatics Journal</em>, 2024.</p><p>[9] Heleen van der Sijs, Jos Aarts, Arnold Vulto, and Marc Berg, &#8220;Overriding of drug safety alerts in computerized physician order entry,&#8221; <em>Journal of the American Medical Informatics Association</em>, 2006.</p><p>[10] &#8220;Use of artificial intelligence to optimize medication alerts generated by clinical decision support systems: a scoping review,&#8221; <em>Journal of the American Medical Informatics Association</em>, 2024.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/p/attention-is-a-budget-not-a-spotlight?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/p/attention-is-a-budget-not-a-spotlight?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p>]]></content:encoded></item><item><title><![CDATA[The Caregiver Is the User]]></title><description><![CDATA[A virtual care agent will not arrive as a robot nurse. It will arrive as the integration of tools that already exist, handed to the most overworked person in the system.]]></description><link>https://johnwbrown.substack.com/p/the-caregiver-is-the-user</link><guid isPermaLink="false">https://johnwbrown.substack.com/p/the-caregiver-is-the-user</guid><dc:creator><![CDATA[John W Brown]]></dc:creator><pubDate>Tue, 14 Jul 2026 06:00:40 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!0yZq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39f05fab-f2fe-4c93-9eb3-3dc92305ccf4_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!0yZq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39f05fab-f2fe-4c93-9eb3-3dc92305ccf4_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!0yZq!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39f05fab-f2fe-4c93-9eb3-3dc92305ccf4_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!0yZq!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39f05fab-f2fe-4c93-9eb3-3dc92305ccf4_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!0yZq!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39f05fab-f2fe-4c93-9eb3-3dc92305ccf4_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!0yZq!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39f05fab-f2fe-4c93-9eb3-3dc92305ccf4_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!0yZq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39f05fab-f2fe-4c93-9eb3-3dc92305ccf4_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/39f05fab-f2fe-4c93-9eb3-3dc92305ccf4_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2261484,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/204337514?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39f05fab-f2fe-4c93-9eb3-3dc92305ccf4_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!0yZq!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39f05fab-f2fe-4c93-9eb3-3dc92305ccf4_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!0yZq!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39f05fab-f2fe-4c93-9eb3-3dc92305ccf4_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!0yZq!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39f05fab-f2fe-4c93-9eb3-3dc92305ccf4_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!0yZq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39f05fab-f2fe-4c93-9eb3-3dc92305ccf4_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Human Factors Brief is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><p>Discussion and Debate by NotebookLM:</p><h3>Discussion:</h3><div class="native-audio-embed" data-component-name="AudioPlaceholder" data-attrs="{&quot;label&quot;:null,&quot;mediaUploadId&quot;:&quot;55ead438-127d-4560-908e-854020ee1482&quot;,&quot;duration&quot;:1241.8873,&quot;downloadable&quot;:true,&quot;isEditorNode&quot;:true}"></div><h3>Debate:</h3><div class="native-audio-embed" data-component-name="AudioPlaceholder" data-attrs="{&quot;label&quot;:null,&quot;mediaUploadId&quot;:&quot;083d8781-1c77-4fd4-afe6-4a7857536a8a&quot;,&quot;duration&quot;:1466.2792,&quot;downloadable&quot;:true,&quot;isEditorNode&quot;:true}"></div><div><hr></div><p><em>Dan and Lee are a composite illustration drawn from the realities of full-time family caregiving. Every capability described is possible with current technology, though fragmented across products and not yet integrated as depicted.</em></p><div><hr></div><p>Lee can leave the house now.</p><p>For the eighteen months before, she could not. Her father, Dan, has heart failure, hypertension, early kidney disease, and a medication list assembled by three specialists who have never spoken to one another. Lee is his full-time caregiver. Not a nurse, though she does things nurses are trained for and she is not. Leaving Dan alone meant the possibility of coming home to a crisis she could have caught, so she stopped leaving. Getting gas meant getting Dan into the car. Shopping meant bringing him along or going without. The job was not only the labor. It was the tether.</p><p>What changed was not that a machine started caring for Dan. What changed was that Lee could go to the store and still be watching.</p><p>This is the part of the virtual care assistant story that the marketing misses and the futurism overshoots. The first serious AI care agent will not walk into the room with soft eyes and a tray of pills. It will assemble, quietly, out of parts that already exist. Continuous physiologic monitoring exists. Home video exists. Trend detection exists. Configurable alerting exists. Remote-authorized smart locks exist. The capacity to push structured data to emergency responders exists. None of it requires a single new invention. What does not yet exist is the integration: one coherent agent, with one accountability structure, governed as a single thing, designed around the person actually holding the arrangement together.</p><p>That person is Lee. And almost nothing in the current conversation about care AI is designed for her.</p><p>This piece is an argument about who the user is. Not the patient, exactly, and not the clinician. In the home, the user is the caregiver, and the design question that matters is whether the care agent extends her reach or quietly replaces her judgment. Those are different things, and a system that blurs them will look like help while functioning as substitution. The difference is the whole game.</p><p>I will refer to the care agent throughout as the agent, or by the name Dan and Lee actually use for it, which is CeCe. They did not name a companion. They sanded an acronym down into something they could say at the kitchen table, the way people always do with the infrastructure they live beside. The nickname is affection. The thing itself is not a character and does not speak in the first person. It is a layer of software stitched through the unglamorous parts of care, and the affection in the nickname does not change what it is.</p><div><hr></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JqdJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd3038a9-7259-4479-9357-1d662d51df50_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JqdJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd3038a9-7259-4479-9357-1d662d51df50_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!JqdJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd3038a9-7259-4479-9357-1d662d51df50_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!JqdJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd3038a9-7259-4479-9357-1d662d51df50_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!JqdJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd3038a9-7259-4479-9357-1d662d51df50_2752x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JqdJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd3038a9-7259-4479-9357-1d662d51df50_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dd3038a9-7259-4479-9357-1d662d51df50_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4725465,&quot;alt&quot;:&quot;Infographic&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/204337514?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd3038a9-7259-4479-9357-1d662d51df50_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Infographic" title="Infographic" srcset="https://substackcdn.com/image/fetch/$s_!JqdJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd3038a9-7259-4479-9357-1d662d51df50_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!JqdJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd3038a9-7259-4479-9357-1d662d51df50_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!JqdJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd3038a9-7259-4479-9357-1d662d51df50_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!JqdJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd3038a9-7259-4479-9357-1d662d51df50_2752x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Infographic created with NotebookLM</figcaption></figure></div><div><hr></div><h2>The math that makes the agent inevitable</h2><p>The case for assembling these fragments is not novelty. It is arithmetic.</p><p>The United States is older than its care systems were built for. People sixty-five and older were 17.3 percent of the population in 2022, and that share is projected to reach 22 percent by 2040.[^1] The people who hold the aging population together at home are mostly unpaid family members. As of 2025, 63 million Americans, nearly one in four adults, are family caregivers.[^2] They are not lightly tasked. Only 11 percent have received any medical training to assist with the activities of daily living they manage, and just over 20 percent have received formal training on the medical and nursing tasks they perform, even though more than half are doing exactly that kind of work.[^3] The formal workforce that is supposed to backstop them is thinning. HRSA projects that the supply of licensed practical nurses will meet only 64 percent of demand by 2037, with registered nurse shortages concentrated in nonmetropolitan areas.[^4]</p><p>So the pressure is real, and it points in one direction. When the human supply cannot scale, the system reaches for the layer that can. The agent is coming not because anyone romanticizes a glowing orb on the nightstand, but because the math leaves a gap and integration is cheaper than nurses.</p><p>The first place that logic already produced a result is worth noting, because it tells you how adoption actually travels. The breakthrough application of clinical AI so far has not been companionship. It has been documentation. Ambient AI scribes that draft clinical notes from the visit conversation were adopted because they solved an immediate, measurable pain. In one analysis, The Permanente Medical Group found that ambient scribes saved Northern California physicians the equivalent of 1,794 working days of documentation time over a single year, while improving the physician-patient interaction.[^5] Healthcare did not adopt that technology because it was futuristic. It adopted it because the pain of the existing workflow had become worse than the risk of the new tool.</p><p>That is the pattern to watch for the home. The agent that merely chats is a toy. The agent that lets an exhausted caregiver leave the house is infrastructure, and infrastructure gets built.</p><h2>Three extensions</h2><p>What CeCe actually does for Lee can be stated precisely, and the precision matters, because it draws the line that the rest of this argument depends on. CeCe extends Lee in three directions she could not reach on her own.</p><p>The first is perceptual. Lee, however devoted, cannot sense Dan&#8217;s overnight heart rate variability. She cannot feel the three pounds of fluid that precede a heart-failure decompensation by days. She cannot perceive the slow downward drift in overnight oxygen saturation or the subtle change in gait that arrives before Dan himself notices anything is wrong. These are not failures of attention or love. They are the limits of human senses. Continuous monitoring extends Lee into a physiological domain she never had access to, and that extension is the clearest, least arguable good the agent provides.</p><p>The second is temporal. Lee has to sleep. A sole caregiver who does not sleep does not last, which is why she eventually breaks and Dan eventually ends up somewhere that costs far more than her exhaustion ever did. CeCe watches at three in the morning so that Lee&#8217;s sleep is not a gap in Dan&#8217;s coverage. The agent extends her across the hours she cannot personally be vigilant.</p><p>The third is spatial, and it is the one that gave Lee back a piece of her life. CeCe lets her be in two places at once, which is the single thing a sole caregiver can never do and the lack of which turns the role into confinement. She takes the monitoring and the live video with her. She is, in a real and not merely sentimental sense, still home while she is at the pump. The tether becomes elastic instead of fixed.</p><p>Name those three plainly, because they are the affirmative case at full strength, and a piece that only catalogued risks would be dishonest. Perceptual, temporal, spatial. CeCe gives Lee back what the role had taken: senses she lacked, the hours she had to sleep through, and the ability to leave the house. The strongest version of the care agent is not the one that does the most. It is the one that extends the caregiver furthest into the places she could not otherwise reach.</p><p>Hold that line in mind, because everything that follows is about protecting it.</p><h2>The line between extension and substitution</h2><p>The same integration that produces the three extensions also produces a second set of capabilities, and they are not the same kind of thing.</p><p>When CeCe detects a pattern and recommends an action, when it drafts a clinical message, when it decides what becomes visible to whom, when it prioritizes one alert and suppresses another, it is no longer extending Lee&#8217;s reach. It is exercising judgment that someone, somewhere, used to exercise. That can be appropriate. It can also be the moment the agent stops augmenting a human and starts replacing one, and the failure mode is that the two kinds of capability arrive in the same product, marketed together, indistinguishable on the same screen.</p><p>So here is the test that should run through every design decision: does this capability extend the caregiver&#8217;s reach into something she could not otherwise do, or does it substitute for judgment she is positioned to exercise? The continuous three-a.m. watch is extension. An autonomous decision about whether Dan&#8217;s shortness of breath warrants escalation is substitution, unless Lee is the one deciding. The line is not always obvious, and that is exactly why it has to be drawn deliberately rather than left to whatever the vendor shipped.</p><p>Most writing about care AI skips this distinction because it skips Lee. It imagines the patient and the clinician and treats the caregiver as a notification endpoint, a phone that buzzes. But in Dan&#8217;s house, Lee is not the recipient of the system&#8217;s decisions. She is the one who makes them. She sets CeCe&#8217;s authorization levels. She decides what the agent may do on its own and what it must bring to her. When an alert fires, she is the one who directs the response, often by phone or text, in the middle of her own day. The agent reports to her before it reports to anyone clinical.</p><p>That configuration is the right one, and it is also the one nobody designed for. We have made an exhausted full-time caregiver the systems administrator of her father&#8217;s care, and the interface she administers it through was built by people who have never done her job.</p><h2>What the caregiver is actually being asked to do</h2><p>Consider the ordinary version of the hard case, which is more instructive than the dramatic one.</p><p>Lee is at the grocery store. Her phone flags an event: CeCe has noticed something in Dan&#8217;s vitals and is asking whether to escalate. Lee has to decide, on a four-inch screen, between the cereal and the checkout, whether to trust the all-clear, abandon her cart and drive home, or push the alert up to the clinic. She authorized CeCe&#8217;s escalation thresholds weeks ago, in a settings menu, under conditions she half remembers. Now the consequences of that configuration arrive while she is buying groceries.</p><p>Every governance question the system raises is present in that moment at once. Provenance: can Lee trust what the screen is showing her, and does she know whether the vitals are device-measured or inferred, fresh or stale? Bounded scope: does she understand what she is actually authorizing when she taps escalate, and how far that decision reaches? Workflow ownership: when she pushes the alert up, who receives it, and who owns it from there? And underneath all of it, the plain human-factors reality that life-affecting decisions are being made on a phone, between other obligations, by someone with no clinical training and no relief.</p><p>This is the cost that the spatial extension carries with it. The freedom to leave the house is real, but it is only real if Lee can trust the system she is relying on to make that freedom possible. A false alarm in the cereal aisle does not just annoy her. It re-tethers her, because a system that cries wolf is a system she cannot leave behind, and a caregiver who has left but cannot stop checking has not really left at all. The quality of the governance is what determines whether the spatial extension is freedom or just guilt with a video feed.</p><p>There is a specific design failure lurking here, and it has a name in the practitioner literature: supervisory burden. The labor of overseeing an automated process is itself work, and it is often invisible to the people who designed the automation. The agent that was supposed to lighten Lee&#8217;s load can quietly convert her from the person who watches Dan into the person who watches the thing that watches Dan. She used to know her father&#8217;s condition by being in the room. Now she configures thresholds, reviews flagged events, and adjudicates escalations. That is a different job, possibly a heavier one, and nobody asked her whether she wanted it. Whether CeCe lightens Lee&#8217;s day or simply changes the shape of its weight depends entirely on design choices she has no part in making.</p><h2>A well-written mistake travels farther</h2><p>Send Dan to the hospital and the provenance problem stops being abstract.</p><p>He arrives through the emergency department after two worsening days. The good version: CeCe has already assembled the week into something the admitting team can use. Weight gain, poor sleep, rising resting heart rate, a missed diuretic dose, climbing shortness of breath, and a note from Lee that says Dad says he is fine but he is not fine. The clinicians see a timeline instead of a fog. The agent has done real work, converting scattered living into a clinical narrative.</p><p>The bad version is just as easy to assemble from the same parts. CeCe imports a medication list that is three weeks stale. It confidently summarizes a symptom Dan denied but Lee mentioned, without marking which is which. A home blood-pressure reading gets labeled clinically verified when it came from a cuff of unknown calibration. The oxygen number, pulled from his late wife&#8217;s old pulse oximeter, arrives with no such caveat. And the discharge instruction reads fluently while describing a regimen Dan cannot actually follow at home.</p><p>In healthcare, fluency is not harmless. A well-written mistake travels farther than an awkward uncertainty, because confidence is contagious and a clean summary invites trust the underlying data may not deserve. This is why provenance is not a technical nicety. The clinician needs to see not just what CeCe knows but where each piece came from, how fresh it is, and whether it was device-measured, patient-reported, caregiver-entered, pulled from the record, or inferred by the model. The interface is where responsibility becomes visible or disappears. AI safety in the home will not be settled by model accuracy alone. It will be settled, or lost, in the handoff.</p><h2>The actuating edge</h2><p>The capabilities that act on the physical world are where the extension-versus-substitution line is brightest and most consequential, and the emergency scenario is the place to examine them, because it is the one everybody finds seductive.</p><p>Dan has a cardiac event. CeCe can do several things, and they are not equivalent.</p><p>It can dispatch EMS as fast as if Lee were standing in the kitchen. This is pure extension, and it is a genuine present-tense good. The spatial reach that let Lee go get gas means the emergency response no longer waits on her physical location. She can summon responders from the parking lot, looking at data that Dan in the room would not have given her. For a sole caregiver of a cardiac patient, collapsing the dispatch latency is not a convenience. It is minutes, and minutes are the whole game in a cardiac emergency.</p><p>It can transmit Dan&#8217;s history and current vitals to the responders. Useful, but this is where provenance has to ride along. A structured, source-tagged data packet that a paramedic can read and interpret is extension. A fluent synthetic voice asserting a clinical conclusion, the way the marketing demos love to stage it, is the audio version of the well-written mistake. The agent should hand responders flagged, sourced, uncertainty-marked data for a human to act on. It should not perform confidence it has not earned. CeCe does not speak Dan&#8217;s diagnosis. It shows its work.</p><p>And it can unlock the door. This is the sharpest one, and it deserves the discomfort, because physical actuation is a categorical step beyond monitoring and messaging. It is the moment the agent acts on the world rather than advising about it. Someone authorized that standing permission. Was it Lee, and did she understand she was granting a piece of software the ability to admit strangers to her father&#8217;s home during a medical event? What happens when it fires on a false reading and unlocks the door while Dan is in the shower and fine? Where is the audit trail? What is the override? None of this is an argument against the capability, which may well save Dan&#8217;s life. It is the design specification that has to exist before the capability ships, and the entire thesis of this piece is that the specification tends to arrive after the capability rather than before it.</p><p>Notice that none of these actions is speculative. Every one is possible with current technology. What is missing is not invention. It is the integration that binds them into one agent and the governance that says, in writing, who authorized what, under what conditions, with what record when it goes wrong.</p><h2>Where the substitution pressure is strongest</h2><p>The place autonomy creeps in first is wherever the human supply is thinnest, because that is where the temptation to let the agent absorb supervision is greatest.</p><p>In the home, that pressure runs through Lee. In a long-term-care facility, it runs through a staffing budget. The familiar pitch for the care agent in the facility is companionship: easing loneliness, running activities, facilitating virtual visits. That framing is the one to be most skeptical of. The honest read on the long-term-care setting is harder: it is the place where alert discipline and workforce-substitution pressure are most acute, because the facility is chronically understaffed and the economic incentive to let the agent stand in for human attention is strongest. A system marketed as reducing isolation can also be the mechanism by which a facility justifies fewer human visits. That is not a reason to keep the agent out. It is a reason to govern, explicitly, the line between extending the staff&#8217;s reach and replacing their presence.</p><p>This is how autonomy actually arrives in institutions. Not with a declaration, but by subtraction. A nurse stops checking one field. A queue gets trusted. A draft gets signed a little faster. A night-shift position goes unfilled because the monitoring is now &#8220;covered.&#8221; Each step is small and defensible on its own, and then a patient is harmed and everyone discovers that no one ever wrote down the moment when not showing something became a decision. Autonomy theater is the performance of human oversight after the human has become ornamental.</p><h2>Respite is the policy version of the same line</h2><p>There is a benefit structure that already encodes everything this piece is arguing, and it is worth naming because it shows the stakes are not hypothetical.</p><p>Full-time caregiving is a confinement, and the system knows it. The VA pays for respite care precisely because someone recognized that an unrelieved caregiver breaks, and a broken caregiver is both a human catastrophe and a fiscal one, since when Lee collapses, Dan goes to a facility that costs the VA far more than her relief ever did. Respite is the institutional admission that the tether is real and that relieving it is both decent and cheaper than the alternative.</p><p>CeCe does something adjacent to respite, and the adjacency is exactly where the danger lives. Respite is scheduled, staffed, and total: someone takes the role for a defined block so the caregiver can genuinely leave it. The agent is ambient and partial: it does not relieve Lee of the role, it loosens the tether enough that she can run to the store without arranging anything. Those are different goods, and a cost-conscious administrator who confuses them will reason that the caregiver who now has continuous monitoring needs fewer authorized respite hours. That is the substitution trap wearing a budget. The elastic tether is not the same as actual relief. CeCe lets Lee get gas. It does not let her sleep for a weekend while someone else carries the weight. A caregiver who is told the monitoring counts as support has been upgraded on paper and left more alone in fact.</p><p>This is the pace-gap that should worry a practitioner. The extension capabilities, perceptual, temporal, spatial, are arriving fast, assembled from parts already on the shelf. The policy understanding needed to keep them from being read as substitution for human support that does different work is lagging behind. Respite is the concrete case where that gap bites, and it bites the person least able to absorb it.</p><h2>What a serious care agent would require</h2><p>If the fragments are going to be integrated, and the arithmetic says they will be, then the integration should carry its governance with it rather than bolting it on after the first harm. Five requirements, each tied to a failure mode already visible in the scenario.</p><p>A bounded scope. The agent should not be &#8220;your AI doctor.&#8221; It needs declared functions, explicit clinical limits, defined escalation rules, and plain-language statements of what it may and may not do on its own. The caregiver authorizing it should be able to see the edges of its authority.</p><p>Data provenance. Every output needs a traceable source. Device-measured, patient-reported, caregiver-entered, record-derived, or model-inferred, marked as such, with its freshness attached. &#8220;CeCe says&#8221; is not a source, and a fluent summary that hides its provenance is a liability dressed as a convenience.</p><p>Alert discipline. More alerts are not more safety. The agent should be judged by whether the right person receives the right signal at the right time with the right action attached, not by the volume of its vigilance. An agent that floods Lee with low-value anomalies does not protect Dan. It re-tethers Lee and trains her to ignore the screen.</p><p>Workflow ownership. Every output needs a home and an owner. If CeCe escalates to the clinic, someone owns that queue. If it asks Lee to act, Lee needs to know whether she is being informed, advised, or made responsible. An alert that lands nowhere is not a safety feature. It is a liability artifact.</p><p>Post-deployment surveillance. Models drift, workflows shift, devices fail, and people adapt around tools in ways their designers never saw coming. A care agent has to be monitored after deployment not only for accuracy but for caregiver burden, alert fatigue, missed escalations, false reassurance, inequity across patients with different devices and bandwidth and language, and the silent workarounds that reveal where the design is failing in practice.</p><p>Notice that all five are really one principle applied five ways: keep the agent extending the human and govern hard against it replacing her judgment, and build the interface so the human can tell which is which and hold the line herself.</p><h2>The morning shift</h2><p>The promise is real, and it is smaller and truer than the robot nurse the marketing keeps selling.</p><p>Done well, the care agent lets Dan stay home longer. It lets Lee carry less of the invisible load and sleep through the night and leave the house. It lets clinicians see the pattern before the crisis instead of meeting it in the emergency department. It turns the patient portal from a mailbox into something closer to a care surface. None of that requires a breakthrough. It requires assembling tools that already exist into one agent that a caregiver can actually govern, and then doing the unglamorous work of governing it.</p><p>The future of the virtual care assistant will not be decided by how compassionate it can sound. It will be decided in the dull places: the consent screens, the audit logs, the escalation pathways, the provenance tags, the reimbursement rules, the staffing decisions a facility makes once it believes the monitoring is covered, and whether the discharge summary matches the life waiting at home. It will be decided by whether the people building these systems remember that, in the home, the user is the caregiver, and by whether the question that matters is whether the agent extends her or replaces her.</p><p>Lee can leave the house now. Whether that freedom is real or just a video feed she cannot look away from depends on choices that have not been made yet by people who have mostly not been thinking about her.</p><p>The dawn was the easy part. Now comes the morning shift.</p><div><hr></div><h2>Notes</h2><p>[^1]: Administration for Community Living, <em>2023 Profile of Older Americans</em> (U.S. Department of Health and Human Services, May 2024). People sixty-five and older numbered 57.8 million in 2022, representing 17.3 percent of the population, projected to reach 22 percent by 2040.</p><p>[^2]: AARP and the National Alliance for Caregiving, <em>Caregiving in the US 2025</em> (Washington, DC: AARP, July 24, 2025). The report documents 63 million family caregivers, nearly one in four adults, a 45 percent increase since 2015.</p><p>[^3]: AARP and the National Alliance for Caregiving, <em>Caregiving in the US 2025</em>. Eleven percent of caregivers have received medical training for activities of daily living, and just over 20 percent have received formal training on medical and nursing tasks, despite more than half performing such tasks.</p><p>[^4]: National Center for Health Workforce Analysis, <em>Nurse Workforce Projections, 2023&#8211;2038</em> (Health Resources and Services Administration, December 2025). HRSA projects LPN supply meeting 64 percent of demand by 2037 and a registered nurse shortage concentrated in nonmetropolitan areas (an 11 percent RN shortage in nonmetro areas versus 2 percent in metro areas by 2038).</p><p>[^5]: Aaron A. Tierney et al., &#8220;Ambient Artificial Intelligence Scribes: Learnings after 1 Year and over 2.5 Million Uses,&#8221; <em>NEJM Catalyst</em> (2025); The Permanente Medical Group, &#8220;Analysis: AI Scribes Save Physicians Time, Improve Patient Interactions and Work Satisfaction,&#8221; permanente.org, 2025. Ambient scribes produced documentation time savings of more than 15,700 hours, equivalent to 1,794 working days, over one year of use.</p><p>[^6]: U.S. Food and Drug Administration, &#8220;Artificial Intelligence-Enabled Medical Devices,&#8221; fda.gov. The FDA maintains a public list of authorized AI-enabled devices and notes it is developing approaches to identify devices incorporating foundation-model and LLM-based functionality.</p><p>[^7]: Office of the National Coordinator for Health Information Technology, &#8220;HTI-1 Final Rule,&#8221; healthit.gov. The rule establishes transparency requirements for AI and predictive algorithms in certified health IT, which supports care delivered by more than 96 percent of hospitals and 78 percent of office-based physicians, enabling clinical users to assess algorithms for fairness, appropriateness, validity, effectiveness, and safety.</p><p>[^8]: World Health Organization, <em>Ethics and Governance of Artificial Intelligence for Health: Guidance on Large Multi-Modal Models</em> (Geneva: WHO, January 2024).</p><p>[^9]: National Institute of Standards and Technology, <em>Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile</em>, NIST AI 600-1 (July 2024).</p><p>[^10]: Telehealth.HHS.gov, &#8220;Billing for Remote Patient Monitoring,&#8221; U.S. Department of Health and Human Services. Remote physiologic monitoring requires an established patient relationship, must monitor an acute or chronic condition, requires at least 16 days of data collection in a 30-day period for device-supply codes, requires patient consent, and requires physiologic data to be automatically and electronically transmitted to the billing practitioner. Note that for 2026, CMS introduced new codes permitting billing at lower data-collection thresholds, so the 16-day rule is in transition.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/p/the-caregiver-is-the-user?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/p/the-caregiver-is-the-user?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Human Factors Brief is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Default Is the Decision]]></title><description><![CDATA[The most consequential choice in your interface is usually the one nobody made.]]></description><link>https://johnwbrown.substack.com/p/the-default-is-the-decision</link><guid isPermaLink="false">https://johnwbrown.substack.com/p/the-default-is-the-decision</guid><dc:creator><![CDATA[John W Brown]]></dc:creator><pubDate>Tue, 07 Jul 2026 06:01:21 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!u98t!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd54f20-f717-42d9-8df3-16ef51e65a25_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!u98t!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd54f20-f717-42d9-8df3-16ef51e65a25_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!u98t!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd54f20-f717-42d9-8df3-16ef51e65a25_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!u98t!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd54f20-f717-42d9-8df3-16ef51e65a25_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!u98t!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd54f20-f717-42d9-8df3-16ef51e65a25_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!u98t!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd54f20-f717-42d9-8df3-16ef51e65a25_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!u98t!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd54f20-f717-42d9-8df3-16ef51e65a25_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/abd54f20-f717-42d9-8df3-16ef51e65a25_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2067149,&quot;alt&quot;:&quot;Computers&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/203718746?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd54f20-f717-42d9-8df3-16ef51e65a25_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Computers" title="Computers" srcset="https://substackcdn.com/image/fetch/$s_!u98t!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd54f20-f717-42d9-8df3-16ef51e65a25_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!u98t!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd54f20-f717-42d9-8df3-16ef51e65a25_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!u98t!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd54f20-f717-42d9-8df3-16ef51e65a25_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!u98t!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd54f20-f717-42d9-8df3-16ef51e65a25_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Image created with generative AI</figcaption></figure></div><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><p>Consider a case. It is a composite, not a specific incident, but it is drawn from a pattern common enough that anyone who has worked near clinical software will recognize it.</p><p>A health system rolls out a new order-entry module. Somewhere in the configuration, a default duration is set for a routine medication order. Fourteen days. No one in the room remembers choosing fourteen. It was probably carried over from the vendor&#8217;s reference build, or copied from a sister facility, or left at whatever the field happened to show when the screen was first assembled. It was not a clinical decision. It was a leftover.</p><p>Eighteen months later, an analyst pulls a report and notices that a striking share of these orders run for exactly fourteen days. Not because fourteen days is clinically indicated in those cases. Because fourteen days is what the screen offered, and clinicians, busy and trusting the system, accepted what the screen offered. The default did not advise. It did not flag. It simply sat there, pre-filled, and the path of least resistance did the rest.</p><p>No one decided that most patients should receive fourteen days. Yet that is what happened, thousands of times, because of a value nobody deliberately chose.</p><p>This is the part of interface design we discuss least and ship most carelessly. The default is the single most consequential decision in most screens, and it is routinely made by accident.</p><h2>Two mechanisms, not one</h2><p>The usual explanation for why defaults matter is inertia. People are busy. Changing a setting takes effort. The pre-filled value is the path of least resistance, so most people keep it. This is true, and it is well documented, but it is only half of what is happening.</p><p>The second mechanism is interpretation. A default is not neutral to the person looking at it. It reads as a recommendation. When a field arrives pre-filled, the user infers that someone, somewhere, with more context than they have, decided this was the right starting point. The pre-selected option carries an implicit message: this is the normal choice, the safe choice, the one most people make. The default speaks even when no one intended it to say anything.</p><p>These two mechanisms compound. Inertia means the user is unlikely to change the value. Interpretation means the user is unlikely to want to, because the value already feels endorsed. A team that thinks of defaults only as a convenience, a way to save a few clicks, has accounted for the first mechanism and missed the second entirely. They have set a value to reduce friction and inadvertently issued guidance.</p><p>This is why a default is not a starting point. It is a position. The interface is not asking a neutral question and waiting for an answer. It is making a suggestion and counting on the user to ratify it. Most of the time, the user does.</p><h2>What the research actually shows</h2><p>The behavioral literature on this is unusually clean. The canonical demonstration is organ donation. Countries with opt-out systems, where citizens are presumed willing donors unless they decline, show participation rates far above countries with opt-in systems that ask people to actively register. The populations are not meaningfully different in their underlying attitudes toward donation. The default is different, and the default governs the outcome.</p><p>Thaler and Sunstein gave this the name choice architecture: the recognition that there is no neutral way to present a choice, that every arrangement of options nudges behavior in some direction, and that the person designing the arrangement is therefore making a decision whether they intend to or not. The designer who declines to think about the default has not avoided shaping behavior. They have shaped it without looking.</p><p>The point for practitioners is narrower and more uncomfortable than the popular version of this idea. It is not simply that defaults are powerful. It is that the power operates regardless of intent. The fourteen-day order default shaped clinical behavior exactly as forcefully as a deliberately optimized default would have. The mechanism does not care whether you were paying attention. The leftover value and the carefully chosen value exert the same kind of pull. The only difference is that one of them was governed and the other was not.</p><h2>A default is a decision that deserves a process</h2><p>If the default carries this much weight, it should clear a deliberate review before it ships. Here are four questions that turn a default from a leftover into a decision.</p><p>First, what behavior does this default produce at scale? Not for the attentive user who reads every field, but for the rushed majority who accept what the screen offers. If most users keep the default, the default is effectively the policy. State the policy out loud and ask whether you would defend it as one.</p><p>Second, whose interest does that behavior serve? Sometimes the default that is easiest for the user and the default that is best for the business are the same value. Often they are not. When they diverge, the direction the team chooses reveals what the product is actually optimizing for, regardless of what the mission statement says.</p><p>Third, what does the default signal? Given that users read the pre-filled value as a recommendation, ask what recommendation you are making. If you would not put the suggestion in words and stand behind it, you should not encode it silently in a default.</p><p>Fourth, what is the recovery cost when the default is wrong? No single default fits every user. For the cases where it does not fit, how hard is it to notice that the value needs changing, and how hard is it to change it? A default with a low recovery cost is forgiving. A default that is hard to detect and hard to reverse is a trap, however reasonable it looked to the team that set it.</p><p>These questions do not take long to ask. The point is that they get asked at all, by someone with the authority to change the answer, before the value reaches production.</p><h2>The line between a good default and a dark pattern</h2><p>There is a temptation to treat dark patterns as a separate, more sinister category of design, the work of bad actors rather than ordinary teams. The mechanism says otherwise. The pre-checked marketing-consent box and the well-chosen clinical default use the identical machinery: inertia plus interpretation, the user&#8217;s tendency to accept and to read acceptance as endorsement. What separates them is not the technique. It is the answer to the second question. Whose interest does the behavior serve?</p><p>A default that serves the user, that produces the outcome an informed user would most likely have chosen for themselves, is good design. A default that serves the business at the user&#8217;s expense, that harvests consent or money or attention the user would have withheld if asked plainly, is a dark pattern. The two can look identical in the interface. They are distinguished entirely by who benefits from the inertia.</p><p>This is worth stating plainly because the legal line and the design line are not the same. Regulators have begun to recognize the mechanism. European data-protection rules, for instance, treat a pre-ticked consent box as invalid, on the grounds that silence and inertia do not constitute genuine agreement. That is the law catching up to what designers have known operationally for years. But a default can be entirely legal and still fail the whose-interest test. The practitioner standard should be higher than the regulatory one. A default that extracts value from the user&#8217;s inattention is a design failure even where no rule forbids it, and a team that ships it has chosen the business over the person on the other side of the screen.</p><h2>When the default is invisible even to the team</h2><p>The hardest version of this problem is now arriving, and it deserves a brief, honest accounting rather than a sweeping one.</p><p>Increasingly, defaults are not fixed values chosen once and set in configuration. They are generated, per user, by a model. The system observes behavior and pre-fills what it predicts this particular user wants. In principle this is the smart default taken to its logical end: a starting point tailored to the individual rather than averaged across everyone.</p><p>The difficulty is that the personalized default is no longer legible to the team that built the system. When the default was fourteen days, you could at least find the number in a configuration file and ask who chose it. When the default is whatever the model emits for this user in this moment, there is no single value to inspect, no field to point at, no one who can say what most users are being shown because every user is being shown something different. The four questions still apply, but the answers are now distributions rather than values, and the team&#8217;s ability to govern the default depends entirely on whether they built the means to observe it. Most have not. They have shipped a system that makes a consequential recommendation to every user and retained no way to see what it is recommending.</p><p>This is not an argument against personalized defaults. It is an argument that the governance burden rises with the sophistication of the mechanism, not the reverse. A leftover static default is at least inspectable. A leftover dynamic default is a recommendation engine no one is reading.</p><h2>The discipline</h2><p>None of this requires new theory. The behavioral science is decades old and the design implications follow directly from it. What it requires is treating a class of decision that currently gets made by default, in the idiomatic sense, as a decision that gets made on purpose.</p><p>The fourteen-day order was not a clinical judgment. It governed clinical behavior anyway. That gap, between the weight a default carries and the attention it receives, is where the work is. Close it by asking the four questions before the value ships. What does this produce at scale. Whose interest does it serve. What does it signal. What does it cost to recover when it is wrong.</p><p>A default is a decision. The only open question is whether anyone made it.</p><h2>Quick reference: the four questions</h2><ol><li><p>Behavior at scale. What happens when the rushed majority accept this value? If most keep it, the default is the policy. Would you defend it as one?</p></li><li><p>Whose interest. Does the resulting behavior serve the user or the business? When they diverge, the choice reveals what the product optimizes for.</p></li><li><p>Signal. Users read a default as a recommendation. Would you state this recommendation in words and stand behind it?</p></li><li><p>Recovery cost. When the default is wrong for this user, how hard is it to notice, and how hard is it to change? Low cost is forgiving. High cost is a trap.</p></li></ol><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/p/the-default-is-the-decision?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/p/the-default-is-the-decision?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p>]]></content:encoded></item><item><title><![CDATA[The Hidden Cost of Workarounds]]></title><description><![CDATA[Why informal user fixes are often the clearest evidence you have that a system is misaligned]]></description><link>https://johnwbrown.substack.com/p/the-hidden-cost-of-workarounds</link><guid isPermaLink="false">https://johnwbrown.substack.com/p/the-hidden-cost-of-workarounds</guid><dc:creator><![CDATA[John W Brown]]></dc:creator><pubDate>Tue, 30 Jun 2026 06:01:42 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!qcfu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F024351fe-28f5-4a76-b994-7c3bb0a24f23_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qcfu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F024351fe-28f5-4a76-b994-7c3bb0a24f23_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qcfu!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F024351fe-28f5-4a76-b994-7c3bb0a24f23_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!qcfu!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F024351fe-28f5-4a76-b994-7c3bb0a24f23_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!qcfu!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F024351fe-28f5-4a76-b994-7c3bb0a24f23_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!qcfu!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F024351fe-28f5-4a76-b994-7c3bb0a24f23_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qcfu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F024351fe-28f5-4a76-b994-7c3bb0a24f23_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/024351fe-28f5-4a76-b994-7c3bb0a24f23_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1886368,&quot;alt&quot;:&quot;Computer&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/202308068?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F024351fe-28f5-4a76-b994-7c3bb0a24f23_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Computer" title="Computer" srcset="https://substackcdn.com/image/fetch/$s_!qcfu!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F024351fe-28f5-4a76-b994-7c3bb0a24f23_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!qcfu!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F024351fe-28f5-4a76-b994-7c3bb0a24f23_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!qcfu!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F024351fe-28f5-4a76-b994-7c3bb0a24f23_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!qcfu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F024351fe-28f5-4a76-b994-7c3bb0a24f23_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Image created with generative AI</figcaption></figure></div><div><hr></div><p><em>The views expressed are my own and do not represent any organization I am affiliated with.</em></p><div><hr></div><p>A nurse on a busy medical-surgical floor scans a patient&#8217;s wristband, then reaches for the medication. The barcode will not read. The label is creased, or the scanner is slow, or the drug came up from pharmacy in a packaging the system does not recognize. The patient is due. Two other patients are also due. The nurse has done this a thousand times and knows the medication is right.</p><p>So she keeps a printed sheet of patient barcodes in her pocket and scans that instead.</p><p>From across the unit, that looks like a violation. It is a violation. It also defeats the entire purpose of a system installed to confirm the right drug reaches the right patient. But the behavior did not come from carelessness. It came from a scanner that fails often enough, at moments busy enough, that following the designed process would mean stopping care to troubleshoot hardware. The nurse is not ignoring the system. She is standing in the gap between the system as designed and the work as it actually arrives, and she is improvising a bridge.</p><p>This is a workaround, and it is one of the most studied behaviors in healthcare human factors. Koppel and colleagues catalogued exactly this pattern in their 2008 study of barcode medication administration, identifying fifteen distinct workarounds and thirty-one causes behind them.[1] The finding that matters is not that clinicians cut corners but that the corners were designed in.</p><p>For human factors, usability, and human systems integration, a workaround is not noise in the data. It is the data.</p><h2>The signal underneath the behavior</h2><p>It is tempting to treat every workaround as user failure, and sometimes that label fits. Workarounds can create risk, weaken controls, degrade data quality, and hide information the official system needs. They can normalize unsafe practice until the unsafe version becomes the trained-in default.</p><p>But the behavior almost always carries a second message. A workaround is an adaptation to friction, a way of completing work when the approved path is too slow, too rigid, too confusing, or poorly matched to the task as it presents itself. People rarely invent workarounds for the pleasure of breaking process. They invent them because something in the work system makes the official process hard to follow under real conditions.</p><p>That something is usually upstream of the user. A poorly designed interface. A policy written for ideal conditions. A workflow that assumes uninterrupted attention. Missing information, duplicated entry, weak integration between systems, alerts that fire so often they become wallpaper, a required field that demands data the user will not have until three steps later. The workaround is the visible behavior. The cause is somewhere the user did not put it.</p><p>This is the core move of the sociotechnical tradition in patient safety, and it is worth naming the lineage because the piece stands on it. Carayon&#8217;s SEIPS model framed the work system, not the individual, as the unit of analysis, and Holden and Carayon later distilled that model into tools a practitioner can pick up without prior training.[4][5] Blijleven&#8217;s review of electronic health record workarounds went further, building a dedicated analysis framework on the premise that a workaround is a valuable point of departure for improving design.[3] Study the behavior as a response to system conditions, the tradition says, not as a character flaw.</p><h2>The formal workflow is cleaner than the work</h2><p>Most organizations carry two workflows. The formal one lives in process maps, standard operating procedures, training decks, and system requirements. It defines expected practice, supports accountability, and makes work teachable and auditable. It is necessary.</p><p>It is also tidier than anything that happens on the floor. Real work includes incomplete information, urgent exceptions, equipment that fails, staffing that varies by shift, and constant pressure to keep moving. Users spend much of their day reconciling what the process expects with what the environment allows.</p><p>The mismatches are specific and they repeat. A design team assumes a user will finish Step 1 before starting Step 2, but in the field the information for Step 4 arrives first while Step 2 waits on someone else, so Step 1 gets documented last because the software demands a linear sequence the work does not have. A policy assumes one person owns a task, while in practice three people carry it across a shift change. A system asks for information at the start of a workflow that nobody can know until the end. Each time, the user improvises: a placeholder, an &#8220;unknown&#8221; entry, a note written somewhere the system cannot see, a phone call that bypasses a step that looked pointless in the moment.</p><p>The workaround may carry real risk. It is also telling the organization something it paid to learn and would otherwise miss. The system does not fit the work as performed. Debono&#8217;s scoping review of nurses&#8217; workarounds found precisely this doubled nature: the behaviors enable care and compromise it at the same time, and they reveal information about clinical work that the formal record does not capture.[2]</p><h2>Not all workarounds are the same</h2><p>The most common analytical mistake is treating workarounds as one category. They are not, and the response to one should not be the response to all. Five rough types cover most of what practitioners actually find.</p><p>Safety-preserving workarounds let a user deliver needed care despite a system outage, a missing field, or a bottleneck that would otherwise block appropriate action. The workaround is compensating for a system that fails the patient if followed to the letter.</p><p>Cut redundant entry, excessive navigation, or duplicate approvals, and you have an efficiency-preserving workaround. These usually point straight at waste in the designed process.</p><p>Information-preserving workarounds are the private spreadsheets, the side notes, the local trackers that exist because the official system does not capture what the team needs to coordinate. The shadow record is doing real work the system of record declined to do.</p><p>When users check a second source because the official output is incomplete, outdated, or simply not believed, the workaround is trust-preserving, and it is a verdict on the system&#8217;s credibility.</p><p>Risk-generating workarounds are the genuinely dangerous ones: bypassed safety checks, shared credentials, unapproved channels, data created outside the system of record. These need attention for the opposite reason from the others.</p><p>The point of the taxonomy is not to excuse the behavior but to refuse a single reflex. A workaround that keeps a patient safe under poor system conditions is not the same problem as one that strips out a safeguard, even when both technically violate policy. Classify first, then respond.</p><h2>The cost that hides</h2><p>Workarounds feel efficient to the person using them, which is part of why they persist and part of why they are dangerous. Local efficiency can carry a system-level cost the local user never sees.</p><p>A private spreadsheet helps one team manage its load while quietly creating version-control problems, privacy exposure, and dependencies no one has mapped. A paper note rescues a user&#8217;s memory in the moment but is not there for the next person in the workflow. A copied template speeds documentation and propagates last month&#8217;s error into this month&#8217;s chart. A dismissed alert saves ten seconds and removes a safeguard on the rare day it mattered.</p><p>The deeper cost is concealment. When skilled users consistently absorb the friction of a bad design, leadership sees a system that appears to work and underinvests in fixing it. The reliability is real, but it is being manufactured by invisible human effort rather than by the system. New users, overloaded users, and users in less supported settings cannot generate that effort on demand. The result is a system that depends on invisible user effort, and a system that depends on invisible user effort is not as stable as it looks, sitting one staffing shortage away from revealing what the workarounds were holding up.</p><p>This is the Safety-II insight that Hollnagel pressed: performance succeeds and fails for the same reasons, and the everyday adjustments that make work succeed are usually invisible precisely because they succeed.[7] Dekker&#8217;s reframing of human error runs parallel. The behavior that looks like a deviation from above often looks like the only reasonable path from inside the work.[8]</p><h2>Finding what users will not volunteer</h2><p>If workarounds are evidence, the practitioner&#8217;s job is to collect them well, and they are unusually hard to collect. Users underreport them, often without meaning to. The behavior has become so routine they no longer notice it. They describe it as just how things are done. They assume everyone already knows the official process does not work. Sometimes they stay quiet because they expect blame rather than repair.</p><p>This means the direct question fails. Asking &#8220;Do you use any workarounds?&#8221; returns very little, because the word implies misconduct and the behavior has stopped feeling like a choice. Concrete questions work better because they treat the behavior as part of the work rather than a confession. What do you do when the system does not have the information you need? Where do you track the things the official tool cannot hold? What do you complete outside the system, and when do you delay documentation? What do experienced users do that new ones have not learned yet? What would break tomorrow if your local spreadsheet disappeared?</p><p>Even good questions miss what users have automated past noticing, which is why observation outperforms interviews. Watching the work reveals what people skip, repeat, double-check, write down, or route around, and it captures the timing that makes a workaround legible. A behavior that sounds like negligence when described often makes complete sense when seen in context. The user who appears to skip a required field turns out to be working at a point where the required information does not yet exist. The user who appears to duplicate data is feeding a downstream team a format the official system never provided. The behavior is visible. The reason is contextual. Human factors work has to capture both, or it mistakes the symptom for the disease.</p><h2>AI will industrialize the workaround</h2><p>Everything above predates the current wave of AI tools, but AI is about to make workarounds more common, more consequential, and harder to see. Some will be familiar in a new costume. Users will paste model outputs into unofficial files, build private prompt libraries, check generated answers against other systems, and pass examples through back channels the way they have always shared shortcuts.</p><p>Others will be genuinely new in their risk profile. Users will paste information into tools never approved for that data. They will lean on AI summaries without reading the source. They will stitch together unofficial automation chains because the sanctioned system is too slow. They will generate documentation that looks complete and was barely verified. None of this is best understood as recklessness. It is the same gap-bridging behavior the nurse with the printed barcodes was doing, now with a far more capable and far less predictable tool in hand.</p><p>The pattern to watch is specific. AI often reduces one burden while quietly creating another, and the workaround grows in the new gap. A tool that hands users an answer while leaving them quietly unsure of it will breed private checking rituals: a second screen kept open, a spreadsheet of known-good cases, a habit of asking a trusted colleague before trusting the model. A tool that is prohibited without a workable alternative will push the work into the dark, outside any visibility at all. Governance that moves slower than the workload pressure will find informal AI use has already become standard practice by the time policy arrives. This is the predictable shape of the next several years, and the organizations that treat these behaviors as evidence rather than as violations will be the ones that learn where their AI tools actually fail the work.</p><h2>What the analysis should capture, and what to do with it</h2><p>The most useful posture toward a workaround is diagnostic. It is a symptom, and the symptom may point to poor usability, weak integration, excessive workload, a policy written for the wrong conditions, missing features, unreliable data, or unclear ownership. The work is to read the symptom accurately, which means documenting more than the behavior itself.</p><p>A workaround analysis worth acting on records what the user does and when, the task or constraint it addresses, who relies on it and how often, the tools and information involved, the part of the official workflow it bypasses, the risk it creates or reduces, and what would happen if it vanished overnight. That last question is the most revealing. A behavior used rarely in a low-stakes situation does not deserve the response owed to one performed daily around a safety control, and documenting the difference prevents both overreaction and neglect.</p><p>This is also where most organizations make their defining mistake. The common response to a workaround is to tell users to stop. Sometimes that is correct, especially when the behavior creates immediate safety, privacy, security, or compliance risk. But stopping the behavior does nothing to the condition that produced it, and the condition is still there the next morning. If the workflow remains slow, incomplete, or misaligned, users will build a new workaround, usually one harder to see than the one before.</p><p>The better response has two parts. Assess the risk of the behavior, deciding whether it must stop now, be controlled temporarily, or be studied further. Then find the system condition underneath it. Ask what need it meets, what burden it removes, what failure it compensates for. None of this excuses unsafe behavior. It makes the correction effective rather than cosmetic, because a workaround is almost never solved by removing the behavior alone. It is solved by fixing the mismatch.</p><p>Some workarounds should be eliminated. Some should be redesigned. And some should be formalized, because the people closest to the work occasionally discover a better workflow than the official one: a cleaner sequence, a smarter coordination practice, a more realistic way to handle exceptions. Organizations should be slow to crush these simply because they were not planned centrally. The test is whether the adaptation is safe, reliable, scalable, auditable, and aligned with the larger system. If it is useful but risky, redesign it. If it is useful and safe, support it. If it is useful only because the official tool is broken, fix the tool. Workarounds are sometimes prototypes built under pressure, and a prototype deserves evaluation before it becomes permanent.</p><h2>The real system speaks through its workarounds</h2><p>In most organizations, the people closest to the work are quietly adjusting the system all day to keep it functioning. They smooth broken handoffs, translate mismatched terminology, reconcile data the systems will not reconcile themselves, and teach new staff the unofficial way to survive the workflow. This is operational resilience, and it is genuinely valuable. It is also a warning, because a system that leans this hard on informal adaptation is more fragile than its dashboards suggest.</p><p>So the question to bring to any workflow, product, or AI-enabled tool is not how to stop users from improvising. The better question is what condition made the improvisation worth it. Find the workarounds, observe them in context, classify them by type and risk, and trace each one back to the need it serves. Then decide, deliberately, whether to eliminate, redesign, formalize, or support it.</p><p>A workaround is not just a deviation from the process. It is a message from the real system, sent by the people who know it best, about where the design has failed the work. The only mistake is refusing to read it.</p><h2>References</h2><p>[1] Ross Koppel et al., &#8220;Workarounds to Barcode Medication Administration Systems: Their Occurrences, Causes, and Threats to Patient Safety,&#8221; <em>Journal of the American Medical Informatics Association</em> 15, no. 4 (2008): 408&#8211;423.</p><p>[2] Deborah S. Debono et al., &#8220;Nurses&#8217; Workarounds in Acute Healthcare Settings: A Scoping Review,&#8221; <em>BMC Health Services Research</em> 13 (2013): 175.</p><p>[3] Vincent Blijleven, Florian Hoxha, and Monique Jaspers, &#8220;Workarounds in Electronic Health Record Systems and the Revised Sociotechnical Electronic Health Record Workaround Analysis Framework: Scoping Review,&#8221; <em>Journal of Medical Internet Research</em> 24, no. 3 (2022): e33046.</p><p>[4] Pascale Carayon et al., &#8220;Work System Design for Patient Safety: The SEIPS Model,&#8221; <em>Quality &amp; Safety in Health Care</em> 15, suppl. 1 (2006): i50&#8211;i58.</p><p>[5] Richard J. Holden and Pascale Carayon, &#8220;SEIPS 101 and Seven Simple SEIPS Tools,&#8221; <em>BMJ Quality &amp; Safety</em> 30, no. 11 (2021): 901&#8211;910.</p><p>[6] AHRQ Patient Safety Network, &#8220;Workarounds and Resiliency on the Front Lines of Health Care.&#8221;</p><p>[7] Erik Hollnagel, <em>Safety-I and Safety-II: The Past and Future of Safety Management</em> (Farnham: Ashgate, 2014).</p><p>[8] Sidney Dekker, <em>The Field Guide to Understanding &#8216;Human Error,&#8217;</em> 3rd ed. (Farnham: Ashgate, 2014).</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/p/the-hidden-cost-of-workarounds?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/p/the-hidden-cost-of-workarounds?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p>]]></content:encoded></item><item><title><![CDATA[AI Does Not Remove Cognitive Load, It Moves It]]></title><description><![CDATA[What a decade of mammography automation should have taught us about AI at work]]></description><link>https://johnwbrown.substack.com/p/ai-does-not-remove-cognitive-load</link><guid isPermaLink="false">https://johnwbrown.substack.com/p/ai-does-not-remove-cognitive-load</guid><dc:creator><![CDATA[John W Brown]]></dc:creator><pubDate>Tue, 23 Jun 2026 06:01:28 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!h7hA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F267ad7f2-4953-4b2a-8809-e636c42ae16b_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!h7hA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F267ad7f2-4953-4b2a-8809-e636c42ae16b_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!h7hA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F267ad7f2-4953-4b2a-8809-e636c42ae16b_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!h7hA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F267ad7f2-4953-4b2a-8809-e636c42ae16b_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!h7hA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F267ad7f2-4953-4b2a-8809-e636c42ae16b_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!h7hA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F267ad7f2-4953-4b2a-8809-e636c42ae16b_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!h7hA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F267ad7f2-4953-4b2a-8809-e636c42ae16b_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/267ad7f2-4953-4b2a-8809-e636c42ae16b_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1521828,&quot;alt&quot;:&quot;Imagery&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/202304186?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F267ad7f2-4953-4b2a-8809-e636c42ae16b_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Imagery" title="Imagery" srcset="https://substackcdn.com/image/fetch/$s_!h7hA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F267ad7f2-4953-4b2a-8809-e636c42ae16b_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!h7hA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F267ad7f2-4953-4b2a-8809-e636c42ae16b_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!h7hA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F267ad7f2-4953-4b2a-8809-e636c42ae16b_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!h7hA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F267ad7f2-4953-4b2a-8809-e636c42ae16b_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Image created with generative AI</figcaption></figure></div><div><hr></div><p><em>The views expressed are my own and do not represent any organization I&#8217;m affiliated with.</em></p><p>AI tools are often introduced with a simple promise: they will reduce work.</p><p>Sometimes they do. They can summarize documents, draft text, classify records, generate code, retrieve information, identify patterns, and reduce the manual effort required to produce an initial output.</p><p>But in complex work environments, the story is rarely that simple.</p><p>AI does not always remove cognitive load. Often, it moves cognitive load from one part of the task to another.</p><p>A person who once wrote a document now reviews an AI-generated draft. A person who once searched for information now evaluates a generated summary. A person who once compared data points now decides whether a recommendation is reasonable. The work did not vanish. It changed shape.</p><p>That may sound like a reduction in burden, but supervision is not effortless. Verification is not passive. Trust calibration is work.</p><p>This matters for human factors, usability, and human systems integration because AI changes the shape of the task. If teams only measure whether AI reduces production time, they may miss the new work created around review, interpretation, accountability, and recovery.</p><p>The real question is not simply, did AI make the task faster? The better question is, where did the human work go?</p><p>We are not the first profession to ask it. Radiology asked it twenty years ago, ran the experiment at national scale, and got an answer worth studying before the rest of us repeat the mistake.</p><h2>The experiment we already ran</h2><p>Computer-aided detection for screening mammography is the closest thing we have to a controlled, decade-long, population-scale trial of putting a pattern-recognition machine next to a trained expert and asking the expert to verify its output.</p><p>The setup was exactly the one most organizations are now building for knowledge work. The FDA approved CAD for mammography in 1998. The Centers for Medicare and Medicaid Services increased reimbursement for it in 2002. The tool spread quickly, until it was used for most screening mammograms in the United States, at a cost of more than four hundred million dollars a year. The machine flagged suspicious regions on the image. The radiologist remained responsible for the final read. A human was in the loop, by design and by regulation.</p><p>Then someone measured the whole system rather than the tool.</p><p>In 2015, Constance Lehman and colleagues at the Breast Cancer Surveillance Consortium published a study in JAMA Internal Medicine comparing the accuracy of digital screening mammography interpreted with CAD against mammography interpreted without it. The dataset was not a laboratory sample. It covered 495,818 mammograms read with CAD and 129,807 read without it, across 323,973 women. The finding was blunt: screening performance was not improved with CAD on any metric the study assessed, and CAD did not improve individual radiologist accuracy [1].</p><p>A tool that worked, in the narrow sense, for over a decade. A tool that reliably produced output. A tool that regulators approved and insurers paid for. And the human-AI system as a whole detected no more cancer than the human alone.</p><p>The reason is the entire subject of this essay. The output was nearly free. The work the output created was not, and the system did not support that work well.</p><p>Earlier observer studies show the mechanism in close detail. When CAD correctly marked a cancer that radiologists had missed, the radiologists still failed to act on the correct prompt in the large majority of those cases [2]. The mark was there. The information was present on the screen. But noticing a mark and integrating it into a confident clinical judgment are different cognitive acts, and the second one is the hard one. Other work on CAD found something more uncomfortable still: when the system was wrong, readers missed more than they would have caught working alone, because the absence of a mark was quietly read as reassurance [2].</p><p>This is the pattern. The tool shifted the radiologist&#8217;s task from searching to verifying. Verifying a confident machine is its own skill, performed under its own pressures, and it is not the skill the tool was marketed to replace. Nobody designed for the new task because everybody assumed the tool had removed work rather than moved it.</p><p>We are now installing the same arrangement across law, finance, medicine, software, and administration, and we are measuring it the way the early CAD adopters did: by whether the tool produces output, not by whether the human-AI system produces better work.</p><h2>The visible task and the hidden task</h2><p>Many AI tools reduce the visible task. A draft appears faster. A summary arrives sooner. A recommendation is generated automatically. A table is populated. A message is categorized.</p><p>This creates an immediate impression of efficiency. The user no longer has to start from a blank page, read every document, or manually perform every classification.</p><p>But the visible task is not the whole task.</p><p>The hidden task may now include checking the output for accuracy, identifying missing context, detecting inappropriate assumptions, deciding whether the system&#8217;s confidence is justified, reconciling the output against other information, documenting the basis for acceptance or rejection, and correcting the result when it fails.</p><p>In some cases, the human does less typing but more judging. Less searching but more validating. Less generating but more monitoring.</p><p>That is not automatically bad. Shifting work can help when the new task is easier, safer, more consistent, and better supported. But shifting work without recognizing it creates risk. A poorly designed workflow may save time at the front end and add burden downstream, make routine cases faster while making exceptions harder to detect, or reduce clerical effort while increasing cognitive responsibility.</p><p>This is why &#8220;AI saves time&#8221; is too blunt as a usability claim. The more useful claim is conditional: AI saves time when the system reduces total work across the task, including verification, correction, coordination, and recovery. CAD failed that conditional test for a decade before anyone added up the total.</p><h2>Cognitive load does not disappear</h2><p>To see why the radiologists struggled, it helps to be precise about what was being moved. Cognitive load refers to the mental effort required to process information, solve problems, decide, remember, and act. The basic concern is practical: people have limited attention and working memory, and poorly designed systems waste those resources [3].</p><p>AI can reduce some forms of cognitive load. It can organize large amounts of information, convert unstructured material into usable form, suggest next steps, and help users get past a blank page. In the right context, that is valuable.</p><p>But AI can also create new cognitive demands. The user may need to understand what the system did, what it did not do, what data it relied on, what it may have missed, whether the output is complete, whether it fits the current case, and whether it can be trusted for the decision at hand.</p><p>This is especially important when the output appears polished. A confident draft, summary, or recommendation can look finished even when it contains omissions or errors. Fluency conceals uncertainty. The user must then perform a difficult kind of review: not just editing what is present, but noticing what is absent. The radiologist who failed to act on a correct mark was not lazy. Noticing what a confident system left out, or got wrong, is the hardest perceptual and cognitive work in the whole task.</p><p>It is not a minor task. It is the work.</p><h2>From production burden to verification burden</h2><p>One of the most common AI shifts is from production burden to verification burden.</p><p>Before AI, a user often produced an artifact directly. That could mean writing a note, preparing a report, searching for policy language, or reviewing a case manually. With AI, the user receives a first draft or recommendation. The production burden drops. But now the user must verify the output, and verification is not a single act. The user has to judge whether the output is factually correct, whether important context is missing, whether it fits the intended purpose, whether the system is overgeneralizing from incomplete information, whether a plausible-sounding recommendation is actually wrong, and whether accepting it creates downstream risk.</p><p>This burden may be manageable for expert users. It may be much harder for novice users, overloaded users, or users working under time pressure. It may also be harder when the system does not explain its basis, show source material, indicate uncertainty, or support easy comparison with the original information.</p><p>In that case, the tool has not eliminated work. It has made a difficult review task appear simple. That is a human factors problem.</p><h2>A second case, closer to the average desk</h2><p>Mammography is a high-stakes example with a clean dataset. Most knowledge work is lower-stakes and messier, but the mechanism is identical. Consider a composite case, drawn from common patterns rather than any specific individual or organization.</p><p>A mid-sized insurer introduces an AI tool that drafts first-pass responses to customer coverage questions. A claims specialist who used to research and write each response now receives a generated draft and approves, edits, or rejects it. The vendor demonstration measures one thing: average handling time per inquiry, which drops by roughly forty percent. On that metric, the rollout succeeds.</p><p>What the metric does not capture is where the specialist&#8217;s work went. The draft is fluent and formatted like the specialist&#8217;s own writing, so its errors do not announce themselves. Most drafts are correct, which is the problem, because a long run of correct drafts trains the specialist to skim. The few drafts that cite a superseded policy version, or quietly omit an exclusion that applies to this customer&#8217;s plan, look exactly like the correct ones. The specialist is now doing low-prevalence error detection, the same task that defeated the radiologists, under a quota that assumes the tool made the job easier.</p><p>Six months in, handling time is still down, the specialist reports that the tool is helpful, and a small but rising number of responses are going out with confident, well-formatted, wrong answers that nobody upstream is positioned to catch. The tool did exactly what it promised. The system around it was never redesigned for the task the tool created.</p><p>This is a hypothetical, and it is deliberately ordinary. You do not need a cancer diagnosis on the line for the dynamic to bite. You need only a fluent machine, a human held responsible for its output, a workflow that rewards speed, and an evaluation that measures production instead of total burden. Those four conditions describe a large share of the AI deployments now underway.</p><h2>The problem of polished uncertainty</h2><p>Human beings are sensitive to presentation. Output that is organized, fluent, and confident tends to feel more credible than output that is hesitant or messy.</p><p>Generative AI creates a special version of this problem. It can produce highly readable material even when the underlying answer is incomplete, poorly grounded, or wrong. A human reviewer must separate fluency from reliability, and that is not always easy. A rough draft written by a person usually carries visible signs of incompleteness: notes, gaps, questions, awkward transitions, uncertain phrasing. A generated draft hides those seams. It can look complete before it has earned that confidence.</p><p>This creates what might be called polished uncertainty. The uncertainty has not disappeared. It has been wrapped in professional-looking output.</p><p>For users, this changes the review task. They cannot only ask whether the text reads well. They have to ask whether it is true, whether it is complete, whether it is appropriate, what evidence supports it, and what the system left out. Those questions take effort. If the interface does not support them, the user supplies that effort alone, or skips it.</p><h2>Automation changes attention</h2><p>Automation does not merely perform tasks. It changes what people attend to.</p><p>When a person performs a task manually, attention is distributed across the steps of the work. The user sees the material, makes small judgments, notices anomalies, and builds a sense of the case through direct interaction. When AI performs part of the task, the human may enter later in the process, receiving a result rather than constructing it. That can be efficient, but it can also reduce situation awareness. The user may know what the system recommends without understanding how the recommendation emerged, see a summary without knowing what was excluded, or approve a draft without having engaged the underlying details.</p><p>This is not a new concern. Automation research has long recognized that people become overreliant on automated aids and less prepared to intervene when those aids fail [4], [5]. The out-of-the-loop problem was named in aviation and process control decades ago. The specific tools are new. The design challenge is familiar. The human must remain meaningfully engaged with the work, not placed at the end of the pipeline as a formal approver.</p><h2>Human-in-the-loop is not enough</h2><p>Many AI systems are defended by saying there is a human in the loop. The phrase can be useful, but it can also obscure more than it explains.</p><p>A human in the loop may be an active decision-maker. They may also be a rushed reviewer, a rubber stamp, a downstream recipient, an exception handler, or a person held accountable for an output they had little practical ability to evaluate. The radiologists in the CAD studies were unambiguously in the loop. Regulation required it. It did not help, because being in the loop and being equipped to catch the machine&#8217;s errors are not the same condition.</p><p>The key question is not whether a human appears somewhere in the workflow. It is whether the human has enough information, time, authority, skill, and interface support to perform the role assigned to them. If the system expects the user to catch errors, the interface has to make error detection realistic. If the user is expected to calibrate trust, the system has to communicate uncertainty and limits. If the user is expected to override the system, the workflow has to make override feasible. If the user is accountable for the final action, the system has to support traceability and review.</p><p>Otherwise, human-in-the-loop becomes a label for accountability without control.</p><h2>The new work of trust calibration</h2><p>Trust is not a simple target. The goal is not maximum trust. The goal is appropriate trust.</p><p>Overtrust leads users to accept poor outputs, ignore contradictions, or stop checking, which is the CAD failure exactly. Undertrust leads users to reject useful support, duplicate effort, or abandon the tool. Both are failures.</p><p>AI systems need to help users calibrate. Users need to understand what the system is good at, where it is limited, when it is uncertain, what evidence it used, and which cases demand extra caution. This is not only an ethics issue. It is a usability issue. A system that delivers every answer in the same confident tone makes calibration harder, and one that hides its source material pushes users toward either blind acceptance or blanket skepticism.</p><p>Good design helps users ask better questions of the output: what is this based on, what evidence supports it, what alternatives were considered, what information was unavailable, and what would make this recommendation wrong. Those questions should not depend entirely on the user&#8217;s memory or skepticism. The system should support them.</p><h2>AI review fatigue</h2><p>As AI systems multiply, another issue will grow: review fatigue.</p><p>If every tool generates drafts, summaries, alerts, classifications, and recommendations, users may spend an increasing share of the day reviewing machine output. Each individual review seems manageable. In aggregate, the burden becomes substantial, especially in environments already crowded with alerts, dashboards, messages, and forms. The user may not experience AI as one helpful assistant. They may experience it as another stream of things to check.</p><p>Review fatigue leads to shallow checking, missed errors, overreliance, irritation, avoidance, and informal workarounds. Some users accept outputs because the cost of careful review is too high. Others stop using the tool because the review burden cancels the promised efficiency.</p><p>The practical lesson is simple. AI-generated output should be treated as workload until proven otherwise. It is not free just because the system produced it automatically.</p><h2>The importance of role clarity</h2><p>AI changes roles, and that change should be explicit. Is the user asking the system for help? Is the system making a recommendation, drafting something for human revision, or making a classification that drives downstream action? Is the human expected to approve, edit, override, monitor, or explain the output?</p><p>These roles are different, and they need different interface support. A reviewer needs access to source material. An approver needs confidence information and traceability. An editor needs clear boundaries between human and machine-generated content. A supervisor needs exception visibility. A downstream recipient needs to know what role AI played in the information they are now relying on.</p><p>If the design does not clarify the human role, the organization will still assign responsibility after something goes wrong. That mismatch creates risk for both the user and the people affected by the system. Role clarity belongs in AI workflow design from the beginning, not in the incident report afterward.</p><h2>What usability testing should measure</h2><p>If AI shifts cognitive load, usability testing has to measure more than speed and satisfaction.</p><p>A test that only asks whether users completed the task can miss the central issue. Users may finish quickly while failing to notice an error, call the tool helpful while misunderstanding its limits, or accept a recommendation because it sounds reasonable rather than because they verified it. The CAD rollout would have passed a naive usability test for years.</p><p>AI usability testing should therefore observe verification behavior, error detection, trust calibration, and recovery. Did users notice incorrect or incomplete output? Did they recognize when the system was uncertain? Could they trace the output back to supporting information? Did they know when to accept, reject, edit, or escalate? Did the output shift their confidence appropriately, or just shift it? Did they over-rely under time pressure? Did the workflow support correction? Did total task burden actually fall, or did it just move to a later step?</p><p>These questions move evaluation from surface usability to operational usability. That is the level at which AI systems should be judged.</p><h2>Design principles for reducing transferred burden</h2><p>If AI shifts work, designers should make the new work visible and supportable. None of what follows is novel; it converges with existing human-AI design guidance [6], [7], [8]. Several principles follow.</p><p>Show the basis for the output, so users can inspect source material, inputs, or evidence in a form appropriate to the task. Communicate uncertainty, because not every output deserves the same confidence, and the system should help users separate routine, well-supported results from uncertain or high-risk ones. Design for comparison, so that verifying a summary, recommendation, or classification against the original is efficient rather than a separate research project. Support exception handling, because AI tends to perform well on common cases and worse on edge cases, and the design should make exceptions visible with clear paths for escalation, correction, or override. Preserve situation awareness, so users are not reduced to final-stage approvers who see only the answer. Measure downstream burden, since a tool that saves time for one role may create work for another, and evaluation should follow the output through the whole workflow. Finally, define accountability honestly: if a human is responsible for an AI-supported decision, the system has to give that human meaningful control.</p><p>These principles are not exotic. Most of them are what the mammography workflow lacked. CAD marked the image but never helped the radiologist weigh the mark, compare it against their own read on equal footing, or treat a missing mark as anything other than reassurance. The design assumed the hard problem was detection. The hard problem was integration.</p><h2>What the radiologists should have told us</h2><p>AI can be useful. It can reduce effort, improve access to information, support drafting, accelerate review, and help people manage complexity. But it is not automatically a cognitive load reducer. In many real systems it changes the user&#8217;s task from doing to checking, from searching to judging, from producing to supervising, and from acting directly to managing uncertainty. That shift may be valuable. It must be designed.</p><p>This is why the first step is never simply adding AI to a workflow. It is understanding the workflow well enough to know what burden the tool will shift, who will inherit it, and whether anything supports them. That means treating the user as an operator inside a system, not a satisfied end user of a feature. For human factors and HSI practice, the essential question is not whether AI can perform a function. It is whether the human-AI system supports safe, effective, understandable, and accountable work. A useful AI system should not simply generate output. It should support the human work required to evaluate that output. It should make uncertainty visible, make verification feasible, preserve situation awareness, clarify roles, and reduce total burden rather than relocate it.</p><p>The promise of AI is not that humans will stop thinking. In high-stakes systems, the promise should be that humans think better, with better support, better context, and better tools for knowing when the system is wrong.</p><p>So the practical question is worth asking before the rollout, not after the audit. When the tool produces its output, where does the cognitive work go? If the answer is &#8220;to the user,&#8221; the next question is whether anyone has designed for that user: whether they can verify the output, see its uncertainty, recover from its errors, override it, explain the final decision, and stay meaningfully in control.</p><p>Computer-aided detection answered those questions by default for more than a decade, at a cost of four hundred million dollars a year, and the answer was no. The technology was never the point. AI should not be evaluated only by what it produces. It should be evaluated by the work it leaves behind.</p><div><hr></div><h2>References</h2><p>[1] Constance D. Lehman, Robert D. Wellman, Diana S. M. Buist, Karla Kerlikowske, Anna N. A. Tosteson, and Diana L. Miglioretti, &#8220;Diagnostic Accuracy of Digital Screening Mammography With and Without Computer-Aided Detection,&#8221; JAMA Internal Medicine 175, no. 11 (2015): 1828&#8211;1837. https://doi.org/10.1001/jamainternmed.2015.5231.</p><p>[2] Robert M. Nishikawa, Robert A. Schmidt, Michael N. Linver, Alan V. Edwards, John Papaioannou, and Michael A. Stull, &#8220;Clinically Missed Cancer: How Effectively Can Radiologists Use Computer-Aided Detection?,&#8221; American Journal of Roentgenology 198, no. 3 (2012): 708&#8211;716, reporting that radiologists failed to recognize a correct CAD prompt in 71 percent of missed-cancer cases. On over-reliance when CAD is absent or incorrect, see Melina A. Kunar et al. on binary CAD and the low-prevalence effect.</p><p>[3] John Sweller, &#8220;Cognitive Load During Problem Solving: Effects on Learning,&#8221; Cognitive Science 12, no. 2 (1988): 257&#8211;285.</p><p>[4] Raja Parasuraman and Dietrich H. Manzey, &#8220;Complacency and Bias in Human Use of Automation: An Attentional Integration,&#8221; Human Factors 52, no. 3 (2010): 381&#8211;410.</p><p>[5] Mica R. Endsley and Esin O. Kiris, &#8220;The Out-of-the-Loop Performance Problem and Level of Control in Automation,&#8221; Human Factors 37, no. 2 (1995): 381&#8211;394.</p><p>[6] Saleema Amershi et al., &#8220;Guidelines for Human-AI Interaction,&#8221; in Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (CHI &#8216;19), Glasgow, Scotland UK, May 4&#8211;9, 2019 (New York: ACM, 2019), 1&#8211;13. https://doi.org/10.1145/3290605.3300233.</p><p>[7] National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (Gaithersburg, MD: U.S. Department of Commerce, January 2023).</p><p>[8] Google, People + AI Guidebook (People + AI Research, PAIR), accessed via pair.withgoogle.com.</p><div><hr></div><p>#UXDesign #HumanFactors #HSI #CognitiveUX #AIUX #HumanCenteredAI #UsabilityTesting</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/p/ai-does-not-remove-cognitive-load?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/p/ai-does-not-remove-cognitive-load?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p>]]></content:encoded></item><item><title><![CDATA[From Persona to Operational Model]]></title><description><![CDATA[How user profiles become practical tools for workflow, risk, and system design]]></description><link>https://johnwbrown.substack.com/p/from-persona-to-operational-model</link><guid isPermaLink="false">https://johnwbrown.substack.com/p/from-persona-to-operational-model</guid><dc:creator><![CDATA[John W Brown]]></dc:creator><pubDate>Tue, 16 Jun 2026 15:37:41 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!BJD6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3182b94d-31ba-4082-9c5d-828af6ca70cc_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!BJD6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3182b94d-31ba-4082-9c5d-828af6ca70cc_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!BJD6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3182b94d-31ba-4082-9c5d-828af6ca70cc_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!BJD6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3182b94d-31ba-4082-9c5d-828af6ca70cc_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!BJD6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3182b94d-31ba-4082-9c5d-828af6ca70cc_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!BJD6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3182b94d-31ba-4082-9c5d-828af6ca70cc_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!BJD6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3182b94d-31ba-4082-9c5d-828af6ca70cc_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3182b94d-31ba-4082-9c5d-828af6ca70cc_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1591693,&quot;alt&quot;:&quot;Persona&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/202299677?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3182b94d-31ba-4082-9c5d-828af6ca70cc_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Persona" title="Persona" srcset="https://substackcdn.com/image/fetch/$s_!BJD6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3182b94d-31ba-4082-9c5d-828af6ca70cc_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!BJD6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3182b94d-31ba-4082-9c5d-828af6ca70cc_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!BJD6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3182b94d-31ba-4082-9c5d-828af6ca70cc_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!BJD6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3182b94d-31ba-4082-9c5d-828af6ca70cc_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Image created with generative AI</figcaption></figure></div><div><hr></div><p><em>The views expressed are my own and do not represent any organization I&#8217;m affiliated with.</em></p><div><hr></div><p>Personas are often introduced as a way to make users memorable. That is useful, but it is not enough.</p><p>A persona that only gives a name, job title, quote, photo, and list of preferences may help a team talk about users more easily. It may even create empathy. But in complex systems, empathy alone does not tell a designer where the system will fail, where training will be needed, where workload will accumulate, or where a decision aid may create new risk.</p><p>For human factors, usability, and human systems integration work, personas should do more than represent users. They should help model the operating conditions under which people interact with systems.</p><p>That means a persona should not simply answer, Who is this user? It should help answer a more practical question: What does this user need to accomplish, under what constraints, with what consequences if the system does not support the work?</p><p>When personas are treated this way, they become more than design communication tools. They become operational models.</p><h2>The limits of the lightweight persona</h2><p>The lightweight persona is familiar. It usually includes a fictional name, role, demographic sketch, goals, frustrations, and a short narrative. In product design, this can be helpful because it gives teams a shared reference point. Instead of saying the user, a team can say this is how Denise would experience the workflow.</p><p>That shift matters. Abstract users are easy to ignore. Specific users are easier to remember.</p><p>The problem is that many personas stop at memorability. They describe the user but do not explain the work. They capture preferences but not constraints. They identify frustrations but not system dependencies. They may state that a user is busy, experienced, or technology hesitant, but they often do not show what that means in a real task environment.</p><p>For low-risk consumer products, this may be sufficient. For healthcare, government, defense, enterprise software, and AI-enabled tools, it is usually not.</p><p>In these settings, users are not simply consumers making individual choices. They are operators inside larger systems. They work with policies, handoffs, legacy tools, documentation requirements, interruptions, time pressure, professional norms, and accountability structures. Their interaction with a system is shaped by more than preference. It is shaped by the conditions of work.</p><p>To be useful, then, a persona must move beyond personality and preference. It must describe the relationship between the person, the task, the environment, and the system.</p><div><hr></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oIV5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F098ba2c6-6bac-453a-b0bf-b980a8524ed2_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oIV5!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F098ba2c6-6bac-453a-b0bf-b980a8524ed2_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!oIV5!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F098ba2c6-6bac-453a-b0bf-b980a8524ed2_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!oIV5!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F098ba2c6-6bac-453a-b0bf-b980a8524ed2_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!oIV5!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F098ba2c6-6bac-453a-b0bf-b980a8524ed2_2752x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oIV5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F098ba2c6-6bac-453a-b0bf-b980a8524ed2_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/098ba2c6-6bac-453a-b0bf-b980a8524ed2_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4225416,&quot;alt&quot;:&quot;Persona&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/202299677?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F098ba2c6-6bac-453a-b0bf-b980a8524ed2_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Persona" title="Persona" srcset="https://substackcdn.com/image/fetch/$s_!oIV5!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F098ba2c6-6bac-453a-b0bf-b980a8524ed2_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!oIV5!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F098ba2c6-6bac-453a-b0bf-b980a8524ed2_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!oIV5!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F098ba2c6-6bac-453a-b0bf-b980a8524ed2_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!oIV5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F098ba2c6-6bac-453a-b0bf-b980a8524ed2_2752x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Image created with NotebookLM</figcaption></figure></div><div><hr></div><h2>Personas as models of work</h2><p>The stronger version begins with context of use. This includes the user&#8217;s tasks, environment, tools, constraints, goals, and organizational setting. In human-centered design, context is not background information. It is part of the design problem.</p><p>For example, two users may share the same job title and still require different design support. One may work in a quiet office with time to review information carefully. Another may work in a high-interruption environment where decisions are made quickly, documentation is fragmented, and attention is divided across multiple systems.</p><p>A demographic persona may treat these users as similar. An operational persona treats them as meaningfully different.</p><p>The difference is important because design failures often appear at the boundary between formal workflow and real work. A process map may show the approved sequence of steps. A policy may describe the expected behavior. A training document may explain the correct way to complete a task. But actual work often includes exceptions, local adaptations, delays, incomplete information, and competing demands.</p><p>Exposing that gap is the operational persona&#8217;s job. It should show what the user is trying to accomplish, what information they need, and where they get it. It should show what they trust, what interrupts them, what tools compete for their attention, and what happens when the system does not fit the workflow.</p><p>The point is not added length for its own sake. A persona built this way earns its detail by making the work more legible.</p><h2>What an operational persona should include</h2><p>A practical operational persona should include several elements that are often missing from traditional persona templates.</p><p>First, it should include the user&#8217;s primary tasks. Not broad goals such as provide quality care or manage workload, but concrete activities the system must support. What does the user actually do? What decisions do they make? What information do they enter, retrieve, verify, interpret, or communicate?</p><p>Second, it should include workflow position. Is this user at the beginning of a process, the middle, or the end? Do they initiate work, review work, approve work, correct work, or inherit work from others? A user who creates information has different needs from a user who must later interpret it.</p><p>Third, it should include constraints. These may include time, staffing, policy, physical environment, competing systems, training variability, documentation burden, and interruptions. Constraints are not side issues. They are often the reason a design succeeds or fails.</p><p>Fourth, it should include information dependencies. What does the user need to know before acting? Where does that information come from? Is it structured or unstructured? Is it trusted? Is it current? Is it visible at the right point in the workflow?</p><p>Fifth, it should include risk exposure. What happens if this user misunderstands the system, misses a signal, accepts a poor recommendation, enters incomplete data, or cannot recover from an error? In many systems, the consequence of poor usability is not simply dissatisfaction. It may be delay, rework, poor coordination, degraded trust, or operational risk.</p><p>Finally, it should include support needs. Does the user need guidance, confirmation, explanation, defaults, alerts, training, job aids, peer review, or better system feedback? A persona that cannot inform support design is probably not detailed enough.</p><h2>From description to prediction</h2><p>The value of an operational persona is not only that it describes users. Its value is that it helps teams make better predictions.</p><p>A team reviewing a proposed design should be able to use the persona to interrogate that design before a single test session is scheduled. The persona tells them whether the user will understand what the system is asking for, and whether the information the user needs will actually be present at the moment of the request. It tells them whether a new workflow adds documentation burden rather than removing it, and whether an alert will arrive when the user can still act on it. It tells them where an AI-generated recommendation is likely to be trusted too much, trusted too little, or trusted for the wrong reason. Above all, it tells them where the design is likely to produce a workaround, and whether a given feature helps the user recover from a real error or only prevents the ideal error imagined by the design team.</p><p>These questions matter because many usability problems are predictable before formal testing. A well-constructed persona gives the team a way to identify likely friction points earlier, especially when paired with scenarios, journey maps, task analyses, and usability findings.</p><p>The goal is not to replace testing. The goal is to improve the quality of design reasoning before testing begins.</p><h2>A case in point</h2><p>Consider a hypothetical that composites a familiar pattern in healthcare informatics work. A health system rolls out an AI feature that drafts the summary portion of a clinical note. The project persona is a lightweight one: Dr. Lang, experienced physician, values efficiency, frustrated by documentation time. On that persona, the feature looks like a clear win, because it removes typing and returns minutes to a hurried clinician.</p><p>Now rebuild the same user as an operational persona. The primary task is not write a note. It is to produce a record that is accurate enough to support the next clinician&#8217;s decision and to withstand later review. The workflow position is consequential: the physician sits in the middle, inheriting structured data from intake and handing a record forward to colleagues and to billing. The binding constraint is interruption, because the note is often finalized between patients, in fragments. The information dependency is that the draft is only as good as the data it summarizes, and the physician cannot always see what the model left out. The risk exposure is that an inaccurate summary, signed under time pressure, propagates downstream as if it were verified.</p><p>Read that way, the persona predicts the failure before testing finds it. The design did not add a writing task. It added a verification task, and it placed that task exactly where the user has the least attention to give it. The fix is not a better draft. It is a display that surfaces what the summary is based on, flags low-confidence content, and makes the act of confirmation deliberate rather than reflexive. A demographic persona praising efficiency would never have surfaced that requirement. The operational persona made it visible while the design was still cheap to change.</p><h2>Why this matters for AI-enabled systems</h2><p>Operational personas become even more important when systems include AI.</p><p>AI tools often change the user&#8217;s role. A person who once produced a document may now review a draft. A person who once searched for information may now evaluate a generated summary. A person who once made a decision from raw data may now decide whether to accept, reject, or question a recommendation.</p><p>That shift changes cognitive work.</p><p>The user may spend less time producing an output, but more time verifying accuracy, identifying missing context, judging confidence, detecting hallucinations, explaining decisions, or managing accountability. The task has not disappeared. It has moved.</p><p>Where a traditional persona may say that the user wants efficiency, an operational persona asks what kind of efficiency is safe, where verification burden lands, and what the user must understand to remain meaningfully in control.</p><p>For AI systems, the persona should clarify the user&#8217;s relationship to the tool. Is the user an operator, reviewer, approver, supervisor, beneficiary, affected party, or downstream recipient of AI-generated output? Each role carries different needs and risks.</p><p>A clinician reviewing an AI-generated note, a benefits specialist evaluating a recommendation, a supervisor monitoring automated case routing, and a patient reading a portal summary are not simply different users. They occupy different positions in a human-AI system.</p><p>Designing for them requires more than empathy. It requires operational clarity.</p><h2>Personas should connect to evidence</h2><p>Personas are most useful when they are grounded in evidence. That evidence may come from interviews, observation, usability testing, workflow analysis, support tickets, training feedback, incident reports, analytics, survey data, or subject matter expert review.</p><p>The point is not that every persona must be statistically representative in the same way as a survey sample. Personas are interpretive models. But they should still be traceable to real findings.</p><p>Teams should be able to answer a short set of questions about any persona they rely on. What evidence supports it? What user groups or roles does it represent? What assumptions in it are still unvalidated? When was it last reviewed? What systems, workflows, or contexts does it apply to, and what should not be inferred from it?</p><p>This is especially important in large organizations where personas can outlive the research that produced them. A persona created for one system, workflow, facility type, or user group may be reused in a setting where it no longer fits.</p><p>A persona without provenance becomes folklore.</p><p>To guard against that, an operational persona should carry enough metadata to support responsible use. It should have a source history, validation status, date of last review, and known limits. This keeps the persona from becoming a decorative artifact or a permanent stereotype.</p><h2>Personas should not replace direct user engagement</h2><p>A good persona helps teams remember what they have learned about users. It does not eliminate the need to keep learning.</p><p>This distinction matters. Personas can be misused when teams treat them as substitutes for current research. A persona should focus attention, organize evidence, and help frame design questions. It should not become an excuse to avoid interviews, observation, usability testing, or field validation.</p><p>The more complex the system, the more important this becomes. In high-stakes domains, user needs change as policies, technologies, staffing models, and workflows change. A persona that was accurate three years ago may still be useful, but it should not be assumed to be current.</p><p>Operational personas should be living artifacts. As the evidence changes, they should be reviewed and refined, and when they no longer fit, retired or replaced.</p><h2>A practical template shift</h2><p>The simplest way to improve persona quality is to change the template. Instead of centering the artifact on biography, center it on operational usefulness.</p><p>A traditional persona might ask: Who is this person? What are their goals? What frustrates them? What technology do they use?</p><p>An operational persona asks a different set. What work must this person accomplish? What conditions shape that work? What information do they need, and what decisions do they make? Which errors or delays are most consequential? What system behaviors help or hinder them, and what support do they need at the point of action? And finally, what assumptions about this user are evidence-based, and what still needs validation?</p><p>This shift does not strip the human out of the persona. It makes the persona more honest about how humans actually interact with systems.</p><h2>Takeaways for practice</h2><p>A persona should be judged by what it helps a team do. If it only helps the team talk about users, it is useful but limited. If it helps the team anticipate workflow friction, design better support, identify risk, frame usability testing, and evaluate design tradeoffs, it has become an operational model. For human factors and HSI work, that should be the goal.</p><p>Personas should not be treated as static portraits. They should be practical representations of people doing work inside systems, connecting user characteristics to tasks, constraints, decisions, information needs, risks, and support requirements.</p><p>Used this way, a persona becomes more than a UX artifact. It becomes a bridge between research and design, between field reality and system requirements, and between what teams imagine users do and what users actually have to accomplish.</p><p>A good persona does not merely make the user memorable.</p><p>It makes the work visible.</p><p>So before building another persona, it is worth asking whether the ones already in use can support better design decisions. Can they predict workflow friction? Can they inform usability test scenarios, identify risk, and clarify how AI changes the user&#8217;s role? Can they distinguish between the user as imagined and the user as operating inside a real system? If not, the persona may need to mature from a profile into an operational model.</p><div><hr></div><p><em>Source basis: Nielsen Norman Group frames personas as a way to make user groups tangible and memorable and treats them as living documents requiring ongoing validation (Taylor Dykes, &#8220;Personas Make Users Memorable,&#8221; NN/g, October 2025). ISO 9241-210 and NIST human-centered design guidance both treat understanding context of use as a central design activity. Alan Cooper&#8217;s goal-directed design work treats personas as tools for understanding needs and prioritizing users. Pruitt and Adlin&#8217;s persona lifecycle work emphasizes keeping users in mind throughout product development. Human systems integration guidance frames human considerations as part of the total system, not just the interface.</em></p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/p/from-persona-to-operational-model?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/p/from-persona-to-operational-model?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p>]]></content:encoded></item><item><title><![CDATA[The Threshold Nobody Set]]></title><description><![CDATA[False positives, false negatives, and the decision usability studies keep making by accident]]></description><link>https://johnwbrown.substack.com/p/the-threshold-nobody-set</link><guid isPermaLink="false">https://johnwbrown.substack.com/p/the-threshold-nobody-set</guid><dc:creator><![CDATA[John W Brown]]></dc:creator><pubDate>Tue, 09 Jun 2026 06:02:03 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!q1RJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0c62a88-3cc0-4e17-be97-0f1b95e1d348_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!q1RJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0c62a88-3cc0-4e17-be97-0f1b95e1d348_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!q1RJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0c62a88-3cc0-4e17-be97-0f1b95e1d348_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!q1RJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0c62a88-3cc0-4e17-be97-0f1b95e1d348_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!q1RJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0c62a88-3cc0-4e17-be97-0f1b95e1d348_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!q1RJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0c62a88-3cc0-4e17-be97-0f1b95e1d348_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!q1RJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0c62a88-3cc0-4e17-be97-0f1b95e1d348_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e0c62a88-3cc0-4e17-be97-0f1b95e1d348_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2073784,&quot;alt&quot;:&quot;False Positives in Usability Testing&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/187395615?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0c62a88-3cc0-4e17-be97-0f1b95e1d348_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="False Positives in Usability Testing" title="False Positives in Usability Testing" srcset="https://substackcdn.com/image/fetch/$s_!q1RJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0c62a88-3cc0-4e17-be97-0f1b95e1d348_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!q1RJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0c62a88-3cc0-4e17-be97-0f1b95e1d348_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!q1RJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0c62a88-3cc0-4e17-be97-0f1b95e1d348_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!q1RJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0c62a88-3cc0-4e17-be97-0f1b95e1d348_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Image created with generative AI</figcaption></figure></div><div><hr></div><p><em>The views expressed are my own and do not represent any organization I am affiliated with.</em></p><div><hr></div><p>Consider the last hour of a usability study, the part that rarely makes it into the report: the finding-review meeting.</p><p>Two evaluators have watched the same recorded session. A participant using a clinical scheduling tool reached the final confirmation screen, paused for roughly forty seconds, backed out to the previous step, returned, and completed the booking. She did not comment on the pause. She did not repeat it across the next three tasks. Afterward she rated her confidence as high.</p><p>The first evaluator logs this as a high-severity finding. The confirmation flow, in this reading, is confusing enough to make a user retreat at the moment of commitment. The second evaluator logs almost nothing: a moment of orientation on first contact, gone by the second task, not worth a redesign. Same clip. Same forty seconds. Opposite severity.</p><p>The meeting resolves the disagreement the way most meetings do, by seniority, by who speaks with more conviction, or by splitting the difference and calling it medium. What it does not do is resolve the disagreement with evidence. And the reason is worth sitting with: the two evaluators are not actually disagreeing about what happened. They are disagreeing about where to set the line between a problem and a non-problem, and neither of them has said so out loud.</p><p>That unstated line is the subject of this piece. Usability testing is usually described as the work of finding problems. It is at least as much the work of deciding which observed behaviors count as problems at all, and that second task is harder, less examined, and more consequential than the first.</p><h2>The finding that looked worse than it was</h2><p>A false positive, borrowing the term from diagnostic testing, is a finding that looks serious under study conditions but does not produce real harm, inefficiency, or risk once the system is in use. The participant stumbles, asks a clarifying question, frowns at a label, then recovers and never encounters the issue again. In the room, it is vivid. The behavior is visible, the reaction is real, the note gets written. At scale, the predicted problem never arrives.</p><p>These behaviors are worth noticing. The cost comes from what happens after they are escalated. Engineering time goes to a confirmation screen that was never going to slow anyone down. The genuine issue two rows down the priority list waits another quarter. And there is a slower, more corrosive cost. When stakeholders watch three findings marked critical get fixed with no measurable change in any metric they care about, they begin to discount the fourth report before they have read it. Credibility, once spent this way, is expensive to rebuild.</p><p>So far this is the familiar argument, and it is correct as far as it goes. But a study that worries only about false positives has solved half a problem and opened another. Every decision to dismiss a finding is also a decision that can be wrong. The behavior you wave off as orientation may be the first visible edge of a defect that will cost thousands of real users real time once the system ships. A practice that takes pride in escalating less is not automatically more disciplined. It may be failing in the opposite direction, and failing quietly, because no one writes a report about the problem they decided not to report.</p><h2>Four outcomes, one threshold</h2><p>It helps to borrow a frame from a field that has spent decades on this exact decision. Signal detection theory, formalized for perception research and radar operators in the middle of the last century, describes what happens whenever someone tries to detect a faint signal against background noise.[^1] There are four possible outcomes. You can catch a real signal, which is a hit. You can fail to catch one that was present, which is a miss. You can raise an alarm at noise that meant nothing, which is a false alarm. Or you can stay quiet when there was in fact nothing to report, which is a correct rejection.</p><p>Map this onto a usability study and the picture sharpens. A real problem you identify is a hit. A false positive is a false alarm. A real problem you dismiss, or never surface at all, is a miss. And the orientation behavior you correctly decline to escalate is a correct rejection, the quiet, unglamorous outcome that good judgment produces all day and no one ever praises.</p><p>The reason this frame earns its place is the part practitioners tend to skip. You cannot drive down false alarms and misses at the same time by simply trying harder or caring more. They trade against each other across a threshold. Lower the threshold, escalate at the faintest signal, and you catch more real problems while also flagging more noise. Raise it, escalate only the unmistakable, and you cut the noise while letting more real problems slip past unreported. The threshold is not a flaw in the method. It is the central decision of the method, and most studies make it by accident.</p><h2>Why two careful evaluators see different things</h2><p>The disagreement in that review meeting reflects the normal condition of the work rather than carelessness on anyone&#8217;s part, and it has been measured repeatedly.</p><p>When Jacobsen, Hertzum, and John had four trained evaluators independently analyze the same four recorded usability sessions, the four together identified ninety-three problems. Only about a fifth of those were caught by all four evaluators, and nearly half were caught by a single evaluator working alone.[^2] When each evaluator then selected the ten problems they considered most severe, the top-ten lists did not converge on a shared core. A later review of eleven studies of this kind found agreement between evaluators ranging from five percent to sixty-five percent, and the effect held for severity judgments specifically, not only for whether a problem got noticed in the first place.[^3]</p><p>This has a direct consequence for how the word &#8220;severe&#8221; should be read in any report. A severity rating is partly a statement about the person who assigned it, not a measurement taken off the system. Which means the inflation of findings does not usually come from bad faith or sloppiness. It comes from a low-reliability instrument being treated as if it were a precise one, and from a threshold that nobody set deliberately doing its work in the background.</p><h2>The half-true story about the lab</h2><p>There is a comfortable explanation that makes the whole problem disappear. The lab is artificial, the story goes, so it manufactures problems the field would never produce, and the remedy is to trust real-world use over the test. The trouble is that this story is half right, which is the most dangerous kind of right.</p><p>Testing environment does shape what a study surfaces. That much is established.[^4] But the further assumption, that more realism reliably means fewer problems, does not survive contact with the evidence. When researchers evaluated a clinical device across lower and higher fidelity conditions, both conditions surfaced the same types of use error, and increasing ecological validity did not dependably shrink the count or reveal that the lab results had been phantoms.[^5] The lab does not conjure problems from nothing; it changes which problems are salient and how often they appear. That is a reason to interpret findings with care, not a license to assume the field will quietly absolve whatever the lab turned up.</p><p>So the false positive is real, and it is worth managing. It is just not the simple artifact of an unrealistic room that the convenient story makes it out to be.</p><h2>A way to triage without pretending</h2><p>If the threshold is the real decision, the practical question is how to set it on purpose. Three diagnostic questions do most of the work, and a fourth step ties them to the stakes.</p><p>First, recurrence. Did the behavior repeat across participants and across sessions, or did it appear once and vanish? A pattern that shows up in four of eight participants is a different object than a single startled pause. Recurrence is the closest thing a study has to a signal-strength reading, and it gets more trustworthy the more participants you ran.</p><p>Second, consequence. Trace what the behavior actually led to. Did it produce a wrong decision, unnecessary rework, an abandoned task, or an unsafe action? Or did the user absorb a few seconds of friction and arrive in the right place with no downstream effect? A finding with no consequence attached to it is a finding waiting for a justification.</p><p>Third, trajectory. Distinguish learning from breakdown. New users orient themselves. They pause, test a control, confirm an assumption, and then proceed, and on the next encounter the pause is gone. Breakdown is the opposite shape. It does not resolve with exposure, it recurs or worsens, and it tends to leave a residue of workarounds behind it. The question is whether the struggle is on its way out or on its way in, not whether the user stumbled once.</p><p>Those three questions estimate how confident you should be that a finding represents a real-use signal. The fourth step decides what to do with that confidence, and it is the one most studies omit. Calibrate the threshold to the cost structure of the system in front of you. In a safety-critical tool, a missed problem can injure someone, so a miss costs far more than a false alarm, and the rational move is to escalate on thinner evidence and accept that you will chase some noise. In a low-stakes, high-volume consumer flow, the arithmetic inverts, and a threshold that escalates everything will bury the team in redesigns that no user needed. Same method, deliberately different line, set before the findings arrive rather than argued about after.</p><h2>Two findings from the same study</h2><p>Picture two findings from a single hypothetical study of that clinical scheduling tool, and run each through the triage.</p><p>The first is the forty-second pause from the opening. One participant, one occurrence, resolved by the second task, no downstream error, confidence high afterward. Recurrence is absent, consequence is none, trajectory points toward learning. The triage answer is to document it honestly and decline to escalate it, with the reasoning written down so that the next evaluator who sees the clip does not relitigate it from scratch. This is a correct rejection, and getting it right is a skill, not a failure of diligence.</p><p>The second finding is quieter, and worse. Three of eight participants accepted an incorrectly pre-populated field, a default value carried over from a prior screen, without hesitation. No frown, no pause, no comment. By the task-completion metric, all three succeeded. In the room, nothing appeared to go wrong. Recurrence is present, consequence is high, and trajectory is irrelevant because the users never registered a problem to learn their way out of. The triage flags this as severe precisely because visible reaction is not one of its criteria.</p><p>That contrast is the whole argument in miniature. The most dangerous finding in a usability study is often the one where nothing appears to go wrong. An evaluation practice tuned to suppress false positives by looking for drama will reliably catch the harmless pause and reliably miss the silent acceptance of a wrong value, which is exactly backward. Friction is easy to see and frequently cheap. The calm error is hard to see and sometimes the one that ships.</p><h2>Putting it to work</h2><p>A few moves turn this from a frame into a practice.</p><p>Decide the severity rubric and the expected error types before the sessions, not during the debrief. Pre-commitment will not eliminate the evaluator effect, but it shrinks the space in which conviction and seniority substitute for evidence after the fact.</p><p>Record two things separately for every finding rather than collapsing them into one number. One is how confident you are that the behavior reflects real use. The other is how often it recurred. A single severity score hides both, and the hidden parts are where the disagreements live.</p><p>For any finding you mean to call severe, get a second evaluator on the tape independently before the recommendation leaves the room. The research is unambiguous that agreement is lowest exactly where it matters most, on the severe calls, so that is where a second pass buys the most.</p><p>Keep the honest annotations the discipline already knows how to write. &#8220;Observed once.&#8221; &#8220;May be orientation.&#8221; &#8220;Did not recur.&#8221; Far from hedging, these annotations are the confidence half of the record, and a report that omits them overstates what it knows.</p><p>Finally, separate what you observed from what you recommend. The observation can be reported with certainty while the recommendation stays provisional. Conflating the two is how a forty-second pause becomes a redesign mandate.</p><h2>Where this can go wrong</h2><p>A triage framework is abusable, and pretending otherwise would undercut the point. &#8220;It is probably just learning&#8221; sits one motivated step away from &#8220;we would rather not fix this,&#8221; and a team under deadline pressure will find the criteria remarkably accommodating. The protection is sequence. The threshold has to be set before the findings are known, not reverse-engineered afterward to fit the roadmap that already existed.</p><p>Recurrence also depends on having enough participants to mean anything. In a five-person study, one occurrence cannot be reliably told apart from a problem that would surface in six of thirty. Small studies should lean harder on consequence and trajectory and treat a single observation as weak evidence rather than a settled non-problem.</p><p>The lab-and-field relationship is genuinely conditional, not directional. Studies have found the two equivalent under favorable conditions and divergent under harder ones, which means you cannot promise a stakeholder that the field will confirm or overturn a given finding. You do not know in advance which way any particular finding will move, and claiming otherwise trades one false certainty for its mirror image.</p><p>And signal detection is a lens here, not a measurement. I am not proposing that anyone compute a sensitivity index from a think-aloud session. The value is in the shape of the tradeoff and the discipline of naming the threshold, not in a number that would imply more precision than the data can carry.</p><h2>Back to the meeting</h2><p>Return to the two evaluators and the forty seconds. They were never really disagreeing about the pause. Each was applying a different threshold and reporting the result as severity, and because neither threshold was visible, the conversation had nowhere to go but conviction. Make the threshold explicit, tie it to what a miss costs against what a false alarm costs in this particular system, and the meeting turns from a negotiation into a decision that can be defended later.</p><p>Strong usability practice depends as much on deciding what not to escalate as on identifying genuine problems. But the decision only holds up when both errors are kept in view at once. The loud finding that wastes a sprint and the silent one that ships are failures of the same instrument, set to the wrong threshold in opposite directions. The findings most worth acting on are usually the quiet ones, recurring and consequential while making no noise in the room.</p><div><hr></div><h2>Quick Reference: Triaging a Usability Finding</h2><p>Before the sessions: set the severity rubric, name the error types you expect, and decide the threshold based on the cost of a miss versus a false alarm in this system.</p><p>For each finding, ask:</p><ol><li><p><strong>Recurrence.</strong> Did it repeat across participants and sessions, or appear once and vanish?</p></li><li><p><strong>Consequence.</strong> Did it cause a wrong decision, rework, abandonment, or an unsafe action, or did it resolve with no downstream effect?</p></li><li><p><strong>Trajectory.</strong> Is the difficulty resolving with exposure (learning) or recurring and breeding workarounds (breakdown)?</p></li></ol><p>For each finding, log two values separately: confidence that it reflects real use, and frequency of recurrence. Do not collapse them into one severity score.</p><p>For any severe call, put a second evaluator on the recording independently before recommending.</p><p>Watch the asymmetry. A practice that never produces a false positive is escalating too little, not succeeding. The quiet, consequential, recurring finding is the one most worth protecting from dismissal.</p><div><hr></div><h2>Notes</h2><p>[^1]: David M. Green and John A. Swets, <em>Signal Detection Theory and Psychophysics</em> (New York: Wiley, 1966).</p><p>[^2]: Niels Ebbe Jacobsen, Morten Hertzum, and Bonnie E. John, &#8220;The Evaluator Effect in Usability Studies: Problem Detection and Severity Judgments,&#8221; in <em>Proceedings of the Human Factors and Ergonomics Society 42nd Annual Meeting</em> (Santa Monica, CA: HFES, 1998), 1336&#8211;40.</p><p>[^3]: Morten Hertzum and Niels Ebbe Jacobsen, &#8220;The Evaluator Effect: A Chilling Fact About Usability Evaluation Methods,&#8221; <em>International Journal of Human-Computer Interaction</em> 15, no. 1 (2003): 183&#8211;204.</p><p>[^4]: Juergen Sauer and Andreas Sonderegger, &#8220;Methodological Issues in Product Evaluation: The Influence of Testing Environment and Task Scenario,&#8221; <em>Applied Ergonomics</em> 42, no. 3 (2011): 487&#8211;94.</p><p>[^5]: Romaric Marcilly, Helen Monkman, Sylvia Pelayo, and Blake J. Lesselroth, &#8220;Usability Evaluation Ecological Validity: Is More Always Better?&#8221; <em>Healthcare</em> 12, no. 14 (2024): 1417.</p><div><hr></div><h2>Bibliography</h2><p>Green, David M., and John A. Swets. <em>Signal Detection Theory and Psychophysics</em>. New York: Wiley, 1966.</p><p>Hertzum, Morten, and Niels Ebbe Jacobsen. &#8220;The Evaluator Effect: A Chilling Fact About Usability Evaluation Methods.&#8221; <em>International Journal of Human-Computer Interaction</em> 15, no. 1 (2003): 183&#8211;204.</p><p>Jacobsen, Niels Ebbe, Morten Hertzum, and Bonnie E. John. &#8220;The Evaluator Effect in Usability Studies: Problem Detection and Severity Judgments.&#8221; In <em>Proceedings of the Human Factors and Ergonomics Society 42nd Annual Meeting</em>, 1336&#8211;40. Santa Monica, CA: HFES, 1998.</p><p>Marcilly, Romaric, Helen Monkman, Sylvia Pelayo, and Blake J. Lesselroth. &#8220;Usability Evaluation Ecological Validity: Is More Always Better?&#8221; <em>Healthcare</em> 12, no. 14 (2024): 1417.</p><p>Sauer, Juergen, and Andreas Sonderegger. &#8220;Methodological Issues in Product Evaluation: The Influence of Testing Environment and Task Scenario.&#8221; <em>Applied Ergonomics</em> 42, no. 3 (2011): 487&#8211;94.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/p/the-threshold-nobody-set?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/p/the-threshold-nobody-set?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p>]]></content:encoded></item><item><title><![CDATA[What Heuristic Violations Actually Predict Real-World Failure?]]></title><description><![CDATA[Severity ratings detect deviations. They do not predict harm. A three-dimension framework for the gap.]]></description><link>https://johnwbrown.substack.com/p/what-heuristic-violations-actually</link><guid isPermaLink="false">https://johnwbrown.substack.com/p/what-heuristic-violations-actually</guid><dc:creator><![CDATA[John W Brown]]></dc:creator><pubDate>Tue, 02 Jun 2026 06:01:04 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!MDMf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ab1a016-029c-455c-9c92-1b6a1aecb06d_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!MDMf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ab1a016-029c-455c-9c92-1b6a1aecb06d_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!MDMf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ab1a016-029c-455c-9c92-1b6a1aecb06d_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!MDMf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ab1a016-029c-455c-9c92-1b6a1aecb06d_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!MDMf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ab1a016-029c-455c-9c92-1b6a1aecb06d_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!MDMf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ab1a016-029c-455c-9c92-1b6a1aecb06d_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!MDMf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ab1a016-029c-455c-9c92-1b6a1aecb06d_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1ab1a016-029c-455c-9c92-1b6a1aecb06d_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2573467,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/187396226?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ab1a016-029c-455c-9c92-1b6a1aecb06d_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!MDMf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ab1a016-029c-455c-9c92-1b6a1aecb06d_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!MDMf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ab1a016-029c-455c-9c92-1b6a1aecb06d_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!MDMf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ab1a016-029c-455c-9c92-1b6a1aecb06d_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!MDMf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ab1a016-029c-455c-9c92-1b6a1aecb06d_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Image created with generative AI</figcaption></figure></div><div><hr></div><p><em>The views expressed are my own and do not represent any organization I am affiliated with.</em></p><div><hr></div><h2>Opening</h2><p>Consider two findings from a hypothetical usability evaluation of an enterprise scheduling platform.</p><p>The first: terminology inconsistency. The same field is labeled &#8220;Start Time&#8221; on the booking screen, &#8220;Begin At&#8221; on the confirmation page, and &#8220;Scheduled For&#8221; in the calendar view. The evaluator flags this as a violation of Consistency and Standards. Severity rating: 3 of 4. Significant problem, fix recommended before release.</p><p>The second: status ambiguity. When a user submits a booking request, the screen returns to the dashboard without explicit confirmation. The booking appears in the user&#8217;s pending list, but only on the next refresh. The evaluator flags this as Visibility of System Status. Severity rating: 2 of 4. Minor problem, fix if time permits.</p><p>Six months after launch, the inconsistency in terminology has resulted in zero help desk tickets. Users adapted within their first two sessions and no longer notice the variation. The status ambiguity has produced a steady stream of duplicate bookings, support escalations, and one near-miss in a clinical environment where a procedure was double-scheduled because the requesting clinician could not confirm her submission had been received.</p><p>The higher-severity finding produced no real friction. The lower-severity finding produced systemic risk.</p><p>Anyone who has tracked an evaluation report into deployment recognizes this pattern. Heuristic severity ratings, taken at face value, are unreliable predictors of which issues will actually hurt users in production. The question is why, and what to use instead.</p><h2>The Gap Between Rating and Impact</h2><p>Heuristic evaluation was developed in an era when interfaces were largely deterministic. A button either responded or it did not. A menu either matched the user&#8217;s mental model or it did not. The classical heuristics, articulated by Nielsen in the early 1990s and reinforced by the engineering psychology tradition associated with Wickens, were calibrated against that world.</p><p>In that world, severity ratings served as reasonable proxies for real-world impact. The gap between &#8220;this violates the guideline&#8221; and &#8220;this will hurt the user&#8221; was narrow. A consistency issue produced predictable confusion. A control failure produced predictable frustration. The mapping was tight enough that rating and prediction looked like the same activity.</p><p>Modern systems have widened the gap.</p><p>Today&#8217;s systems are adaptive. They respond differently to different users and to the same user over time. They infer intent. They generate content. They surface information selectively based on relevance models that the user cannot inspect. A violation in this environment is not a static deviation from a guideline. It is a behavior that interacts with user adaptation and workflow context under conditions that the evaluator cannot fully model from the screen alone.</p><p>The severity rating remains a useful signal. It is no longer sufficient.</p><p>What experienced practitioners do, often without articulating it, is supply a second layer of judgment on top of the rating. They ask whether the friction will persist. They ask whether it will appear often enough to matter. They ask whether the user is being set up to act on something they cannot verify. The framework below names what the second layer is doing.</p><h2>What Changed: The New Failure Modes</h2><p>Three shifts in how systems behave have reduced the predictive power of severity alone.</p><p>The first is adaptation speed. In practitioner experience, users adapt to interfaces faster than the heuristic tradition implicitly assumes, particularly in tools they use daily. Minor inconsistencies and aesthetic imperfections that look serious to a fresh evaluator often vanish into procedural memory within a week. The severity rating reflects first-encounter impact. The deployed system is judged by long-run impact.</p><p>The second is opacity. In systems that incorporate machine learning, recommendation, or generative behavior, the user is asked to act on outputs whose provenance is not visible. Whether a suggested diagnosis came from a rule-based decision tree or a statistical model trained on data from another institution changes how the user should weight it. Heuristics designed for the visible state do not cleanly extend to the inferred state.</p><p>The third is asymmetry of consequence. A trivial-looking issue in a high-stakes workflow can cause harm disproportionate to its apparent severity, while a glaring problem in a low-stakes workflow can be absorbed without damage. The classical severity scale, applied uniformly across screens, treats these as equivalent. They are not.</p><p>Each of these shifts pushes the practitioner toward judgment that the heuristic itself cannot supply. The framework below organizes that judgment into something teachable.</p><h2>A Predictive Framework: Three Dimensions</h2><p>A heuristic violation predicts real-world failure to the extent that it scores high on three dimensions.</p><h3>Dimension 1: Persistence</h3><p>Will the issue survive user adaptation?</p><p>Some findings decay. Users encounter them, develop a workaround, and proceed without further friction. Inconsistent terminology, minor layout irregularities, and most aesthetic issues fall into this category. Their first encounter cost is real. Their long-run cost approaches zero.</p><p>Other findings do not decay. The user encounters them on every cycle, cannot route around them, and experiences the same friction at session 500 that they did at session 5. Status ambiguity, recovery dead-ends, and any condition that requires the user to verify what the system should have communicated belong here. The friction does not fade because the underlying ambiguity does not fade.</p><p>The practical question for the evaluator: Can a competent user, given a week of regular use, develop a stable workaround? If yes, the issue likely overpredicts long-run impact. If no, it predicts persistent failure.</p><h3>Dimension 2: Stakes-Weighted Frequency</h3><p>How often does the violation occur, and what is at stake when it does?</p><p>Frequency without stakes is annoyance. Stakes without frequency is risk. The product of the two is what predicts impact in deployment.</p><p>A finding that surfaces once per month in a low-consequence workflow is tolerable, even at high severity. A finding that surfaces dozens of times per day in a low-consequence workflow becomes systemic friction through sheer repetition. A finding that surfaces rarely but appears at a decision point where the consequence is irreversible carries weight regardless of how clean the rest of the interface is.</p><p>The test in practice: where in the user&#8217;s daily workflow does this appear, and what is the cost of the worst plausible misstep when it does? The answer reweights severity in ways the original rating cannot.</p><h3>Dimension 3: Inference Load</h3><p>Does the violation force the user to infer system state, intent, or output provenance that should be directly available?</p><p>This dimension captures the new failure modes most cleanly. When a system produces a suggestion, a ranking, a flagged record, or a generated output, the user must decide how to weight it. That decision depends on what the user can know about how the output was produced and how confident the system is in it.</p><p>When that information is absent, ambiguous, or buried, the user is forced into inference. Inference under uncertainty is where serious mistakes accumulate. The user assumes the system is more certain than it is, or less certain than it is, and acts accordingly. Neither error is visible in the interface itself.</p><p>Inference load violations often score low on classical heuristics because the interface is consistent, responsive, and well-labeled. They predict real-world failure anyway, because the failure happens in the user&#8217;s reasoning rather than in the user&#8217;s interaction. This is the most common blind spot in heuristic reports on AI-assisted systems today.</p><p>The diagnostic question: at any decision point, is the user being asked to act on something the system has not made fully legible? Where the answer is yes, the issue predicts failure even when the surface heuristics are clean.</p><h2>Worked Cases</h2><h3>Case 1: The Cosmetic Inconsistency</h3><p>A logistics platform displays delivery windows in twelve-hour format on the dispatch screen and twenty-four-hour format on the route detail screen. The evaluator flags Consistency and Standards. Severity 3 of 4.</p><ul><li><p>Persistence: Low. Drivers and dispatchers adapt within the first shift. The translation becomes automatic.</p></li><li><p>Stakes-Weighted Frequency: Low. Time confusion produces a clarifying call, not a missed delivery. Frequent but recoverable.</p></li><li><p>Inference Load: Low. The information is fully present. No system reasoning is hidden.</p></li></ul><p>Predicted real-world impact: minor. The classical severity overstates the issue. Fix on the next release cycle, not before launch.</p><h3>Case 2: The Quiet Status Failure</h3><p>An electronic health records system saves a draft note when the user navigates away, but does not display a save indicator. The user assumes the note is preserved because there is no error message. The evaluator flags Visibility of System Status. Severity 2 of 4.</p><ul><li><p>Persistence: High. The ambiguity cannot be adapted around because the user has no reliable way to verify save state without leaving and returning to the note.</p></li><li><p>Stakes-Weighted Frequency: High. Clinical documentation is frequent and consequential. A lost note creates a documentation gap with downstream effects on billing, continuity of care, and medicolegal exposure.</p></li><li><p>Inference Load: Moderate to high. The user must infer save status from the absence of an error message, which is a logically unreliable signal.</p></li></ul><p>Predicted real-world impact: severe. The classical severity understates the issue by roughly a full grade. This is the finding that should block release, not the one rated higher.</p><h3>Case 3: The Modern Failure Mode</h3><p>A clinical decision support tool surfaces a ranked list of differential diagnoses. The interface is clean and responsive. Labels are consistent. No classical heuristic violations are detected. The evaluator returns a clean report on the relevant screen.</p><ul><li><p>Persistence: Not applicable in the classical sense. There is no surface friction to adapt around.</p></li><li><p>Stakes-Weighted Frequency: Maximum. Diagnostic reasoning is high-frequency in the workflow and high-consequence in outcome.</p></li><li><p>Inference Load: Maximum. The clinician is being asked to act on a ranking whose underlying confidence, training distribution, and known failure modes are not visible at the decision point.</p></li></ul><p>Predicted real-world impact: severe, despite zero classical findings. This is the case heuristic evaluation alone will miss, and it is increasingly the case that matters most.</p><h2>Implementation: Using This in Real Evaluation Work</h2><p>The three-dimension framework is not a replacement for heuristic evaluation. It is a layer applied after the heuristic pass.</p><p>In practice this means three changes to how findings are documented and communicated.</p><p>First, in the evaluation report itself, each finding above a chosen severity threshold should be annotated with a predicted-impact rating derived from the three dimensions. This can be a simple low-medium-high tag on each axis, or a weighted composite depending on the formality of the engagement. The annotation makes visible what experienced evaluators are already doing privately.</p><p>Second, in stakeholder meetings, the predicted-impact rating should lead the conversation, not the classical severity. Engineering leaders and product owners need to know what will hurt users in production. Telling them that a rating-3 finding will produce no real friction while a rating-2 finding will produce systemic risk is the actual practitioner contribution. It is what distinguishes evaluation from inspection.</p><p>Third, in the test plan that follows the evaluation, the issues with the highest predicted impact should be the ones targeted in user testing, not the ones with the highest classical severity. The point of the framework is to direct scarce research budget at the conditions most likely to fail.</p><p>A short version of this annotation, suitable for embedding in any evaluation report, appears at the close of this piece.</p><h2>Where the Framework Breaks</h2><p>The three-dimension framework has known limits and should not be presented as universal.</p><p>It assumes the evaluator has enough domain knowledge to estimate stakes and frequency in the user&#8217;s actual workflow. In greenfield environments or novel domains, this estimate may be unreliable, and classical severity may be the more honest starting point until field data accumulates.</p><p>It assumes that adaptation is observable and that workarounds are stable. In systems undergoing rapid iteration, where the interface shifts under the user faster than adaptation can occur, persistence estimates lose meaning. Every finding begins to behave as if it is high-persistence because the user never reaches steady state.</p><p>It assumes that inference load can be assessed by inspection. In systems where the user&#8217;s reasoning is partially externalized through configuration, intelligent agents, or multi-step workflows, inference load may be distributed across the system in ways a single-screen evaluation cannot capture.</p><p>Each of these conditions calls for paired user research rather than expanded heuristic work. The framework is a sharper inspection tool. It is not a substitute for talking to users in their actual context.</p><h2>Closing</h2><p>Heuristic evaluation remains a strong inspection method. Its weakness is not in detection but in prediction. A rating that captures deviation from a guideline does not, by itself, capture which deviations will produce real harm in deployment.</p><p>The three dimensions named here, persistence, stakes-weighted frequency, and inference load, do not replace classical severity. They translate it into something closer to a forecast. The forecast is what stakeholders need, what testing budgets should be aligned against, and what experienced evaluators have been quietly supplying for years without naming.</p><p>The contribution of this framework is to name it, so that it can be taught rather than only practiced.</p><div><hr></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kv6_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e896a0e-ff58-400c-a859-d2b36688a4db_711x425.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kv6_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e896a0e-ff58-400c-a859-d2b36688a4db_711x425.png 424w, https://substackcdn.com/image/fetch/$s_!kv6_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e896a0e-ff58-400c-a859-d2b36688a4db_711x425.png 848w, https://substackcdn.com/image/fetch/$s_!kv6_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e896a0e-ff58-400c-a859-d2b36688a4db_711x425.png 1272w, https://substackcdn.com/image/fetch/$s_!kv6_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e896a0e-ff58-400c-a859-d2b36688a4db_711x425.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kv6_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e896a0e-ff58-400c-a859-d2b36688a4db_711x425.png" width="711" height="425" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6e896a0e-ff58-400c-a859-d2b36688a4db_711x425.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:425,&quot;width&quot;:711,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:72349,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/187396226?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e896a0e-ff58-400c-a859-d2b36688a4db_711x425.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!kv6_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e896a0e-ff58-400c-a859-d2b36688a4db_711x425.png 424w, https://substackcdn.com/image/fetch/$s_!kv6_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e896a0e-ff58-400c-a859-d2b36688a4db_711x425.png 848w, https://substackcdn.com/image/fetch/$s_!kv6_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e896a0e-ff58-400c-a859-d2b36688a4db_711x425.png 1272w, https://substackcdn.com/image/fetch/$s_!kv6_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e896a0e-ff58-400c-a859-d2b36688a4db_711x425.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>References</h2><p>Nielsen, Jakob. &#8220;Enhancing the Explanatory Power of Usability Heuristics.&#8221; In <em>Proceedings of the SIGCHI Conference on Human Factors in Computing Systems</em>, 152&#8211;158. New York: ACM, 1994.</p><p>Wickens, Christopher D., Justin G. Hollands, Simon Banbury, and Raja Parasuraman. <em>Engineering Psychology and Human Performance</em>. New York: Routledge.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/p/what-heuristic-violations-actually?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/p/what-heuristic-violations-actually?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div>]]></content:encoded></item><item><title><![CDATA[Why Users Don’t Use Help, Even When They Need It]]></title><description><![CDATA[The Help Credibility Loop and Why Documentation Quietly Loses Its Audience]]></description><link>https://johnwbrown.substack.com/p/why-users-dont-use-help-even-when</link><guid isPermaLink="false">https://johnwbrown.substack.com/p/why-users-dont-use-help-even-when</guid><dc:creator><![CDATA[John W Brown]]></dc:creator><pubDate>Tue, 26 May 2026 06:01:43 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!CVKj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206af2ec-cad1-4502-a594-bf668d753e5d_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CVKj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206af2ec-cad1-4502-a594-bf668d753e5d_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CVKj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206af2ec-cad1-4502-a594-bf668d753e5d_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!CVKj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206af2ec-cad1-4502-a594-bf668d753e5d_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!CVKj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206af2ec-cad1-4502-a594-bf668d753e5d_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!CVKj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206af2ec-cad1-4502-a594-bf668d753e5d_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CVKj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206af2ec-cad1-4502-a594-bf668d753e5d_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/206af2ec-cad1-4502-a594-bf668d753e5d_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1973300,&quot;alt&quot;:&quot;Why Users Don&#8217;t Use Help, Even When They Need It&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/187395810?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206af2ec-cad1-4502-a594-bf668d753e5d_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Why Users Don&#8217;t Use Help, Even When They Need It" title="Why Users Don&#8217;t Use Help, Even When They Need It" srcset="https://substackcdn.com/image/fetch/$s_!CVKj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206af2ec-cad1-4502-a594-bf668d753e5d_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!CVKj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206af2ec-cad1-4502-a594-bf668d753e5d_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!CVKj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206af2ec-cad1-4502-a594-bf668d753e5d_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!CVKj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206af2ec-cad1-4502-a594-bf668d753e5d_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Image created with generative AI</figcaption></figure></div><div><hr></div><p><em>The views expressed are my own and do not represent any organization I am affiliated with.</em></p><div><hr></div><p>Consider a clinician working in an electronic health record system at 7:42 in the morning. The patient list is loading slowly. A new banner has appeared at the top of the chart for a recently transferred patient, with an unfamiliar status code attached. The clinician needs to know whether the code requires a chart action before rounds begin in eighteen minutes.</p><p>A &#8220;Help&#8221; icon sits in the upper-right corner of the screen. The clinician does not click it. Instead, she opens a Teams chat to a colleague she knows is already on shift. The colleague does not know either, but checks a third source. Two minutes pass. The status code is decoded by triangulation.</p><p>The help button was right there. It was never considered.</p><p>This is not laziness. It is the rational behavior of a user who has learned, across thousands of prior interactions, that consulting in-product help is a poor return on attention.</p><p>Help systems have a credibility problem. And credibility, once spent, is hard to rebuild.</p><h2>What the Field Already Knew</h2><p>The argument that users avoid documentation is not new. Marc Rettig made it directly in 1991, with an article whose title remains its thesis: &#8220;Nobody Reads Documentation.&#8221;&#185; John Carroll&#8217;s <em>The Nurnberg Funnel</em>, published a year earlier, was a sustained study of why users abandon training material and what minimalist instruction would have to do to survive their actual habits.&#178; Jakob Nielsen included help and documentation as the tenth and last of his usability heuristics, with a measured note that the best systems should not need them, and the next-best should make them easy to search, task-focused, and concrete.&#179;</p><p>The conventional reading of this literature is that help is a fallback. Build the interface well enough that help is rarely needed. When it is needed, make it good. This is reasonable advice. It is also incomplete.</p><p>What the literature underdiagnoses is the running ledger users keep. Every help interaction either earns credit or spends it. After enough withdrawals, the user stops opening the account.</p><p>Call this the Help Credibility Loop.</p><h2>The Help Credibility Loop</h2><p>The Help Credibility Loop describes the dynamic by which a help system&#8217;s perceived value rises or falls in the user&#8217;s mind based on the running history of their interactions with it. The loop has five steps.</p><ol><li><p><strong>Encounter.</strong> The user hits friction. Something is unclear, unexpected, or unfamiliar. A status code, an error message, a new feature.</p></li><li><p><strong>Threshold.</strong> The user assesses, usually below conscious deliberation, whether consulting help is worth the interruption. This assessment is shaped almost entirely by prior experience with the help system, the product, and the broader category of help systems.</p></li><li><p><strong>Action.</strong> The user either consults help or routes around it. The route-around may be a colleague, a search engine, trial and error, or feature avoidance.</p></li><li><p><strong>Resolution.</strong> The friction is resolved, partially resolved, or abandoned. If help was consulted, it either delivered or did not.</p></li><li><p><strong>Update.</strong> The user&#8217;s running credibility balance for that help channel adjusts. Resolution earns a small credit. Failure spends a larger debit. Abandonment without consulting help, when help would have worked, does nothing to the balance, because the user never finds out.</p></li></ol><p>The loop is asymmetric. Bad help experiences leave a sharper trace than good ones. This is consistent with Lee and See&#8217;s broader account of trust in automation, where calibration errors and unmet expectations degrade reliance faster than competent performance restores it.&#8308; It is also consistent with Parasuraman and Riley&#8217;s framework on the misuse, disuse, and abuse of automation: disuse, in their terms, is the disengagement of a capable system because the operator has stopped trusting it.&#8309; Help disuse is automation disuse, applied to the documentation layer.</p><p>The implication is structural. A help system that has lost credibility cannot be repaired by adding more help. It can only be repaired by interactions that begin returning credit to the balance, one consultation at a time.</p><h2>Why Help Spends Credit Faster Than It Earns It</h2><p>Four common patterns drain the credibility balance. Practitioners who run usability evaluations will recognize all of them.</p><p><strong>Generic content.</strong> Help is written at a level of abstraction one or two steps removed from the user&#8217;s actual situation. The user encounters a status code; help explains that status codes communicate patient transfer state. The user already knew that. What they wanted was to know what this code, in this chart, means for what to do next. The translation work required to bridge the gap is itself the cost the user is trying to avoid.</p><p><strong>Poor timing.</strong> Inline hints surface during onboarding, when the user has no context to attach them to. Static documentation sits behind a separate navigation path that requires the user to leave the workflow, find the relevant article, and then return to the chart. The help that would have been useful at minute three of the encounter arrives at minute eight, after the user has already worked around the problem.</p><p><strong>Intent mismatch.</strong> Help explains what the system does. The user wants to know what they should do. These are different questions. A help article titled &#8220;About Patient Status Codes&#8221; describes the feature. A help article titled &#8220;What to Do When a Transfer Status Code Appears&#8221; describes the task. Most help systems are organized around the first kind of title. Users come looking for the second.</p><p><strong>Verbosity.</strong> A user in the middle of a task does not want to read four hundred words to find a sentence. They want a sentence. Help that buries the answer in context is, for the in-task user, indistinguishable from help that does not have the answer at all.</p><p>Each pattern is a debit against the credibility balance. None of them, individually, is fatal. In combination, sustained over months of product use, they teach the user that consulting help is not worth the interruption.</p><p>This is a usability failure, not a user failure.</p><h2>What Users Do Instead</h2><p>When users avoid help, they compensate. The compensations are largely invisible to designers and evaluators because they happen outside the product surface.</p><p>They ask colleagues. They keep personal cheat sheets in OneNote or in a desk drawer. They develop superstitions about which actions are safe and which are risky. They route around features they do not trust themselves to operate correctly. They search the web for the question they would have asked the help system, and they often find better answers there than they would have inside the product.</p><p>These adaptations have costs. Institutional knowledge distributes into individual heads and does not survive turnover. Practice diverges across teams, and feature avoidance suppresses product value.</p><p>The system appears usable in aggregate metrics because users have adapted. Adaptation creates the appearance of usability without producing it. The gap is where most help-credibility debt accumulates.</p><p>What follows is the loop in three settings: an enterprise software migration, a consumer application with conversational help, and a regulated environment.</p><h2>Case: Enterprise Software Migration</h2><p>A mid-sized financial services firm migrates its case management system. The new system has a contextual help panel that opens as a sidebar when the user clicks an information icon next to any field.</p><p>Six weeks after rollout, an internal survey finds that 71 percent of analysts have never opened the help panel. Of the 29 percent who have, more than half describe the content as &#8220;not what I needed&#8221; or &#8220;too generic to use.&#8221; The product team adds more help articles. Open rates do not recover.</p><p>The diagnosis was wrong. The problem was not coverage. The problem was that the first three help articles the average analyst encountered, during the first week of the migration when stress was highest, had failed them. The credibility balance had gone negative early and the loop had closed. Additional articles, however well-written, were not being read because the channel itself had lost its audience.</p><p>The intervention that eventually moved the metric was not more help. It was a redesign of three specific high-friction workflows so that the most common questions were answered by interface elements at the moment of decision. The help panel was retained for genuine edge cases, but it was no longer load-bearing. Open rates remained low. Resolution rates improved.</p><p>The help channel was the wrong unit of optimization. The interaction was the right one.</p><h2>Case: Consumer Application with Conversational Help</h2><p>A consumer banking application introduces an in-product assistant that uses a large language model to answer user questions in natural language. Early metrics are positive. Engagement is high.</p><p>By month four, the engagement curve has flattened and complaint volume related to the assistant has grown. Investigation shows that the assistant has been confidently giving incorrect answers to a small but consistent set of edge cases involving account types the model was not well calibrated on. Users who encountered these incorrect answers stopped using the assistant. Some told colleagues and family members not to use it either. Word of mouth in this direction travels faster than in the other.</p><p>Generative systems accelerate the credibility loop. Conventional help fails by being unhelpful. Generative help can fail by being confidently wrong, which is a more expensive category of debit. A user who consults static documentation and finds nothing useful loses time. A user who consults a conversational assistant and acts on a wrong answer loses time, money, or, in clinical contexts, more.</p><p>Lee and See&#8217;s framework anticipates this. Trust calibration depends on the operator&#8217;s ability to assess the automation&#8217;s competence in context.&#8308; Generative help systems often present output in a uniform register of confidence regardless of whether the underlying answer is grounded or invented. The user has no signal to tell the two apart. When the signal is absent, calibration collapses, and disuse follows.</p><p>The design implication is that conversational help must invest heavily in uncertainty signaling, source attribution, and bounded scope. A system that says &#8220;I do not know&#8221; credibly is more usable, over time, than a system that always answers.</p><h2>Case: Regulated Environment</h2><p>In aviation and clinical settings, help consultation is sometimes mandated. Procedural references, checklists, and decision support tools are not optional in the same way they are in consumer software.</p><p>This shifts the credibility loop but does not eliminate it. Mandated consultation produces compliance behavior, which is not the same as use. A clinician required to acknowledge a decision support alert may dismiss it without reading it. A pilot required to consult a quick reference handbook may skim for the heading that matches the current situation and ignore the surrounding context. The credibility balance still operates, but it has migrated from the question &#8220;Should I consult this?&#8221; to the question &#8220;How carefully should I read it?&#8221;</p><p>In these environments, the design problem is not driving consultation. It is earning attention within consultation. The same diagnostic principles apply, with the same fixes: specificity, timing, intent alignment, brevity.</p><h2>What Changes When Help Becomes an Agent</h2><p>The migration of help into agentic systems, where the help layer can act on the user&#8217;s behalf rather than merely instruct, changes the stakes of the credibility loop in two ways.</p><p>First, the consequences of a debit grow. An agent that misreads the user&#8217;s intent and takes an incorrect action is not a help failure. It is a workflow failure that the user must then unwind. One bad agent interaction costs more credibility than several bad documentation interactions.</p><p>Second, the loop becomes harder to repair. A user who learns not to trust a documentation panel can still read the interface and recover. A user who learns not to trust an agent has been taught something about the product&#8217;s reliability that generalizes beyond the help layer. The credibility debt does not stay contained.</p><p>Agentic help systems require more conservative design than documentation panels for structural reasons. The technology is just as capable. The credibility economics are less forgiving.</p><p>The principle that follows is straightforward. Help should reduce cognitive burden, not add to it. An agent that adds burden, by introducing recovery work or by requiring the user to verify its actions, has inverted the value proposition.</p><h2>Boundary Conditions</h2><p>The Help Credibility Loop applies most cleanly in environments where the user has discretion over whether to consult help and where the product surface is the primary site of work. It applies with some modification in regulated environments, as noted above.</p><p>The loop applies less cleanly in three situations.</p><p>In high-novice populations, where users do not yet have a credibility balance because they have not yet had enough interactions to form one, help consumption is driven more by social cues, training context, and first-week onboarding design than by accumulated experience. The loop begins to operate by the end of the first month.</p><p>In short-tenure populations, where users churn out of the product before their credibility balance stabilizes, the diagnostic value of the loop is limited. Aggregate metrics in these environments will show low help engagement that is not actually a credibility problem. It is a tenure problem.</p><p>In compliance-heavy environments, the loop operates on attention quality rather than consultation frequency, as described above.</p><p>Practitioners running diagnostics on help systems should ask which condition their user population is in before applying the framework.</p><h2>What This Means in Evaluation</h2><p>Usability evaluations routinely note whether help exists and whether it is discoverable. These are necessary conditions and not sufficient ones. The harder and more diagnostic question is whether users choose to use it, and what their reasoning is when they do not.</p><p>Evaluators should ask three questions of the user after observing a help avoidance moment.</p><p>What did you expect help would tell you?</p><p>What did you do instead, and why?</p><p>When was the last time you opened help in this product, and what happened?</p><p>The answers reveal the credibility balance. They reveal which channel has been spent down and which still has credit. They reveal whether the problem is coverage, timing, intent, verbosity, or accumulated trust debt from prior interactions the evaluator did not observe.</p><p>Help usage is behavior, not a checkbox. It is observable, measurable, and diagnosable, but only if the evaluation is structured to surface it.</p><h2>What Practitioners Can Do This Quarter</h2><p>Three moves are available now without rebuilding the help layer from scratch.</p><p>First, audit the first three help articles a new user is most likely to encounter in the first week. These are the credibility-establishing interactions. If they are generic, poorly timed, or off-intent, the rest of the help system is operating in a structural deficit.</p><p>Second, identify the two or three highest-friction workflows in the product and ask whether the questions users have at those moments are better answered by help content or by interface changes. The help channel is often the wrong unit of optimization. The interaction is the right unit.</p><p>Third, instrument help avoidance, not just help consumption. Track moments where users hesitate at a friction point, do not consult help, and either resolve through workaround or abandon. These moments are the credibility loop running. They are also the largest single source of unrecognized usability debt in most mature products.</p><p>When help earns trust, users use it. When it does not, they quietly work around the system instead. The work of designing help, increasingly, is the work of restoring the channel&#8217;s credibility one interaction at a time.</p><h2>Quick Reference</h2><ul><li><p>Help avoidance is a credibility problem, not a discoverability problem.</p></li><li><p>The Help Credibility Loop has five steps: encounter, threshold, action, resolution, update.</p></li><li><p>The loop is asymmetric. Bad help interactions spend credit faster than good ones earn it.</p></li><li><p>Generative and agentic help systems raise the cost of each debit and require more conservative design.</p></li><li><p>Evaluate help usage as behavior, not as a feature checkbox.</p></li></ul><h2>References</h2><ol><li><p>Rettig, Marc. &#8220;Nobody Reads Documentation.&#8221; <em>Communications of the ACM</em> 34, no. 7 (July 1991): 19-24.</p></li><li><p>Carroll, John M. <em>The Nurnberg Funnel: Designing Minimalist Instruction for Practical Computer Skill.</em> Cambridge, MA: MIT Press, 1990.</p></li><li><p>Nielsen, Jakob. &#8220;10 Usability Heuristics for User Interface Design.&#8221; Nielsen Norman Group, 1994. https://www.nngroup.com/articles/ten-usability-heuristics/.</p></li><li><p>Lee, John D., and Katrina A. See. &#8220;Trust in Automation: Designing for Appropriate Reliance.&#8221; <em>Human Factors</em> 46, no. 1 (2004): 50-80.</p></li><li><p>Parasuraman, Raja, and Victor Riley. &#8220;Humans and Automation: Use, Misuse, Disuse, Abuse.&#8221; <em>Human Factors</em> 39, no. 2 (June 1997): 230-253.</p></li></ol><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/p/why-users-dont-use-help-even-when?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/p/why-users-dont-use-help-even-when?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[What Belongs in a Usability Report Besides the Metrics]]></title><description><![CDATA[How to Report the Mess Without Losing the Meaning]]></description><link>https://johnwbrown.substack.com/p/what-belongs-in-a-usability-report</link><guid isPermaLink="false">https://johnwbrown.substack.com/p/what-belongs-in-a-usability-report</guid><dc:creator><![CDATA[John W Brown]]></dc:creator><pubDate>Tue, 19 May 2026 06:02:16 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!n3Ou!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42b743d7-2de7-4de4-9b61-19427176abdc_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!n3Ou!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42b743d7-2de7-4de4-9b61-19427176abdc_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!n3Ou!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42b743d7-2de7-4de4-9b61-19427176abdc_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!n3Ou!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42b743d7-2de7-4de4-9b61-19427176abdc_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!n3Ou!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42b743d7-2de7-4de4-9b61-19427176abdc_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!n3Ou!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42b743d7-2de7-4de4-9b61-19427176abdc_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!n3Ou!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42b743d7-2de7-4de4-9b61-19427176abdc_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/42b743d7-2de7-4de4-9b61-19427176abdc_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2179201,&quot;alt&quot;:&quot;Taking notes&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/190653304?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42b743d7-2de7-4de4-9b61-19427176abdc_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Taking notes" title="Taking notes" srcset="https://substackcdn.com/image/fetch/$s_!n3Ou!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42b743d7-2de7-4de4-9b61-19427176abdc_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!n3Ou!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42b743d7-2de7-4de4-9b61-19427176abdc_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!n3Ou!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42b743d7-2de7-4de4-9b61-19427176abdc_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!n3Ou!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42b743d7-2de7-4de4-9b61-19427176abdc_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Image created with generative AI</figcaption></figure></div><div><hr></div><p><em>The views expressed are my own and do not represent any organization I&#8217;m affiliated with.</em></p><div><hr></div><p>Consider a usability session that ends with a participant completing all assigned tasks within the time threshold, logging zero critical errors, and rating confidence above the midpoint on each item. On paper, the session is a success.</p><p>Then the facilitator&#8217;s notes surface. The participant hesitated visibly before each high-consequence action. She said twice, unprompted, that she would &#8220;double-check with the team&#8221; before submitting anything through the system. When asked to walk through her reasoning, she described calling a colleague as a routine step: not a workaround, just how things actually worked. The system was completing the tasks. The participant was not trusting the results.</p><p>If the report contains only the metrics, the session looks fine. The product looks ready. The mismatch between what the interface produced and what the user believed about those results goes unaddressed until it surfaces elsewhere, making it less convenient.</p><p>That is not a reporting problem. That is a reporting choice.</p><div><hr></div><p>Usability reports often narrow too quickly around the most convenient forms of evidence. Task completion rates, time on task, error counts, confidence ratings, and severity rankings all matter. They help organize findings, support comparison across participants, and give stakeholders something familiar to scan.</p><p>But anyone who has moderated real sessions knows the most important parts of a study do not always fit neatly into those boxes.</p><p>A participant completes the task but clearly does not trust the result. Another goes off script and starts describing who they would call for help in real life. A third hesitates repeatedly, not because the next step is hidden, but because the consequences of being wrong feel unclear. The observation team then debates what the conduct meant. Was it navigation failure, instruction ambiguity, learned local workflow, or simple distrust of the interface?</p><p>Those moments are not reporting problems. They are part of the work.</p><p>The mistake is not that usability testing produces messy evidence. The mistake is pretending it does not.</p><p>A good report should absolutely include metrics when they are appropriate. But it should also capture the parts of testing that explain what the metrics cannot. The challenge is doing that without turning the report into a running transcript of every opinion, every debate, or every moment of facilitator intuition.</p><p>The goal is not to report everything. The goal is to report the right non-metric insight in a way that remains useful, credible, and disciplined.</p><div><hr></div><h2>Metrics are useful, but they do not explain themselves</h2><p>There is a reason metrics dominate usability reports. They are compact. They travel well in presentations. They support comparison. They help answer familiar questions. Could users complete the task? How long did it take? Where did errors occur? How often did the issue appear?</p><p>Those are important questions. In many studies, they are central.</p><p>But metrics are not self-explanatory. A completion rate does not tell you whether the participant trusted what they just did. A timing measure does not tell you whether hesitation came from confusion, caution, prior bad experience, or fear of consequences. Even a severe failure may remain poorly understood unless the report explains what the participant thought was happening.</p><p>This connects to a problem Hertzum and Jacobsen documented in their work on the evaluator effect: different evaluators observing the same session routinely surface different sets of problems, and neither set fully captures what happened. Metrics appear precise because they are standardized. But standardization compresses exactly the kind of contextual variation that often explains the finding.</p><p>Two participants can produce the same metric outcome for very different reasons. One may fail because the navigation structure is misleading. Another may fail because the labels conflict with local terminology. A third may succeed only because prior training compensates for weak design. If the report includes only the numeric outcome, the real design problem may stay blurred.</p><p>That is why a report built entirely on metrics often looks cleaner than the underlying evidence really is. The numbers may be precise, but the interpretation may still be incomplete.</p><p>Good reporting does not discard metrics. It puts them in context.</p><div><hr></div><h2>What belongs in the report besides the metrics</h2><p>The answer is not &#8220;everything the team noticed.&#8221; It is the subset of observations and interpretations that materially help explain participant conduct, clarify a finding, or guide action.</p><p>Four kinds of non-metric content are especially useful.</p><p>The first is <strong>behavioral context</strong>: hesitation patterns, visible confidence drops, fallback responses, workaround tendencies, and the point where a participant appears to shift from solving the task to escaping the problem. These observations often explain why a finding matters.</p><p>The second is <strong>facilitator observation</strong>: the moderator is the person closest to the session's live dynamics. That perspective adds value by capturing aspects that may not be clear in the task table, especially around trust, uncertainty, and coping patterns.</p><p>The third is <strong>off-script conduct</strong>: when participants move outside the intended task path, that should not always be treated as noise. Sometimes it reveals how the task would actually be completed in the real world, including reliance on coworkers, support staff, memory, notes, or external search.</p><p>The fourth is <strong>observer disagreement</strong>: teams do not always interpret what they saw the same way. When that disagreement is meaningful, it belongs in the report because it signals that a conclusion should be framed with more caution.</p><p>Each of these can belong in a usability report. The question is how to present them so the report stays readable and credible.</p><div><hr></div><h2>Separate observation from interpretation</h2><p>One of the simplest ways to report the mess without losing the meaning is to separate what happened from what the team thinks it means.</p><p>At a minimum, the report should distinguish:</p><p><strong>Observed behavior:</strong> What the participant actually did or said.</p><p><strong>Interpretive context:</strong> What the facilitator or observation team thinks the conduct may indicate.</p><p><strong>Analytic takeaway:</strong> What conclusion, recommendation, or follow-up question the report ultimately draws.</p><p>That structure matters because it prevents observation and inference from collapsing into a single unsupported claim.</p><p>Consider a participant who stops pursuing the workflow and says she would call a coworker for help.</p><p>A weak report might say: &#8220;The participant did not trust the system.&#8221;</p><p>That may be true, but it is still an inference.</p><p>A stronger version would look more like this:</p><p><strong>Observed behavior:</strong> The participant stopped pursuing the in-system workflow and stated that in actual practice, she would call a coworker for help.</p><p><strong>Interpretive context:</strong> The facilitator viewed this as a likely loss of confidence in the system&#8217;s ability to support unaided completion. One observer agreed. Another thought the response may have reflected a learned local habit rather than distrust of the interface itself.</p><p><strong>Analytic takeaway:</strong> Regardless of the cause, the session indicates that independent task completion may depend on external support, suggesting the interface does not provide sufficient clarity or confidence for self-sufficient use.</p><p>That version is more useful because it preserves the evidence, acknowledges interpretation, and still gives the reader a clear conclusion. The discipline of labeling what is observation and what is inference is not bureaucratic overhead. It is how the report earns the reader&#8217;s trust.</p><div><hr></div><h2>Facilitator comments can add value if they stay in bounds</h2><p>Some practitioners avoid facilitator comments because they worry that those comments make a report feel subjective. That concern is fair. A facilitator&#8217;s perspective should not replace evidence or overpower it.</p><p>Still, facilitator observations add value when they serve a clear purpose.</p><p>The moderator sees things that may not be clear in a transcript or score table: confidence drops, repeated hesitation before high-consequence actions, changes in tone when participants stop trusting the interface, and moments when a participant appears to give up mentally before the formal outcome is recorded. The moderator also knows when and why a session shifted from structured task completion into exploratory discussion.</p><p>Those observations can help explain the formal findings.</p><p>The key is to label them correctly. Facilitator comments should be presented as informed observations, not as objective task data. They should stay close to observable actions and statements. They should not drift into personality judgments, overconfident conclusions, or ungrounded speculation.</p><p>A short section titled <strong>"Facilitator Observations"</strong>&nbsp;or&nbsp;<strong>"Facilitator Insights"</strong> can work well for this. Used carefully, it gives the report a place to capture context without confusing interpretation with measurement.</p><div><hr></div><h2>Off-script conduct is often part of the findings</h2><p>One common reporting mistake is treating off-script sessions as disruptions to clean out of the final write-up.</p><p>Sometimes participants go off script because they are distracted or misunderstood the task. But sometimes the departure is the data.</p><p>A participant says she would search externally. Another says he would call Robert. A third explains that no one actually completes the task this way and describes the workaround her team uses in practice. In a narrow metric frame, these look like lost data points. In actual usability work, they may be among the most important observations in the study.</p><p>They reveal dependency. They reveal external recovery strategies. They reveal the gap between formal workflow and lived workflow. That gap is often where the real design problem is located.</p><p>A report should say that plainly.</p><p>That does not mean pretending the session still supports clean task metrics. It means being honest that the session shifted, and that the shift produced qualitative insight worth preserving. If a participant&#8217;s first instinct is to leave the system and seek external support, the interface is only part of the real task environment. That belongs in the report.</p><div><hr></div><h2>Briefly noting disagreement can increase credibility</h2><p>Many teams smooth over disagreements because it looks messy. It can make the team seem uncertain. It complicates the summary. It adds nuance where stakeholders often want simplicity.</p><p>But not all disagreements should be hidden.</p><p>One observer may think a participant failed because the label was unclear. Another may think the prompt itself caused the confusion. The facilitator may think the deeper issue was not comprehension at all, but reluctance to act without confidence. Each explanation may be plausible. In that situation, briefly acknowledging the disagreement makes the report more credible, not less.</p><p>A line such as, &#8220;The observation team differed on whether this reflected navigation failure or instruction ambiguity,&#8221; is often enough. Another useful line: &#8220;Because interpretations varied, this finding should be treated as directional and explored further in follow-up testing.&#8221;</p><p>That tells the reader two things. First, the team did not force certainty where certainty did not exist. Second, the ambiguity itself may matter for the next steps.</p><p>The important word is <em>briefly</em>. Not every internal debate belongs in the report. But when disagreement changes the meaning, confidence level, or recommended action for a finding, it is worth noting.</p><div><hr></div><h2>Keep the report disciplined</h2><p>Once reports begin to include facilitator comments, off-script moments, and observer disagreement, there is a risk that they will become loose, overlong, or overly personal. That risk is real. It can be managed.</p><p>Anchor everything to observable evidence. Even interpretive material should point back to what the participant actually did or said.</p><p>Label interpretive material clearly. Readers should not have to guess whether they are reading a metric, a quote, a facilitator note, or a team inference.</p><p>Use disagreement selectively. Include it when it changes understanding, not merely because it happened.</p><p>Avoid personality judgments. Findings should remain focused on actions, context, and design implications, rather than on assessments of individual participants.</p><p>Still provide an analytic takeaway. A report should not become a catalog of unresolved opinions. Even when evidence is mixed, the reader needs to know what the most responsible conclusion is, or whether the right conclusion is that follow-up testing is needed.</p><p>And finally, edit. Honest reporting is not the same as unfiltered reporting. The goal is to reduce uncertainty for the reader, not to transfer it.</p><div><hr></div><h2>A structure that works in practice</h2><p>For teams that want to capture more than metrics without losing coherence, a simple pattern can help. Each major finding can include:</p><p><strong>Observed behavior:</strong> A concise description of what happened.</p><p><strong>Why it matters:</strong> A short explanation of why the pattern is significant.</p><p><strong>Additional context:</strong> A brief facilitator note, off-script moment, or relevant disagreement, if it materially affects interpretation.</p><p><strong>Recommendation or next step:</strong> What the team thinks should happen in response.</p><p>This keeps the report readable while preserving nuance.</p><p>Another option is a short section near the end of the report titled <strong>Facilitator Observations</strong>, <strong>Session Context</strong>, or <strong>Interpretive Notes</strong>. That section can hold patterns that do not belong in metric tables but still deserve attention: repeated confidence drops, recurring external dependencies, and moments when the observation team saw the same actions differently.</p><p>The exact structure matters less than the discipline behind it. The reader should be able to tell what happened, what the team thinks it means, and how confident the team is in that interpretation.</p><div><hr></div><h2>Why this matters</h2><p>Usability work is often expected to do two things at once: produce defensible evidence and help teams understand reality. Those goals overlap, but they are not identical.</p><p>Metrics support defensibility. Context supports understanding.</p><p>A report that contains only the first may be neat but shallow. A report that contains only the second may be vivid but hard to act on. Good reporting requires both.</p><p>That is especially true in complex settings, where user responses are shaped not just by screens and labels, but by fear of consequences, role expectations, prior training, institutional memory, and outside support structures. In those contexts, the most important finding may not be that the user took twelve seconds too long. It may be that they never intended to trust the interface on its own in the first place.</p><p>If the report excludes that kind of insight because it does not fit cleanly into a metric format, it may satisfy the template while missing the real lesson.</p><p>Usability findings are often cleaner in spreadsheets than they are in sessions. Reports that pretend otherwise do not protect the reader from the mess. They just move it downstream, where it costs more to fix.</p><div><hr></div><p>#UXDesign #HumanFactors #HSI #UsabilityTesting #UXResearch</p><div><hr></div><p></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/p/what-belongs-in-a-usability-report?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/p/what-belongs-in-a-usability-report?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[What UX Metrics Assume, and When Those Assumptions Break]]></title><description><![CDATA[Five hidden conditions in the numbers we trust most]]></description><link>https://johnwbrown.substack.com/p/what-ux-metrics-assume-and-when-those</link><guid isPermaLink="false">https://johnwbrown.substack.com/p/what-ux-metrics-assume-and-when-those</guid><dc:creator><![CDATA[John W Brown]]></dc:creator><pubDate>Tue, 12 May 2026 06:02:06 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!R1h2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4b5713-9ef4-483e-93bb-a6133d04bb34_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!R1h2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4b5713-9ef4-483e-93bb-a6133d04bb34_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!R1h2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4b5713-9ef4-483e-93bb-a6133d04bb34_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!R1h2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4b5713-9ef4-483e-93bb-a6133d04bb34_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!R1h2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4b5713-9ef4-483e-93bb-a6133d04bb34_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!R1h2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4b5713-9ef4-483e-93bb-a6133d04bb34_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!R1h2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4b5713-9ef4-483e-93bb-a6133d04bb34_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9a4b5713-9ef4-483e-93bb-a6133d04bb34_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2925669,&quot;alt&quot;:&quot;The Hidden Assumptions Inside Common UX Metrics&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/187396406?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4b5713-9ef4-483e-93bb-a6133d04bb34_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The Hidden Assumptions Inside Common UX Metrics" title="The Hidden Assumptions Inside Common UX Metrics" srcset="https://substackcdn.com/image/fetch/$s_!R1h2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4b5713-9ef4-483e-93bb-a6133d04bb34_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!R1h2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4b5713-9ef4-483e-93bb-a6133d04bb34_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!R1h2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4b5713-9ef4-483e-93bb-a6133d04bb34_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!R1h2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4b5713-9ef4-483e-93bb-a6133d04bb34_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Image created with generative AI</figcaption></figure></div><div><hr></div><p><em>The views expressed are my own and do not represent any organization I am affiliated with.</em></p><h2>A familiar dashboard</h2><p>Consider a usability dashboard reviewed in a quarterly product meeting. Task completion is up. Time on task is down. Error counts are stable. Satisfaction scores have crept past 4 out of 5. The team congratulates itself on a quarter of steady gains and moves on to the next agenda item.</p><p>Three months later, support volume rises. A handful of customers describe checking the system&#8217;s outputs against a separate spreadsheet before acting on them. Two enterprise accounts have built parallel tools. The product team is puzzled. The metrics did not move.</p><p>Nothing on the dashboard is wrong. The numbers are what they are. The problem is that each of those numbers was answering a question the team had stopped asking, under conditions the team had stopped checking. Every metric carried an assumption about how the system worked, how users behaved, and what counted as success. When those assumptions held, the metrics meant what they appeared to mean. When the assumptions stopped holding, the metrics kept producing numbers anyway.</p><p>This is the part of measurement work that gets the least attention. We talk about which metrics to collect. We talk less about what those metrics quietly take for granted.</p><h2>The conditions metrics were built for</h2><p>Most standard usability metrics were developed in largely deterministic environments. A task had a defined start and end. The system gave the same response to the same input. Errors were observable, often immediate, and almost always attributable to a discrete moment. Success could be defined as reaching the correct endpoint within an acceptable window of time and effort.</p><p>Time on task reflected efficiency. Task success reflected effectiveness. Error counts reflected breakdowns. Satisfaction loosely tracked perceived quality. The instruments were calibrated for the environment in which they were built.</p><p>Lisanne Bainbridge&#8217;s &#8220;Ironies of Automation&#8221; warned more than four decades ago that automating parts of a system does not eliminate the human role. It changes that role, often in ways the original measurement scheme is not equipped to see. The skills the operator needs change. The failures shift in shape. The places where understanding has to live shift, too. A measurement program that keeps recording the same numbers in the same way after that shift is not measuring the new system. It is measuring what the old system used to do.</p><p>Metrics are not wrong. They are conditional. Their meaning depends on conditions that are rarely documented and even less often audited.</p><div><hr></div><h2>The five hidden assumptions</h2><p>Five assumptions sit underneath the metrics most usability programs collect. Naming them helps practitioners audit when and where their dashboards stop telling the truth.</p><h3>1. Stable outcomes</h3><p>Standard task-success metrics assume that a successful interaction generalizes. Complete the task once, and the system behaves the same way the next time. The user&#8217;s understanding of what just happened remains valid the day after.</p><p>Adaptive and inferential systems weaken this assumption. A user may complete a task successfully and receive a different outcome the next time under conditions the user cannot identify, and the system does not surface. The completion still registers. The confidence does not carry forward. Task success captures what happened. It does not capture whether the result can be relied upon.</p><p>This assumption breaks down earliest in systems where the model behind the interface changes more frequently than the interface itself.</p><h3>2. Visible failure</h3><p>Most usability metrics are built around visible signals. An error message. An abandoned task. A complaint. A drop in completion. The implicit model is that when something goes wrong, something registers.</p><p>Failures in adaptive systems often go unnoticed. A user who completes a task but verifies the result somewhere else has succeeded by the dashboard&#8217;s reckoning. A user who distrusts an output but proceeds anyway leaves no trace. A user who quietly stops using a feature generates no error event, only an absence that the system was not asked to look for.</p><p>Elsewhere, I have argued that self-checking is an unreported usability signal. The same principle generalizes here. The dashboard sees the actions that succeeded. It does not see the verifications that made those actions feel safe.</p><h3>3. Speed as quality</h3><p>Time on task is treated as a proxy for efficiency. Faster numbers point to a smoother interface. Slower numbers point to friction. This holds when tasks are well-defined, and the consequences of getting them wrong are bounded.</p><p>In high-stakes or interpretive work, the relationship inverts. Slower performance can reflect the careful judgment the work actually requires. Faster performance can reflect overconfidence in an output that the user did not have time to evaluate. A time-on-task speed-up in a clinical decision-support tool is not unambiguously good news. It may mean the interface is clearer. It may also mean that clinicians have stopped scrutinizing the recommendation.</p><p>A metric that rewards speed without context will eventually punish the users who slow down for the right reasons.</p><h3>4. Satisfaction as trust</h3><p>Satisfaction scores are sensitive to framing, timing, and emotional state. A user surveyed immediately after completing a task may report high satisfaction even while harboring serious doubts about the outcome. Politeness inflates scores. Relief at finishing inflates scores. Low expectations inflate scores more than anything else.</p><p>Lee and See, in their work on trust in automation, drew a distinction worth importing into usability practice. Satisfaction is an affective response to the immediate experience. Trust is a calibrated belief about whether the system can be relied upon for a class of decisions over time. The two are related but distinct, and they diverge in exactly the situations where the difference matters most.</p><p>A system can score above 4 out of 5 in satisfaction while users quietly engineer their workflows to depend on it as little as possible. Satisfaction reports the session. Whether the user will rely on the system tomorrow, recommend it to a colleague, or accept its output under pressure is a separate question, asked by a different instrument.</p><h3>5. Independent action</h3><p>Standard metrics assume that the recorded interaction is the whole interaction. The user sat down, used the system, finished or did not, and that was the session. The data the system captured is the data that is to be captured.</p><p>In practice, professional users routinely supplement primary systems with side channels. A second monitor is open to a reference tool. A colleague consulted on Slack. A spreadsheet is maintained on the side because the system of record cannot be trusted for one specific calculation. These behaviors are invisible to the system that thinks it owns the workflow.</p><p>When supplementary channels do the cognitive work the primary system claims to do, the primary system's metrics look better than it deserves. The work is happening. It is just not happening where the dashboard is looking.</p><div><hr></div><h2>A composite case</h2><p>Consider a clinical decision-support module deployed in an outpatient practice. The module reviews medication orders and surfaces potential interactions, contraindications, and renal-dosing concerns. After two quarters in production, the data looks favorable. Alert acknowledgment rates are within target ranges. Override rates have stabilized. Time per medication order is down by twelve percent. Provider satisfaction with the module sits at 4.2 out of 5.</p><p>A focused observational study, run a quarter later, finds a different picture under each of the five assumptions.</p><p>The override rate held steady through two retrainings of the model behind it. Several physicians describe being surprised by alerts that previously would not have fired and missing alerts they had come to expect. The number on the dashboard stayed flat. What the number was measuring did not.</p><p>Failure is not visible in the dashboard. Three of the eight observed physicians keep a printed reference card for high-risk drug classes near the workstation. Two describe pulling up an external interaction checker for any medication they have not prescribed in the past month. None of this verification work is captured by the module. It is captured nowhere.</p><p>The twelve percent improvement is real. The shape of it is the issue. The gain is concentrated in the highest-volume prescribers, several of whom describe acknowledging alerts without reading them when the workflow is busy. The interface is faster. The decision behind it is shallower.</p><p>Survey scores have outrun reliance. The numbers sit above target. In conversation, several of the same physicians describe the module as &#8220;useful, but I always double-check the ones that matter.&#8221; Asked whether they would feel comfortable if the module ran in a more autonomous mode, none said yes.</p><p>Independent action is the largest blind spot. The printed cards, the external checker, the colleague consulted in the hallway, and the pharmacist queried by phone are all doing decision work that the dashboard credits to the module.</p><p>The module is not failing. It is also not performing the way the dashboard says it is. The gap between those two statements is exactly the territory standard metrics cannot reach.</p><div><hr></div><h2>Stating the conditions</h2><p>The fix is not to abandon the metrics. The instruments work for what they were designed to do. The fix is to state the conditions under which they hold and to write those conditions next to the numbers.</p><p>A task-success rate is meaningful when outcomes are stable, when failure is observable, and when the user is acting alone with the system of record. When any of those conditions weakens, the rate continues to compute. Its meaning does not.</p><p>This is the kind of practitioner discipline that does not show up well in a slide. It shows up in the small footnote next to the number that says: this metric assumes the model has not been retrained this quarter, that users have access to the same external tools they did at baseline, and that workflow conditions match the conditions under which the benchmark was set. A report with that footnote is harder to read at a glance and considerably harder to mislead with.</p><p>The most useful thing a usability program can do with its metrics is to write those footnotes in advance, audit them on a regular cadence, and treat the unstated assumption as the place where the next failure is most likely to hide.</p><div><hr></div><h2>Quick reference</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!P2Kx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0dd7ae19-daa6-498f-b6f7-fcb20fff896a_701x564.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!P2Kx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0dd7ae19-daa6-498f-b6f7-fcb20fff896a_701x564.png 424w, https://substackcdn.com/image/fetch/$s_!P2Kx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0dd7ae19-daa6-498f-b6f7-fcb20fff896a_701x564.png 848w, https://substackcdn.com/image/fetch/$s_!P2Kx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0dd7ae19-daa6-498f-b6f7-fcb20fff896a_701x564.png 1272w, https://substackcdn.com/image/fetch/$s_!P2Kx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0dd7ae19-daa6-498f-b6f7-fcb20fff896a_701x564.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!P2Kx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0dd7ae19-daa6-498f-b6f7-fcb20fff896a_701x564.png" width="701" height="564" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0dd7ae19-daa6-498f-b6f7-fcb20fff896a_701x564.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:564,&quot;width&quot;:701,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:86921,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/187396406?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0dd7ae19-daa6-498f-b6f7-fcb20fff896a_701x564.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!P2Kx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0dd7ae19-daa6-498f-b6f7-fcb20fff896a_701x564.png 424w, https://substackcdn.com/image/fetch/$s_!P2Kx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0dd7ae19-daa6-498f-b6f7-fcb20fff896a_701x564.png 848w, https://substackcdn.com/image/fetch/$s_!P2Kx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0dd7ae19-daa6-498f-b6f7-fcb20fff896a_701x564.png 1272w, https://substackcdn.com/image/fetch/$s_!P2Kx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0dd7ae19-daa6-498f-b6f7-fcb20fff896a_701x564.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>Close</h2><p>The numbers on the dashboard are not the problem. The conditions underneath them are. A usability program that audits its assumptions as carefully as it tracks its metrics learns where its instruments stop measuring the system and start measuring an older version of it.</p><p>The most consequential usability findings are usually the ones the dashboard was not built to see.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/subscribe?"><span>Subscribe now</span></a></p>]]></content:encoded></item><item><title><![CDATA[From Canonical UI to Canonical Substrate]]></title><description><![CDATA[Why AI may personalize the computer, not just the screen]]></description><link>https://johnwbrown.substack.com/p/from-canonical-ui-to-canonical-substrate</link><guid isPermaLink="false">https://johnwbrown.substack.com/p/from-canonical-ui-to-canonical-substrate</guid><dc:creator><![CDATA[John W Brown]]></dc:creator><pubDate>Tue, 05 May 2026 06:02:23 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!_NXe!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bb628fa-9512-41f8-847d-b8b58dd8f0f1_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_NXe!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bb628fa-9512-41f8-847d-b8b58dd8f0f1_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_NXe!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bb628fa-9512-41f8-847d-b8b58dd8f0f1_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!_NXe!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bb628fa-9512-41f8-847d-b8b58dd8f0f1_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!_NXe!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bb628fa-9512-41f8-847d-b8b58dd8f0f1_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!_NXe!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bb628fa-9512-41f8-847d-b8b58dd8f0f1_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_NXe!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bb628fa-9512-41f8-847d-b8b58dd8f0f1_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6bb628fa-9512-41f8-847d-b8b58dd8f0f1_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2338709,&quot;alt&quot;:&quot;Interface stack&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/195054869?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bb628fa-9512-41f8-847d-b8b58dd8f0f1_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Interface stack" title="Interface stack" srcset="https://substackcdn.com/image/fetch/$s_!_NXe!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bb628fa-9512-41f8-847d-b8b58dd8f0f1_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!_NXe!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bb628fa-9512-41f8-847d-b8b58dd8f0f1_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!_NXe!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bb628fa-9512-41f8-847d-b8b58dd8f0f1_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!_NXe!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bb628fa-9512-41f8-847d-b8b58dd8f0f1_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Image created with generative AI</figcaption></figure></div><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/p/from-canonical-ui-to-canonical-substrate?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/p/from-canonical-ui-to-canonical-substrate?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><p>Consider two clinicians in the same hospital, same shift, each opening the same patient&#8217;s record at a workstation. They see different screens. Not just different content. Different structure. One sees a dense table of labs, vitals, and medication history anchored at the top of the view, an information layout she has built up over ten years and trusts to surface what matters without further navigation. The other, three years into practice, sees a staged summary with the same underlying data disclosed progressively, sequenced around a workup the system infers he is preparing to perform. The facts beneath both views are identical. The permissions, the audit trail, the safety logic, the accountability chain: all identical. Only the presentation has moved.</p><p>Is this better care, or a coordination risk waiting to happen?</p><p>The answer depends on something the screens themselves do not reveal. It depends on whether the hospital has explicitly decided what may vary between the two views and what must not. It depends on whether a nurse handing off the patient can know, with confidence, what either clinician saw. It depends on whether &#8220;the interface&#8221; is still the main thing being designed, or whether the real design work has moved somewhere else.</p><p>Jakob Nielsen has argued that AI represents the first genuinely new UI paradigm in decades.[^1] The framing is somewhat dramatic, but the underlying point is worth taking seriously. If users increasingly tell systems what they want rather than step through fixed menus, windows, and commands, then the center of design begins to shift. The interface does not disappear, but it stops being the sole or even primary object of concern. What matters more is the system&#8217;s ability to interpret intent, assemble capabilities, and present an interaction model that fits the person, the role, and the task. Nielsen Norman Group has extended that argument through its work on generative UI, describing a shift away from designing one static experience for the average user and toward defining constraints within which systems can generate tailored experiences in real time.[^2]</p><p>That distinction matters because most personalization today is fairly shallow. Systems remember a preference, recommend content, save a layout, or offer a few display options. Useful, yes, but still operating inside a fixed shell. The more consequential possibility is not personalization at the content layer, but individualization at the interaction layer and perhaps even at the software layer. In that future, the user is not simply selecting from a pre-authored application. The system is helping assemble the application around the user&#8217;s goals and circumstances. A person who wants a word processor may not open one canonical product that everyone else uses. They may instead receive a writing environment assembled around the kind of writing they do, the tools they rely on, the structure they need, and the level of guidance they prefer. NN/g&#8217;s recent distinction between generative UI and vibe coding helps here: in one case the system generates the interface or product in response to a need; in the other the user directs the AI to build it.[^3] Either way, the visible artifact becomes more contingent and less fixed than the conventional software product.</p><p>This is why claims about the death of the UI are both overstated and useful. They are overstated because users will still need representations, controls, feedback, constraints, recoverability, and help. Those do not vanish. Nielsen Norman Group&#8217;s work on the articulation barrier makes the point directly: prompt-based interaction by itself often places too much burden on the user, and hybrid interfaces that combine prompt input with visible controls can reduce cognitive load and improve discoverability.[^4] AI is unlikely to abolish interface design. What it does threaten is the assumption that the stable screen is always the main thing being designed.</p><p>Healthcare is one of the clearest places to test this idea because it shows both the need for adaptation and the danger of getting it wrong. AHRQ has highlighted research showing that allowing clinicians to customize what they see in the EHR can reduce cognitive burden, save time, and support decision making in busy clinical environments.[^5] In one study led by Yalini Senathirajah at the University of Pittsburgh School of Medicine, a widget-based environment called MedWISER allowed clinicians to compose the information they wanted onto a single screen, pulling together lab results, imaging, notes, and medication data in the arrangement most useful to them. Clinicians could also draw on preset views created by colleagues in the same health system, effectively sharing the curation work rather than each rebuilding their own dashboard. The research found that this approach reduced cognitive burden, saved time, and supported decision-making in high-stress clinical environments. Importantly, the customization in that study was driven by clinicians rather than imposed on them, which is a different situation from generating an interface around inferred intent. AHRQ has also linked EHR usability problems directly to patient safety concerns,[^6] and ONC frames usability and provider burden as ongoing policy and operational issues, not cosmetic matters.[^7] Healthcare already demonstrates that interface design is part of the cognitive and safety environment of work.</p><p>That point becomes stronger once we move beyond traditional role-based design. It is obvious that a physician, pharmacist, nurse, scheduler, and patient should not all face the same interface. The less obvious claim is that even within a single role, the same screen may not serve all users equally well. Return to the two clinicians from the opening. The experienced physician benefits from dense summarization and quick access to exceptions, because pattern recognition is doing most of her work. The early-career clinician benefits from scaffolded sequencing, because he is still building the mental models that let dense views become useful. A third clinician with a different specialty may care about an entirely different subset of the record. The historical compromise of enterprise software has been to build one interface that is acceptable for many users while being ideal for very few. AI makes a different model more plausible: stable meaning underneath, more adaptive presentation on top. That is not merely a nicer UI. It is a different operating assumption about what software is. The product becomes less a fixed artifact and more a governed capability space.</p><p>The strongest argument against this picture comes from outside the AI literature. In 1983, Lisanne Bainbridge observed what she called the ironies of automation: that automating a process does not eliminate the human operator&#8217;s cognitive demands so much as redistribute them, often making the remaining human work harder and less supported.[^8] The same irony applies here. Generating the interface does not remove the need for design judgment. It moves that judgment upstream, into the specification of what the system may and may not generate. The practitioner&#8217;s cognitive burden does not disappear either. It shifts from &#8220;learn this screen&#8221; to &#8220;understand what this screen is allowed to be,&#8221; which is a subtler demand. Hertzum and Jacobsen&#8217;s work on the evaluator effect, showing that different evaluators applied to the same system surface different usability issues, points to a related concern: when an interface is generated rather than fixed, what exactly are we evaluating, and across how many instances does that evaluation need to hold?[^9]</p><p>These concerns are not decorative. In healthcare, handoffs, supervision, training, audit, and safety all depend on some degree of common structure. If every user sees a radically different world, coordination becomes harder and risk increases. Discussions of adaptive UI often become too loose here. They assume that if the visible interface varies, standardization disappears. It does not. It simply moves down a layer.</p><h2>Where the real risks live</h2><p>A clean-eyed look at the substrate model surfaces at least six risks worth naming, because they are what practitioners will actually have to manage.</p><p><strong>Inconsistency and coordination.</strong> Two clinicians discussing the same patient may be looking at materially different presentations. Without a shared reference frame, miscommunication becomes more likely, particularly in time-pressured environments.</p><p><strong>Training burden.</strong> Onboarding a new staff member to the EHR is already hard. Onboarding them to an EHR whose interface is generated per user is harder, unless the training target shifts from the screen to the underlying capability model. In practice this means teaching the system&#8217;s grammar rather than its screens: what the system can do, what it will not do, how it announces a mode change, and how to recover a canonical view when the generated one falters.</p><p><strong>Audit and accountability.</strong> When an adverse event is reviewed months later, the first question is often what the user actually saw at the moment of decision. In a fixed-interface system, reconstructing the view is straightforward. In a generated-interface system, it depends on whether the specific arrangement, the prompt or inferred intent that triggered it, the model version, and the data snapshot at that moment were all captured. Regenerating the exact view becomes a nontrivial engineering requirement, and one with legal weight.</p><p><strong>Handoff and supervision.</strong> Residents working under attending supervision share work through a shared representation. If the representation varies, supervision becomes more expensive and less reliable.</p><p><strong>Adversarial surface.</strong> A generated interface can, in principle, be manipulated by carefully crafted inputs or contextual cues in a way that a fixed interface cannot. Data in the record, content in an attached document, or even another user&#8217;s notes could, under the wrong architecture, influence what the system chooses to display to a downstream clinician. This is a new failure mode, not a reframed old one.</p><p><strong>Fallback design.</strong> When adaptation becomes confusing, unsafe, or simply unhelpful, there must be a documented path back to a stable, canonical view. That fallback itself has to be designed, tested, and trained on.</p><p>None of these risks argue against the substrate model. They argue for disciplined implementation of it.</p><p>That lower layer, where standardization actually lives, is what I would call the canonical substrate. In traditional software, the baseline is usually imagined as the default interface: the common navigation model, the shared workflow, the familiar menu structure, the screen everyone can be trained on. In a more AI-mediated future, that may no longer be the right place to anchor consistency. The stable reference point may instead be the underlying data model, action model, permission structure, safety logic, provenance model, and audit trail. In healthcare, that means the clinical meaning of data, the allowed actions, the conditions under which automation is permitted, the logging of recommendations, the explanation of system outputs, and the accountability chain all need to remain stable even if the presentation does not. The baseline, in other words, may no longer be a canonical interface. It may be a canonical substrate. This framing fits well with both the NIST AI Risk Management Framework, which treats trustworthy AI as a socio-technical matter shaped by validity, safety, accountability, transparency, explainability, privacy, and fairness in context,[^10] and ONC&#8217;s HTI-1 rule, which adds transparency and risk-management expectations around predictive decision support interventions in certified health IT.[^11]</p><h2>The substrate audit</h2><p>If the substrate is where standardization now lives, then evaluation needs a matching structure. Traditional heuristics still apply. Clarity, feedback, error prevention, cognitive load: none of these go away. But they have to be paired with a second set of questions aimed at invariants. In practical work, three tiers tend to be useful.</p><p><strong>Tier 1: Always visible.</strong> What must always be present, in any generated variant, before a high-risk action can be taken? In a clinical setting this might include the patient&#8217;s current medication list, known allergies, pending orders, and any active safety flags. These are load-bearing signals that cannot be summarized away without creating real risk.</p><p><strong>Tier 2: Adaptable within limits.</strong> What may the system adapt, and within what boundaries? Density, sequencing, terminology level, disclosure pattern, and visual grouping are strong candidates. What should not be adaptable without explicit user action: the definition of a clinical term, the threshold for an alert, the permissions structure of the workflow.</p><p><strong>Tier 3: Logged and recoverable.</strong> What must the system record, in enough fidelity, that a reviewer can later reconstruct what the user saw, what the system recommended, and why? This includes the prompt or inferred intent that drove a generated view, the model version, the data snapshot, and any user-initiated overrides.</p><p>A substrate audit applies these three tiers to a specific workflow and produces an answer: here is what we guarantee, here is what we allow to vary, and here is what we log. Vendors who cannot answer those questions are not ready to deploy adaptive interfaces into high-consequence work.</p><p>This shift changes how organizations should think about procurement and governance as well. If future software is assembled around users rather than merely configured by them, then requirements cannot be limited to feature lists and mock screens. Organizations will need to specify the allowable grammar of adaptation. What kinds of user-specific variation are acceptable? Which user traits may the system adapt to, and which should it avoid inferring or acting on? What logging is required when the interface changes? How are training materials, SOPs, and support processes structured when there is no single authoritative screen? What must remain common across all generated variants so that teams can still coordinate safely? NIST&#8217;s framework is especially useful here because it treats AI risk management as continuous, contextual, and organizational, not merely technical.[^10] These are not side questions for a user research team. They are part of the organizational risk architecture.</p><p>There is also a broader professional implication. UX and HSI have often treated the interface as the main artifact and the backend as a supporting concern. That division becomes harder to maintain when the interface is partly generated from underlying models, rules, and orchestration logic. At that point, meaning is no longer guaranteed by the screen alone. It depends on the integrity of the structures beneath the screen. The enduring design object is not just the page, form, or workflow, but the governed system that can produce many acceptable variants without violating meaning, safety, accountability, or coherence. AI does not eliminate UI work. It expands it downward. It forces design and human factors practice to engage more directly with underlying models, constraint frameworks, and the conditions under which variation remains safe and useful.</p><p>The argument, finally, runs this way. AI may not kill the UI, but it may demote the UI from primary artifact to generated expression. The stable thing in the system may no longer be the screen. It may be the substrate from which many screens, workflows, and interaction patterns can be composed. For consumer tools, this may produce highly individualized software experiences assembled around personal needs. For healthcare and other high-consequence domains, it is more likely to produce selective individualization built on tightly governed invariants. Either way, the design center shifts. The future challenge is not simply how to make one interface usable for everyone. It is how to govern variation without losing shared meaning, safe action, and organizational coherence. That is not the end of design work. It is the beginning of a more demanding form of it.</p><h2>For practitioners</h2><p><strong>Ask what is invariant.</strong> For any adaptive or generative interface under consideration, require an explicit answer to: what must always be visible, what may adapt, what must be logged. If the vendor cannot answer, you are not ready.</p><p><strong>Design the fallback first.</strong> The stable, canonical view that users can return to when adaptation fails is the floor of the system. Design it before designing the variants.</p><p><strong>Evaluate across variants, not just instances.</strong> A generated interface that passes heuristic evaluation in one session may fail across the range of views it produces. Evaluation now has a distribution, not a target.</p><p><strong>Move governance upstream.</strong> The design decisions that matter most are no longer about which button goes where. They are about which rules constrain what the system may generate. That is where practitioner expertise is most needed, and where it is currently least represented.</p><div><hr></div><h2>References</h2><p>[^1]: Jakob Nielsen, &#8220;AI: First New UI Paradigm in 60 Years,&#8221; Nielsen Norman Group, June 18, 2023, https://www.nngroup.com/articles/ai-paradigm/.</p><p>[^2]: Kate Moran and Sarah Gibbons, &#8220;Generative UI and Outcome-Oriented Design,&#8221; Nielsen Norman Group, March 22, 2024, https://www.nngroup.com/articles/generative-ui/.</p><p>[^3]: Kate Moran, &#8220;GenUI vs. Vibe Coding: Who&#8217;s Designing?&#8221; Nielsen Norman Group, March 27, 2026, https://www.nngroup.com/articles/genui-vs-vibe/.</p><p>[^4]: Tarun Mugunthan, &#8220;Overcoming the Articulation Barrier in Generative AI Using Hybrid Interfaces,&#8221; Nielsen Norman Group, August 13, 2023, https://www.nngroup.com/articles/ai-articulation-barrier/.</p><p>[^5]: Agency for Healthcare Research and Quality, &#8220;Choosing What Clinicians See in an Electronic Health Record Can Reduce Cognitive Burden and Improve Decision Making,&#8221; AHRQ Digital Healthcare Research, https://digital.ahrq.gov/program-overview/research-stories/choosing-what-clinicians-see-electronic-health-record-can-reduce-cognitive-burden-and-improve.</p><p>[^6]: Agency for Healthcare Research and Quality, &#8220;Improving Electronic Health Record Usability for Patient Safety,&#8221; AHRQ Digital Healthcare Research, https://digital.ahrq.gov/program-overview/research-stories/improving-electronic-health-record-usability-patient-safety.</p><p>[^7]: Office of the National Coordinator for Health Information Technology, &#8220;Usability and Provider Burden,&#8221; HealthIT.gov, https://www.healthit.gov/usability-and-provider-burden/.</p><p>[^8]: Lisanne Bainbridge, &#8220;Ironies of Automation,&#8221; <em>Automatica</em> 19, no. 6 (1983): 775&#8211;779.</p><p>[^9]: Morten Hertzum and Niels Ebbe Jacobsen, &#8220;The Evaluator Effect: A Chilling Fact About Usability Evaluation Methods,&#8221; <em>International Journal of Human-Computer Interaction</em> 13, no. 4 (2001): 421&#8211;443.</p><p>[^10]: National Institute of Standards and Technology, <em>Artificial Intelligence Risk Management Framework (AI RMF 1.0)</em>, NIST AI 100-1 (Gaithersburg, MD: U.S. Department of Commerce, January 2023), https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf.</p><p>[^11]: Office of the National Coordinator for Health Information Technology, <em>Health Data, Technology, and Interoperability: Certification Program Updates, Algorithm Transparency, and Information Sharing (HTI-1) Final Rule</em>, 89 Fed. Reg. 1192 (January 9, 2024), https://www.healthit.gov/topic/laws-regulation-and-policy/health-data-technology-and-interoperability-certification-program.</p><div><hr></div>]]></content:encoded></item><item><title><![CDATA[Lexical Precision as an AI Interaction Skill]]></title><description><![CDATA[Natural language makes AI accessible. Precision is what makes it effective.]]></description><link>https://johnwbrown.substack.com/p/lexical-precision-as-an-ai-interaction</link><guid isPermaLink="false">https://johnwbrown.substack.com/p/lexical-precision-as-an-ai-interaction</guid><dc:creator><![CDATA[John W Brown]]></dc:creator><pubDate>Tue, 28 Apr 2026 06:00:36 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!aLLw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92b8bae0-312a-4b11-8c03-756b58d2897b_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!aLLw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92b8bae0-312a-4b11-8c03-756b58d2897b_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!aLLw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92b8bae0-312a-4b11-8c03-756b58d2897b_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!aLLw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92b8bae0-312a-4b11-8c03-756b58d2897b_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!aLLw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92b8bae0-312a-4b11-8c03-756b58d2897b_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!aLLw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92b8bae0-312a-4b11-8c03-756b58d2897b_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!aLLw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92b8bae0-312a-4b11-8c03-756b58d2897b_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/92b8bae0-312a-4b11-8c03-756b58d2897b_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2243644,&quot;alt&quot;:&quot;Laptop&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/194823840?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92b8bae0-312a-4b11-8c03-756b58d2897b_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Laptop" title="Laptop" srcset="https://substackcdn.com/image/fetch/$s_!aLLw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92b8bae0-312a-4b11-8c03-756b58d2897b_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!aLLw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92b8bae0-312a-4b11-8c03-756b58d2897b_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!aLLw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92b8bae0-312a-4b11-8c03-756b58d2897b_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!aLLw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92b8bae0-312a-4b11-8c03-756b58d2897b_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Image created with generative AI</figcaption></figure></div><div><hr></div><p>Consider a hypothetical scenario. Two analysts at the same organization are asked to produce a short summary of a new internal policy using the same AI drafting tool. They have comparable domain knowledge, comparable training, and comparable access. They sit down with the same source document and the same deadline.</p><p>The first analyst asks the tool to make the summary &#8220;shorter and more professional.&#8221; The second asks for a summary that is &#8220;concise, procedural, and written for frontline staff, with action steps placed before background.&#8221; Both receive fluent output within seconds. Both versions are grammatically correct. Both read as if a competent person wrote them.</p><p>But the first output reads like a press release. The second reads like a standard operating procedure. The first analyst revises three or four times before the draft is usable. The second revises once.</p><p>No reasonable observer would say one analyst is more skilled at AI. The model is the same. The task is the same. The domain knowledge is the same. What differs is vocabulary. The second analyst could name what they wanted at a level of specification the first analyst did not reach. The difference, small in the wording, was decisive in the output.</p><p>This is a hypothetical, but it describes a pattern that is increasingly observable in professional settings where AI tools are being adopted at scale. It points to something that tends to be under-discussed in the current AI literacy conversation: lexical precision.</p><div><hr></div><p>AI systems are often described as easy to use because they accept natural language. That description is broadly true, but it can also be misleading. Natural language lowers the barrier to entry, yet it does not eliminate the need for precision. In many practical settings, especially in knowledge work, analysis, drafting, and summarization, the quality of an AI interaction depends not only on the model&#8217;s capabilities but also on the user&#8217;s ability to express intent clearly and evaluate results accurately.</p><p>Lexical precision refers to the ability to select words that closely match one&#8217;s intended meaning, level of constraint, tone, scope, and purpose. In ordinary conversation, people communicate adequately through approximation. Context, shared assumptions, body language, and follow-up questions help fill in the gaps. In AI interaction, those supports are weaker or absent. The system generates its response from the user&#8217;s language, the system&#8217;s training, and the statistical patterns that govern response generation. Under those conditions, small differences in wording can produce materially different results.</p><div><hr></div><p>This matters because many AI tasks are not merely conversational. They are acts of specification. When a user asks a model to summarize a policy, rewrite a passage, generate a draft, compare options, explain a technical concept, or identify risks, the user is not simply chatting. The user is shaping an output for a purpose. That purpose may involve a particular audience, a required level of formality, a target length, a specific structure, or a preferred style of reasoning. The more precisely those conditions can be expressed, the more likely the interaction is to produce a useful result.</p><p>The common framing of AI as a natural-language interface can obscure this point. It encourages the belief that because the interface is conversational, the interaction itself is informal or low-demand. In practice, many successful AI interactions depend on the user&#8217;s ability to make distinctions that are linguistic in nature.</p><p>A request to make a paragraph shorter is not the same as a request to make it more concise. Shorter compresses by deletion. Concise compresses by density. One operation removes content; the other reorganizes it so the same meaning occupies less space. A system responding to &#8220;shorter&#8221; will typically cut clauses, examples, or qualifiers. A system responding to &#8220;more concise&#8221; will more often rewrite sentences to carry more meaning per word. The outputs diverge in ways that matter when the text is going to a specific audience for a specific purpose.</p><p>The same pattern applies to tone. &#8220;Professional&#8221; is not a specification. It is an approximation that collapses several distinct registers, each signaling a different set of expectations and each likely to produce a different result. The user who can distinguish among those registers has a finer instrument than the user who cannot.</p><div><hr></div><p>The human-factors literature has documented a structurally similar phenomenon for decades, though not in the context of AI. Hertzum and Jacobsen (2003), reviewing eleven studies of three widely used usability evaluation methods, found what they called the evaluator effect: multiple evaluators examining the same interface with the same method detected markedly different sets of usability problems. The effect held across novice and experienced evaluators, across cosmetic and severe problems, and across methods. In one study reviewed in the paper, four evaluators analyzing the same four usability sessions detected ninety-three problems collectively. Only twenty percent of those problems were identified by all four. Forty-six percent were identified by only one.</p><p>The methods were the same. The interfaces were the same. The users being studied were the same. What differed was how each evaluator noticed, framed, and described the problems in front of them. Variability in perception and articulation produced materially different findings from a shared input.</p><p>Lexical precision is the adjacent phenomenon in a new setting. Same system. Same task. Same domain. But the output varies because the expression of intent varies. In AI interaction, the user is not only the evaluator. The user is also the specifier. The act of specifying, like the act of evaluating, depends heavily on the language the user has available for it.</p><p>Parasuraman and Riley (1997) made a related point about automation more broadly. Whether an automated system produced good outcomes depended not only on the system&#8217;s capability but on how the human chose to use, misuse, disuse, or abuse it. The human side of the interaction shaped the outcome at least as much as the machine side did. In AI interaction, the primary instrument of that shaping is language.</p><div><hr></div><p>Five dimensions are worth teaching explicitly, because they are the places where lexical precision most often determines whether an AI interaction produces something useful.</p><p><strong>Scope.</strong> What is in the request and what is deliberately excluded. &#8220;Summarize this policy&#8221; leaves scope open. &#8220;Summarize the enforcement provisions of this policy, without restating the background rationale&#8221; narrows it sharply. The common failure mode is assuming the system will infer relevance from surrounding context. The precise move is to state what should be left out, not only what should be included.</p><p><strong>Constraint.</strong> The hard limits on output: length, format, structure, inclusion or exclusion of specific elements. &#8220;Shorter&#8221; is vague. &#8220;Under two hundred words, no bullet points, no headings, action steps first&#8221; is a constraint. The common failure mode is treating constraint as an afterthought rather than part of the specification. The precise move is to state the constraint before describing the content.</p><p><strong>Tone.</strong> The register of the output. &#8220;Professional&#8221; collapses several distinct options: formal, neutral, executive, clinical, plain-language, warm, instructional. Each produces a different voice. The common failure mode is reaching for tone words that signal competence without specifying direction. The precise move is to choose the tone word that points to the output actually wanted and, where possible, to supply a sample.</p><p><strong>Audience.</strong> Who the output is for. A compliance brief written for legal counsel is a different document than one written for frontline staff. Audience is often the single most underspecified dimension in practitioner requests. The common failure mode is assuming the system will reach for a default audience that matches the user&#8217;s own frame of reference. The precise move is to name the audience and, where useful, the audience&#8217;s likely prior knowledge.</p><p><strong>Uncertainty.</strong> How the system should handle what it does not know. Some tasks call for hedged language and explicit caveats. Others call for a confident recommendation with the hedging done elsewhere. A request that specifies &#8220;flag any claim you are less than highly confident about&#8221; behaves differently from a request that does not mention uncertainty at all. The common failure mode is accepting the system&#8217;s default uncertainty posture without noticing what it is. The precise move is to specify the posture the output should take.</p><p>These five are not exhaustive. They are the dimensions most likely to separate useful output from fluent output, and they are the ones most often collapsed when a user reaches for a single adjective to carry the whole specification.</p><div><hr></div><p>Lexical precision also affects evaluation, which is the second half of most AI interactions. The first output is rarely the final product. The user must judge the response, identify where it falls short, and refine the request. This requires more than a general sense that something feels off. It requires naming the problem.</p><p>An output may be accurate but too abstract. Polished but too generic. Concise but too blunt. Helpful but overly categorical. Well organized but too detached from the actual purpose. The ability to recognize and describe those shortcomings determines how efficiently the user can guide the system toward a better result. A user who can say &#8220;this reads as detached when it should read as instructive&#8221; is working with a sharper instrument than a user who can only say &#8220;try again.&#8221;</p><p>Treating lexical precision as an AI interaction skill helps explain why some users get more from the same systems than others. The difference is not always deeper technical expertise or secret prompting technique. In many cases, it is verbal control. A user who can ask for something to be procedural rather than descriptive, or skeptical rather than balanced, works with a finer instrument than a user whose requests stay broad and underspecified. What looks like superior AI fluency is often a combination of vocabulary, judgment, and iterative refinement.</p><div><hr></div><p>For UX and HSI practitioners, this raises a useful question. When an AI system appears to perform better for some users than for others, what exactly is being measured? Is the difference attributable to the model, the interface, the task, the user&#8217;s domain knowledge, or the user&#8217;s ability to express distinctions in language? In variable-output systems, those factors are often intertwined. If one user consistently gets more useful results because they are better able to specify intent and diagnose response quality, then some portion of the observed performance difference belongs to the human side of the interaction.</p><p>That observation should influence both evaluation and design. In evaluation, teams should be careful not to treat user performance with AI as a simple measure of model quality. A system can appear highly effective in the hands of language-precise users while frustrating others who struggle to formulate and revise requests. Without accounting for that variability, assessments will overestimate usability for the broader population. Testing should consider how users with different levels of verbal confidence and lexical range interact with the same system. It is also useful to observe not only task completion and satisfaction, but the nature of the revisions users make, the specificity of their requests, and the types of wording changes that lead to improved outcomes.</p><p>In design, the implication is not that systems should require users to become rhetoricians. The better response is to reduce unnecessary dependence on vocabulary alone. Interfaces can make distinctions visible so users do not have to generate them unprompted. A writing assistant might let users choose between plain-language, formal, executive summary, technical explanation, and patient-friendly explanation rather than rely solely on free-form prompting. A revision panel might surface dimensions such as brevity, certainty, tone, structure, and audience fit as adjustable controls rather than leaving those dimensions buried in whatever adjective the user happened to supply. The goal is not to remove language from the interaction. It is to make sure the interaction does not silently punish the user who lacks the language.</p><p>This is especially relevant in organizational settings where AI use is becoming normalized among people with varied backgrounds. Some users will arrive with strong writing habits and broad vocabularies. Others will be highly competent in their roles but less practiced in describing stylistic or structural requirements. If success depends too heavily on lexical precision without interface support, the tool may inadvertently privilege one kind of user while appearing neutral on the surface. From an HSI perspective, that is not just a training issue. It is part of the interaction design problem.</p><div><hr></div><p>There is also a developmental dimension worth noting. Repeated AI use tends to strengthen lexical discrimination over time. Users learn by iteration. They try one phrase, inspect the result, adjust the wording, and observe the change. Over many cycles, they start to notice which distinctions matter. A user who once asked only for something to be better starts asking for it to be more concise, more neutral, more audience-specific, more procedural, or more cautious in its claims. AI use contributes to skill development in addition to depending on it. The interaction becomes a feedback loop in which vocabulary is both an input and an outcome.</p><p>That possibility is significant because it suggests that AI adoption is not only changing workflows. It may also be shaping the language habits users bring to work. If so, then training and evaluation should not focus solely on tool operation. They should also consider how users learn to specify intent, assess fit, and refine language over time. A mature AI literacy program might therefore include not only platform guidance and risk awareness, but practical instruction in phrasing, comparison, refinement, and response critique.</p><div><hr></div><p>None of this means vocabulary alone determines success. Domain knowledge still matters. Judgment still matters. Task framing still matters. A precise request aimed at the wrong objective will not produce a good result simply because it is well worded. Even so, lexical precision remains an important and often overlooked part of effective human-AI interaction. It helps users express what they want, recognize when the system has drifted from that intent, and guide the interaction toward a more useful outcome.</p><p>The interface is bilingual. Natural language gets the user in the door. Precision is what the system responds to once they are inside. As AI systems become more common in professional environments, the language used to interact with them should not be treated as incidental. It is part of the interface. Natural language may make these systems accessible, but precision is what often makes them effective. For that reason, lexical precision deserves more attention as a practical interaction skill, with direct implications for usability, AI literacy, and system design.</p><div><hr></div><p><strong>References</strong></p><p>Hertzum, M., &amp; Jacobsen, N. E. (2003). The evaluator effect: A chilling fact about usability evaluation methods. <em>International Journal of Human-Computer Interaction</em>, 15(1), 183&#8211;204.</p><p>Parasuraman, R., &amp; Riley, V. (1997). Humans and automation: Use, misuse, disuse, abuse. <em>Human Factors</em>, 39(2), 230&#8211;253.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/p/lexical-precision-as-an-ai-interaction?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/p/lexical-precision-as-an-ai-interaction?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[The Most Important Usability Signal No One Writes Down]]></title><description><![CDATA[But they should]]></description><link>https://johnwbrown.substack.com/p/the-most-important-usability-signal</link><guid isPermaLink="false">https://johnwbrown.substack.com/p/the-most-important-usability-signal</guid><dc:creator><![CDATA[John W Brown]]></dc:creator><pubDate>Tue, 21 Apr 2026 06:02:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!HxDH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42fc7fcc-ec6a-4c4b-ba1f-319c3d7225cf_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HxDH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42fc7fcc-ec6a-4c4b-ba1f-319c3d7225cf_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HxDH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42fc7fcc-ec6a-4c4b-ba1f-319c3d7225cf_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!HxDH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42fc7fcc-ec6a-4c4b-ba1f-319c3d7225cf_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!HxDH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42fc7fcc-ec6a-4c4b-ba1f-319c3d7225cf_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!HxDH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42fc7fcc-ec6a-4c4b-ba1f-319c3d7225cf_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HxDH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42fc7fcc-ec6a-4c4b-ba1f-319c3d7225cf_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/42fc7fcc-ec6a-4c4b-ba1f-319c3d7225cf_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1950792,&quot;alt&quot;:&quot;The Most Important Usability Signal No One Writes Down&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/187394265?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42fc7fcc-ec6a-4c4b-ba1f-319c3d7225cf_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The Most Important Usability Signal No One Writes Down" title="The Most Important Usability Signal No One Writes Down" srcset="https://substackcdn.com/image/fetch/$s_!HxDH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42fc7fcc-ec6a-4c4b-ba1f-319c3d7225cf_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!HxDH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42fc7fcc-ec6a-4c4b-ba1f-319c3d7225cf_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!HxDH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42fc7fcc-ec6a-4c4b-ba1f-319c3d7225cf_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!HxDH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42fc7fcc-ec6a-4c4b-ba1f-319c3d7225cf_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Image created with generative AI</figcaption></figure></div><div><hr></div><p><em>The views expressed are my own and do not represent any organization I am affiliated with.</em></p><div><hr></div><p>Consider a pharmacy technician reviewing a medication order flagged by the system as complete. The order looks correct. The dosage matches the prescription. The patient information is accurate. Nothing on the screen suggests a problem.</p><p>She checks it again anyway.</p><p>She scrolls back to the original prescription, compares the dosage field a second time, tabs forward to the confirmation screen, then returns to confirm the patient&#8217;s date of birth. The entire sequence takes about twenty seconds. She submits the order. Task complete, no errors, no flags.</p><p>From a metrics standpoint, nothing looks wrong.</p><p>But something happened in those twenty seconds that most usability reports will never capture. The technician was not confused. She was not lost. She was filling a gap the system left open: it had given her no reason to doubt the output and no reason to trust it either.</p><p>That behavior has a name, though it rarely appears in formal findings.</p><p>It is called self-checking.</p><h2>What Self-Checking Looks Like</h2><p>Self-checking occurs when users pause to verify that they have done the right thing, even when the system has provided no indication of error. They reread instructions, review entered data, scroll back to confirm earlier steps, or mentally replay what they just did.</p><p>Unlike confusion, self-checking does not interrupt task flow. Unlike errors, it does not require correction. Users often complete the task successfully while engaging in repeated verification behaviors. The task succeeds. The practitioner moves on. The evaluator notes nothing remarkable.</p><p>This is precisely the problem.</p><p>Self-checking is quiet. It hides inside competent performance. It looks like thoroughness rather than friction, and because usability practice is oriented toward capturing breakdowns, behaviors that occur within successful task completion tend to fall below the reporting threshold.</p><h2>Why It Rarely Gets Written Down</h2><p>Self-checking is easy to dismiss because it feels reasonable. In many professional domains, careful practitioners are praised for reviewing their work. Evaluators may assume the behavior reflects domain complexity, professional diligence, or personal style rather than a usability concern.</p><p>It is also difficult to quantify. Self-checking does not always extend task time significantly. It may appear as brief pauses or small backward movements that are easy to overlook unless specifically watched for.</p><p>As a result, self-checking often remains an unspoken observation. Evaluators see it, recognize it, perhaps mention it to a colleague in the hallway afterward. But it seldom appears in the report with the same rigor as a missed button or a misunderstood label. It lacks the clean narrative that findings require: user attempted X, failed, because Y.</p><p>Self-checking does not fail. That is what makes it invisible.</p><h2>What Self-Checking Actually Signals</h2><p>Research on trust calibration in automated systems offers useful framing here. Lee and See&#8217;s foundational work on trust in automation established that appropriate reliance depends on users having access to the system&#8217;s performance characteristics, its process, and its purpose.[^1] When any of these dimensions is hidden, users adjust. They do not stop using the system. They monitor it, layering their own review on top of the system&#8217;s output.</p><p>Self-checking is rarely about user uncertainty alone. It is a response to systems that do not sufficiently support trust. Practitioners understand what they did, but they are unsure whether the system interpreted their input correctly or whether the output reflects their intent.</p><p>This behavior is especially common in systems that summarize, infer, or automate decisions. When practitioners cannot see how an outcome was produced, they fill the gap by reviewing everything they can see.</p><p>Parasuraman and Riley&#8217;s research on automation misuse and disuse describes a related pattern: operators who cannot calibrate their reliance on a system default to either over-trust or persistent monitoring, depending on the stakes involved.[^2] Self-checking is the behavioral signature of the monitoring response.</p><p>In this way, self-checking signals that the system has shifted cognitive burden back onto the user, not through poor design in the conventional sense, but through insufficient support for trust.</p><h2>The Cost of Ignoring It</h2><p>Self-checking imposes real cognitive cost. It consumes attention, slows work, and increases mental fatigue over time. In high-frequency tasks, even small review behaviors accumulate into significant inefficiency. A twenty-second check performed forty times per shift is more than thirteen minutes of unacknowledged cognitive labor per day.</p><p>More importantly, persistent self-checking erodes trust. Practitioners may rely on the system only provisionally, treating it as something to be monitored rather than something to depend on. Over time, this can lead to avoidance of features, parallel workflows, or reliance on informal checks outside the system. Clinicians printing records they could view on screen, analysts rebuilding calculations in a personal spreadsheet, claims processors keeping handwritten notes alongside a digital queue: these are all downstream expressions of assurance failure that self-checking predicted.</p><p>None of these outcomes appear in standard usability metrics. Task completion rates remain high. Error rates remain low. Satisfaction scores may register mild dissatisfaction, but not the kind that triggers redesign.</p><p>The system works. The users cope. And the gap between working and trusted widens without documentation.</p><h2>A Framework for Observing Self-Checking</h2><p>Capturing self-checking requires more than noting that it happened. Evaluators need a structured way to classify what they observe, assess its severity, and connect it to design response. The following framework organizes self-checking behaviors into four categories based on what the practitioner is reviewing and why.</p><p><strong>1. Input Echo Checking.</strong> The practitioner re-examines data they entered to confirm it was captured correctly. This includes rereading form fields after entry, toggling between input and confirmation screens, or scrolling back to previously completed steps. Input echo checking signals that the system does not adequately confirm what it received. The fix is typically straightforward: persistent input summaries, inline confirmation, or edit-in-place visibility.</p><p><strong>2. Output Verification.</strong> The practitioner examines system-generated results to determine whether the output reflects their intent. This is common in systems that calculate, summarize, or recommend. The practitioner knows what they asked for but cannot determine whether the system interpreted the request correctly. Output verification signals that the system&#8217;s reasoning is hidden. Fixes involve progressive disclosure of logic, trust indicators, or traceable input-to-output mappings.</p><p><strong>3. Commit-Point Hesitation.</strong> The practitioner pauses at the moment of submission, approval, or irreversible action. They may reread a summary screen, hover over a submit button, or navigate backward one final time before proceeding. Commit-point hesitation signals that the perceived cost of error exceeds the assurance the system has provided. Fixes include clear undo pathways, pre-commit summaries that highlight consequential fields, and explicit confirmation of what will happen next.</p><p><strong>4. Ambient Re-confirmation.</strong> The practitioner periodically checks elements of the interface that have not changed, such as a patient banner, a file name, or a project identifier. This behavior signals concern about context rather than content: the practitioner is confirming they are still working on the right record in the right place. Ambient re-confirmation is especially common in systems that support multiple concurrent records or contexts. Fixes include persistent contextual anchors, differentiated visual environments per context, and clear state indicators.</p><p>Not all self-checking is equal. Input echo checking is the mildest form and often the easiest to resolve. Ambient re-confirmation, particularly in safety-critical systems, can indicate serious trust failures that warrant design priority.</p><h2>Self-Checking in Practice</h2><p><strong>The Order That Was Already Cleared.</strong> Suppose a hospital pharmacy implements an automated order review system that cross-references prescriptions against patient records, formulary data, and interaction databases. The system flags potential conflicts and clears orders that meet safety thresholds. Post-deployment observation reveals that pharmacists routinely re-examine cleared orders manually, spending an average of fifteen additional seconds per order reviewing information the system has already checked. Standard metrics show no increase in error rates and only modest increases in processing time. A self-checking-aware evaluation would classify this as output verification (Category 2) and investigate whether the system provides sufficient transparency about what it checked and why the order was cleared. The pharmacists are not doubting the system&#8217;s accuracy. They are covering for its opacity.</p><p><strong>The Approver Who Clicked Through.</strong> Consider an expense management platform that auto-categorizes submitted expenses using transaction metadata. Approving managers are shown a summary with category assignments already applied. Observation reveals that managers frequently click into individual line items to review categories, even when the summary view provides all necessary information. Task completion is unaffected, but managers report in follow-up interviews that they &#8220;just want to make sure it got it right.&#8221; This is a blend of output verification and commit-point hesitation (Categories 2 and 3). The system&#8217;s auto-categorization removed a step the managers previously performed manually, but it did not replace the certainty that manual categorization provided. Showing the basis for each categorization decision, even briefly, would address the underlying signal.</p><p><strong>The Clinician Who Watched the Banner.</strong> Imagine a clinician working in an electronic health record system that supports rapid switching between patient charts. Observation during usability evaluation reveals that the clinician checks the patient name banner after nearly every navigation action, even when the system has not changed contexts. The pattern is most pronounced when the clinician returns from a brief interruption or switches between tasks within the same chart. This is ambient re-confirmation (Category 4) and signals that the system&#8217;s contextual anchoring is insufficient for the cognitive demands of multi-patient workflows. Differentiated color schemes per patient, persistent identity elements, and interruption-recovery cues would reduce this form of self-checking.</p><h2>How to Observe Self-Checking Intentionally</h2><p>Capturing self-checking requires deliberate attention during evaluation sessions. It will not surface in task-completion data or standard think-aloud protocols unless evaluators are specifically watching for it.</p><p>Watch for repeated scanning of previously viewed information, backward navigation after apparent task success, pauses that occur at decision points rather than confusion points, and any moment where the practitioner appears to be confirming rather than progressing. These actions should be noted explicitly, timestamped, and classified using the framework above.</p><p>Follow-up questions during or after sessions can clarify intent. Useful prompts include: &#8220;What were you checking there?&#8221; &#8220;Did anything feel uncertain at that point?&#8221; &#8220;Would you normally double-check this step, or was something about the system prompting that?&#8221; Practitioners often surface concerns only when asked directly, and their explanations frequently reveal trust gaps that observation alone cannot fully diagnose.</p><p>Treating self-checking as data, rather than incidental behavior, changes how findings are framed. A report that notes &#8220;users successfully completed all tasks with no errors&#8221; reads very differently from one that adds &#8220;however, practitioners engaged in repeated review behaviors at three critical decision points, suggesting that task success required additional cognitive effort the system did not support.&#8221;</p><h2>Designing to Reduce Unnecessary Self-Checking</h2><p>Reducing self-checking does not mean encouraging blind trust. It means designing systems that help practitioners understand when review is necessary and when it is not. Several design patterns address this directly.</p><p><strong>Input echo and persistent summaries.</strong> When practitioners enter data that will drive downstream decisions, the system should confirm what it received in the practitioner&#8217;s terms, not just the system&#8217;s internal representation. A persistent summary that reflects the practitioner&#8217;s input, visible at decision points, reduces input echo checking.</p><p><strong>Transparent reasoning.</strong> When systems automate decisions, even partial transparency about the basis for those decisions reduces output verification. This does not require full algorithmic explanation. A brief statement of the factors considered, such as &#8220;Cleared: no interactions found with current medications (checked 4 active prescriptions),&#8221; is often sufficient.</p><p><strong>Proportional commit-point support.</strong> The level of confirmation support at submission points should match the consequences of the action. Low-stakes actions need minimal confirmation. High-stakes or irreversible actions benefit from pre-commit summaries that highlight consequential fields, clear statements of what will happen next, and accessible undo or correction pathways.</p><p><strong>Contextual anchoring.</strong> In systems that support multiple concurrent records, persistent and visually differentiated contextual indicators reduce ambient re-confirmation. Color-coding, prominent identity banners, and context-change alerts help practitioners maintain orientation without manual checking.</p><p>The goal is to align practitioner effort with actual risk, not to eliminate review entirely, but to ensure that the system supports trust where it can so that practitioners reserve their checking for moments that genuinely warrant it.</p><h2>Boundary Conditions: When Self-Checking Is Not a Usability Signal</h2><p>Not all self-checking indicates a design problem. In certain contexts, review behavior is appropriate and even expected.</p><p>Regulatory environments may require documented human verification regardless of system assurance. Training scenarios often involve deliberate checking as a learning behavior. Genuinely novel or ambiguous situations may warrant review even in well-designed systems. And individual differences in baseline review tendency vary across practitioners and should be accounted for in evaluation design rather than attributed entirely to the interface.</p><p>The evaluator&#8217;s task is to distinguish between self-checking that arises from the system&#8217;s failure to support trust and self-checking that arises from the domain&#8217;s legitimate demands. The framework above helps make that distinction by linking observed behavior to specific system characteristics rather than treating all checking as equivalent.</p><h2>Practical Takeaway</h2><p>The most important usability signals are not always failures or errors. They are often subtle patterns that indicate where practitioners are doing extra cognitive work to cover for system opacity.</p><p>Self-checking deserves to be written down. It reveals where systems technically work but fail to support trust. Ignoring it leads to designs that appear usable while quietly taxing practitioners in ways that metrics never surface and reports never capture.</p><p>Usability work that documents self-checking moves closer to the real experience of work: not just what practitioners did, but what they felt they had to review before trusting it was done.</p><div><hr></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cT80!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84828b38-bcd9-4a38-949a-727b0e58df3a_619x506.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cT80!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84828b38-bcd9-4a38-949a-727b0e58df3a_619x506.png 424w, https://substackcdn.com/image/fetch/$s_!cT80!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84828b38-bcd9-4a38-949a-727b0e58df3a_619x506.png 848w, https://substackcdn.com/image/fetch/$s_!cT80!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84828b38-bcd9-4a38-949a-727b0e58df3a_619x506.png 1272w, https://substackcdn.com/image/fetch/$s_!cT80!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84828b38-bcd9-4a38-949a-727b0e58df3a_619x506.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cT80!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84828b38-bcd9-4a38-949a-727b0e58df3a_619x506.png" width="619" height="506" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/84828b38-bcd9-4a38-949a-727b0e58df3a_619x506.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:506,&quot;width&quot;:619,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:73382,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/187394265?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84828b38-bcd9-4a38-949a-727b0e58df3a_619x506.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!cT80!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84828b38-bcd9-4a38-949a-727b0e58df3a_619x506.png 424w, https://substackcdn.com/image/fetch/$s_!cT80!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84828b38-bcd9-4a38-949a-727b0e58df3a_619x506.png 848w, https://substackcdn.com/image/fetch/$s_!cT80!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84828b38-bcd9-4a38-949a-727b0e58df3a_619x506.png 1272w, https://substackcdn.com/image/fetch/$s_!cT80!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84828b38-bcd9-4a38-949a-727b0e58df3a_619x506.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><p>[^1]: John D. Lee and Katrina A. See, &#8220;Trust in Automation: Designing for Appropriate Reliance,&#8221; <em>Human Factors</em> 46, no. 1 (2004): 50&#8211;80.</p><p>[^2]: Raja Parasuraman and Victor Riley, &#8220;Humans and Automation: Use, Misuse, Disuse, Abuse,&#8221; <em>Human Factors</em> 39, no. 2 (1997): 230&#8211;253.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/p/the-most-important-usability-signal?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/p/the-most-important-usability-signal?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p>]]></content:encoded></item><item><title><![CDATA[Error Messages That Don’t Trigger Errors]]></title><description><![CDATA[Why the most dangerous usability failures in AI-integrated tools are the ones that never announce themselves]]></description><link>https://johnwbrown.substack.com/p/error-messages-that-dont-trigger</link><guid isPermaLink="false">https://johnwbrown.substack.com/p/error-messages-that-dont-trigger</guid><dc:creator><![CDATA[John W Brown]]></dc:creator><pubDate>Tue, 14 Apr 2026 06:01:58 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!CSI7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7262fad1-51d0-4ed3-8e0e-350efad1b303_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CSI7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7262fad1-51d0-4ed3-8e0e-350efad1b303_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CSI7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7262fad1-51d0-4ed3-8e0e-350efad1b303_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!CSI7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7262fad1-51d0-4ed3-8e0e-350efad1b303_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!CSI7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7262fad1-51d0-4ed3-8e0e-350efad1b303_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!CSI7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7262fad1-51d0-4ed3-8e0e-350efad1b303_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CSI7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7262fad1-51d0-4ed3-8e0e-350efad1b303_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7262fad1-51d0-4ed3-8e0e-350efad1b303_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2050451,&quot;alt&quot;:&quot;Error Messages That Don&#8217;t Trigger Errors&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/187393820?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7262fad1-51d0-4ed3-8e0e-350efad1b303_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Error Messages That Don&#8217;t Trigger Errors" title="Error Messages That Don&#8217;t Trigger Errors" srcset="https://substackcdn.com/image/fetch/$s_!CSI7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7262fad1-51d0-4ed3-8e0e-350efad1b303_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!CSI7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7262fad1-51d0-4ed3-8e0e-350efad1b303_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!CSI7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7262fad1-51d0-4ed3-8e0e-350efad1b303_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!CSI7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7262fad1-51d0-4ed3-8e0e-350efad1b303_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Image created with generative AI</figcaption></figure></div><div><hr></div><p><em>The views expressed are my own and do not represent any organization I am affiliated with.</em></p><h2>When the System Stays Quiet</h2><p>A clinician finishes a fifteen-minute encounter and opens the AI-assisted documentation tool. The draft note is already waiting. It is well-formatted, grammatically clean, and structured to the template her department prefers. The chief complaint is correct. The history of present illness reads naturally. The assessment and plan look reasonable.</p><p>She makes a few edits, signs the note, and moves to the next patient.</p><p>What the note does not say is that the patient mentioned, in passing, that she had stopped taking her anticoagulant three days earlier. The model captured the medication in the active list. It did not capture the discontinuation. The omission is invisible in the finished document because nothing in the document points to it. There is no flag. No warning. No asterisk. No error message.</p><p>The next clinician who reads that note will assume the medication list is current. The decision support rules that fire from that list will assume the same. The patient is now in a record that believes something untrue, and the belief was introduced by a tool that performed exactly as designed.</p><p>This is the failure mode I want to discuss. Not the kind that interrupts. The kind that does not.</p><h2>The Classical Model of Error Handling</h2><p>Most usability guidance around errors assumes a discrete breakdown. An invalid input is entered. A constraint is violated. A process fails. The interface responds with an error message, ideally one that is clear, actionable, and polite.</p><p>This model is the inheritance of decades of careful work. Don Norman&#8217;s distinction between slips and mistakes, James Reason&#8217;s taxonomy of human error, the entire tradition of forcing functions and constraint-based design, all of it assumes that errors are events. Something measurable happens. The tool can detect it. The user can be notified. The evaluator can test for it.</p><p>In deterministic systems this works. Errors are rule violations. The software knows when something has gone wrong and can respond immediately. Usability evaluators can probe error prevention and recovery by deliberately triggering known failure conditions and observing what the interface does.</p><p>In this model, no error message generally implies no error.</p><p>That assumption is now actively dangerous.</p><h2>Where the Model Breaks</h2><p>Adaptive and inferential tools do not produce errors the way deterministic systems do. They produce outputs. The outputs are interpretations, not retrievals. They are based on probabilistic reasoning over inputs that may be incomplete, ambiguous, or themselves uncertain. The model has no internal definition of correctness against which to compare its own answer.</p><p>From the platform&#8217;s perspective, nothing has failed. The model ran. The pipeline completed. The output was generated. There is no rule violation to detect because there was no rule to violate. The notion of &#8220;error&#8221; as a discrete event does not map onto a process whose entire job is to make a guess.</p><p>From the user&#8217;s perspective, the tool has quietly led them somewhere wrong.</p><p>This gap is widening, and it is widening fastest in the domains where the consequences are heaviest. Clinical decision support. Legal review. Financial risk modeling. Code generation. Intelligence analysis. The places we are most eager to deploy inferential tools are the places where quiet failure costs the most.</p><h2>Plausible Wrongness</h2><p>The most dangerous failure mode in inferential tools is what I will call plausible wrongness. The model returns an answer that looks correct enough to accept without scrutiny. The formatting is right. The tone is confident. The structure matches expectations. Nothing about the surface of the output invites doubt.</p><p>The user proceeds, trusting the result, and only later discovers that something important was missed, misclassified, or misinterpreted. By then the answer has been incorporated into a document, a decision, a downstream process, or a chain of inference that other people are now relying on.</p><p>Plausible wrongness is not a single failure pattern. It comes in at least five recognizable forms.</p><p><strong>Confident misclassification.</strong> The model assigns an input to the wrong category and presents the assignment with no indication of difficulty. A risk score is produced. A diagnosis is suggested. A document is tagged. The user sees the result, not the closeness of the call.</p><p><strong>Fluent fabrication.</strong> The model generates content that has no grounding in source material but reads as if it does. Citations that do not exist. Quotes that were never said. Statistics that look reasonable but were never measured. The fluency of the output is precisely what makes it hard to question.</p><p><strong>Scope drift.</strong> The model answers a question adjacent to the one that was asked. The user requested an analysis of Q3 returns and received an analysis of Q3 revenue. The response is internally coherent. The mismatch only becomes visible if the user reads carefully enough to notice that the output does not actually address the question.</p><p><strong>Stale-context output.</strong> The tool relies on cached, outdated, or incomplete context and produces a result that would have been correct earlier. A medication list that is current as of last week. A policy reference that has since been superseded. A patient record that omits the most recent encounter.</p><p><strong>False completeness.</strong> The tool returns what it found and implies, by the form of its response, that what it found is all there is. A search result presented without acknowledgment of what was excluded. A summary that omits the contradictory passage. A recommendation that does not mention the alternatives it considered and rejected.</p><p>What unites these patterns is the absence of a signal. The tool does not say &#8220;I am uncertain.&#8221; It does not say &#8220;I may have missed something.&#8221; It states the answer, in the same voice it uses for answers it has high confidence in, and the user has no way of knowing the difference.</p><h2>Error Signaling Versus Uncertainty Signaling</h2><p>The classical error message is a binary. Something is wrong, or nothing is. This binary made sense when software could actually tell. It does not make sense for tools whose outputs are guesses by design.</p><p>The more honest signal in inferential tools is uncertainty. Rather than presenting every result with the same confident affect, a well-designed interface can indicate when outputs are based on incomplete information, ambiguous inputs, or low-confidence inference. This is fundamentally different from error messaging. It is not a report of failure. It is a report of epistemic state.</p><p>Uncertainty signaling, done well, helps users calibrate trust. It gives them a reason to pause, verify, or seek a second source, without implying that they did anything wrong. It treats the user as a reasoning partner rather than as a passive recipient of pronouncements.</p><p>Uncertainty signaling, done badly, becomes noise. Confidence scores attached to every output get ignored within a week. Color-coded warnings that fire indiscriminately produce the alarm fatigue that has plagued clinical decision support for two decades. The challenge is not whether to signal uncertainty but when, and how legibly, and at what threshold.</p><p>The best implementations I have seen share a few traits. They are sparing. They distinguish meaningfully between confidence levels rather than spraying gradients. They tell the user what specifically is uncertain, not just that something is. They give the user something to do with the information, even if that something is only &#8220;look more carefully at this one field.&#8221;</p><p>The worst implementations treat uncertainty signaling as a liability waiver. Caveats are appended to every result so that no individual caveat carries weight. The signal becomes a legal artifact rather than a usability feature, and users learn to scroll past it.</p><p>A tool that never acknowledges uncertainty is making an implicit claim it cannot back up. A tool that acknowledges it constantly is making no claim at all. The interesting design space is in between, and most current AI-integrated platforms have not yet found it.</p><h2>Three Quiet Failures</h2><p>A diagnostic decision-support tool was evaluated for a specialty practice. The tool produced differential diagnoses ranked by likelihood, drawing on the patient&#8217;s chart and reported symptoms. In testing, it performed well on common presentations. What the testing did not surface was that on edge cases, the tool would silently exclude diagnoses for which the chart lacked sufficient structured data, even when those diagnoses were clinically obvious from the unstructured notes. The exclusion was not flagged. The ranked list looked complete. Clinicians who trusted the list as a working differential were, in a small but meaningful percentage of cases, working from a list with the right answer missing. No error was ever logged. The tool was, by its own measures, functioning correctly.</p><p>An LLM-assisted code review tool was deployed on a development team to flag security concerns in pull requests. It was fast, articulate, and produced summaries that read like the work of an experienced reviewer. Over the first quarter it caught real issues and built credibility. What the team noticed only after a near-miss in production was that the tool consistently failed to flag a particular class of injection vulnerability when the unsafe input was passed through more than two function calls before reaching the sink. The summary for those pull requests said the code looked good. It did not say the tool had not traced the data flow that far. The team had been treating the absence of flags as evidence of safety.</p><p>A clinical summarization tool was used to condense long inpatient stays into discharge-ready narratives. The summaries were well-organized and readable. Over time, audit revealed that the tool tended to drop comorbidities mentioned only once in the chart and never repeated. For most patients this was harmless. For patients whose single mention was the most clinically significant fact in the entire record, the omission propagated into the discharge summary, the primary care handoff, and the patient&#8217;s own understanding of their condition. The tool never indicated that anything had been left out. Its measure of success was the readability of the summary, and by that measure it succeeded.</p><p>In each case the tool performed as designed. In each case the design did not include any mechanism for telling the user what the tool might be wrong about. In each case the user only discovered the failure through a path that bypassed the interface entirely: a near-miss, an audit, a downstream reviewer who happened to know what to look for.</p><h2>What to Look For During Evaluation</h2><p>Quiet failures evade most standard usability testing because the user completes the task. The session ends. Notes record success. The risk only surfaces in real use, and by then the evaluator has moved on.</p><p>Catching plausible wrongness requires probing for it deliberately. The following are the heuristics I have found most useful when evaluating inferential tools for silent failure.</p><p><strong>1. Probe with known-incomplete inputs.</strong> Deliberately give the tool inputs that are missing information it has no way to know is missing. Watch what it produces. A well-designed tool will indicate that its answer is conditional on what it was given. A poorly designed one will produce the same confident output it would have produced with complete information.</p><p><strong>2. Test for fabrication on plausible-but-fictional prompts.</strong> Ask about entities, cases, or references that do not exist but that sound like they might. A model that fabricates fluently in testing is a model that will fabricate in production.</p><p><strong>3. Compare outputs across paraphrased inputs.</strong> Ask the same question several ways. If the responses vary substantially, the tool is making guesses it is presenting as facts. The variance itself is the signal.</p><p><strong>4. Watch for mismatches between question and answer.</strong> Track whether the output actually addresses what was asked, or whether it addresses something adjacent. Scope drift is one of the easiest silent failures to miss because the response looks responsive.</p><p><strong>5. Check what the tool omits.</strong> Ask not only whether the output is correct but whether it is complete. What did the tool have access to that it did not include? What did it consider and reject? Can the user tell?</p><p><strong>6. Observe user reactions, not just task outcomes.</strong> Hesitation, second-guessing, expressions of mild surprise, reflexive double-checking against another source. These are often the only behavioral evidence that something silent went wrong.</p><p><strong>7. Audit downstream.</strong> Where possible, follow the outputs into the next step of the workflow and the step after that. Silent failures often become visible only when their consequences propagate, and the propagation path is where the real evaluation happens.</p><p>The heuristics share a common premise. Task completion is not the same as task success. In inferential tools, the two have come apart, and evaluation methods need to come apart with them.</p><h2>Implementation Guidance</h2><p>For evaluators and design leaders integrating inferential tools into existing practice, a few practical steps are worth building into the standard protocol.</p><p>Add silent-failure probing to acceptance testing. Before a tool is approved for production, it should be tested not only on its happy path but on its quiet-wrong path. Write test cases whose correct answer is &#8220;I am not sure&#8221; or &#8220;this question cannot be answered with the information available.&#8221; Check whether the tool produces those answers, or whether it produces a confident wrong one.</p><p>Require uncertainty signaling in vendor specifications. When evaluating tools for procurement, ask vendors directly how their product communicates uncertainty to end users. Ask for examples. A vendor who cannot answer this question has not designed for it, and the absence is itself information.</p><p>Pair AI outputs with verification surfaces. Where a tool produces a recommendation, design the interface so that the user has a clear, low-friction way to see what the recommendation is based on. Source highlighting. Inline citations. Visible confidence indicators when meaningful. The goal is not to force verification on every interaction but to make verification possible when the user wants it.</p><p>Train users on the failure modes. Practitioners who understand what plausible wrongness looks like are more likely to catch it. The training does not need to be technical. It needs to be specific. Show real examples of fluent fabrication and confident misclassification from the actual tools the team will be using.</p><p>Build feedback loops. Quiet failures get caught downstream or not at all. The practitioners closest to the consequences need an easy way to report what they found, and the reports need to feed back into evaluation rather than disappearing into a ticket queue.</p><h2>Risks and Limitations</h2><p>The case for uncertainty signaling has limits worth naming.</p><p>Not every domain tolerates probabilistic outputs. Some tasks require a single confident answer because the downstream process cannot proceed without one. In those domains, the right design move may be to constrain the tool&#8217;s scope rather than to expose its uncertainty.</p><p>Uncertainty signaling can also be overused. A tool that hedges every output trains its users to ignore the hedges, which is worse than not hedging at all. Calibration matters more than coverage.</p><p>Some failure modes are not detectable by the tool itself. A model that does not know what it does not know cannot signal the gap. In those cases, the corrective has to come from outside: from process design, from human review, from audit. No interface convention will rescue a tool whose blind spots are invisible to itself.</p><p>And there is the alarm fatigue history to respect. Clinical decision support has spent two decades learning, painfully, that warning the user about everything is the same as warning them about nothing. Whatever uncertainty signaling looks like in the next generation of inferential tools, it will have to learn that lesson without having to relive it.</p><h2>Conclusion</h2><p>The most dangerous usability failures are not the ones that stop users. They are the ones that let users proceed confidently in the wrong direction.</p><p>Classical error handling was built for systems that knew when they were wrong. Inferential tools do not have that knowledge. Their outputs are guesses dressed in the visual grammar of facts, and the visual grammar is what makes them hard to challenge. Evaluating these tools means evaluating not only what they produce but what they fail to communicate about how they produced it.</p><p>Error messages that never trigger are not a sign of success. They are often a sign that the system does not know when it is wrong. Usability work that focuses only on visible errors will consistently miss the failures that matter most. The next generation of evaluation practice has to make the silent ones audible.</p><div><hr></div><h2>Quick Reference: Silent Failure Evaluation Checklist</h2><ul><li><p>Does the tool distinguish between high-confidence and low-confidence outputs?</p></li><li><p>Does it tell the user what it does not know, or only what it knows?</p></li><li><p>Does it indicate when its answer is based on incomplete or ambiguous input?</p></li><li><p>Can the user see what was considered and excluded?</p></li><li><p>Does the same question, asked different ways, produce consistent results?</p></li><li><p>Does the tool fabricate fluently when prompted with plausible-but-fictional inputs?</p></li><li><p>Are there test cases in the acceptance protocol whose correct answer is &#8220;I do not know&#8221;?</p></li><li><p>Is there a feedback path for downstream reviewers who catch silent failures in production?</p></li><li><p>Does the vendor have a clear answer for how the tool communicates uncertainty?</p></li><li><p>Are users trained on the specific failure modes of the specific tools they use?</p></li></ul><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/p/error-messages-that-dont-trigger?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/p/error-messages-that-dont-trigger?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p>]]></content:encoded></item><item><title><![CDATA[The Productive Detour in Usability Testing]]></title><description><![CDATA[When Going Off Script Produces Better Findings]]></description><link>https://johnwbrown.substack.com/p/the-productive-detour-in-usability</link><guid isPermaLink="false">https://johnwbrown.substack.com/p/the-productive-detour-in-usability</guid><dc:creator><![CDATA[John W Brown]]></dc:creator><pubDate>Tue, 07 Apr 2026 06:00:54 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!WeBW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff362554f-ebdb-46ea-b276-bcd14731baca_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WeBW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff362554f-ebdb-46ea-b276-bcd14731baca_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WeBW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff362554f-ebdb-46ea-b276-bcd14731baca_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!WeBW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff362554f-ebdb-46ea-b276-bcd14731baca_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!WeBW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff362554f-ebdb-46ea-b276-bcd14731baca_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!WeBW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff362554f-ebdb-46ea-b276-bcd14731baca_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WeBW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff362554f-ebdb-46ea-b276-bcd14731baca_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f362554f-ebdb-46ea-b276-bcd14731baca_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1854130,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/190650178?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff362554f-ebdb-46ea-b276-bcd14731baca_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!WeBW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff362554f-ebdb-46ea-b276-bcd14731baca_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!WeBW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff362554f-ebdb-46ea-b276-bcd14731baca_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!WeBW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff362554f-ebdb-46ea-b276-bcd14731baca_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!WeBW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff362554f-ebdb-46ea-b276-bcd14731baca_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Image created with generative AI</figcaption></figure></div><div><hr></div><p><em>The views expressed are my own and do not represent any organization I am affiliated with.</em></p><div><hr></div><p>Consider a moderator facilitating a usability study on a clinical documentation workflow. The assigned scenario is straightforward: locate a specific patient record and update a data field. The participant finds the correct screen, pauses, and then opens a different module entirely. &#8220;I would never actually do this here,&#8221; she says. &#8220;I would go to this other system first, check the note from last week, and then come back.&#8221; She is four minutes into a twenty-minute session.</p><p>The moderator has about three seconds to decide: redirect her to the protocol, or follow.</p><p>Most training would say redirect. Keep the study on track. Protect the standardized data. That instinct is sound, and in many situations it is the right call.</p><p>But in this case, the researcher has just learned something more valuable than completion data. The official scenario assumes the user works from within a single platform. She does not. Her workflow depends on an external source, an implicit verification step, and a mental model that the design team never accounted for. The formal workflow and the lived workflow are not the same thing.</p><p>Redirect too quickly, and that finding disappears.</p><div><hr></div><h2>The confusion between deviation and contamination</h2><p>A strict view of usability testing treats any departure from the planned path as a threat to validity. Under that model, the cleanest session is the best session. Standardization protects comparability. Consistency protects rigor. The more testers follow the same route, the more confidently findings can be aggregated and reported.</p><p>That logic is not wrong. It is incomplete.</p><p>The same standardization that makes cross-participant comparison reliable can suppress the findings that matter most. When users deviate from the intended path, they are not necessarily being uncooperative. They are revealing something: that the information architecture does not match how they understand the task, that a label carries a meaning different from what the design team assumed, that a trust threshold is higher than the interface accounts for, or that the real-world workflow depends on a dependency the system never acknowledged.</p><p>Deviation is not noise by default. In many cases, it is the signal the study was designed to detect, surfacing through a path no one planned.</p><p>The distinction worth preserving is not between compliant and non-compliant behavior. It is between the deviation that is diagnostically rich and the deviation that is not. Most moderation training addresses the first distinction well. The second is where experienced judgment makes the difference.</p><div><hr></div><h2>Five conditions for a productive detour</h2><p>Not every departure from the protocol deserves follow-through. A productive detour is a meaningful divergence that reveals something important about user understanding, system design, task context, or decision-making in practice. Recognizing the difference between that and ordinary drift is the core skill.</p><p>Five conditions tend to indicate that a departure is worth following.</p><p><strong>1. Diagnostic value.</strong> The divergence reveals something about the user&#8217;s mental model or the design&#8217;s failure to support it. The path taken shows where the interface pointed in the wrong direction, not merely that the person went the wrong way.</p><p><strong>2. Ecological validity.</strong> The behavior reflects what this user would plausibly do under real operational conditions. A workaround that appears to be an error under controlled conditions may be standard practice in actual conditions.</p><p><strong>3. Structural generalizability.</strong> The finding is likely to apply beyond this one person. If the alternative path makes logical sense given the interface, other users will probably take it too.</p><p><strong>4. Trust or safety signal.</strong> The user is managing risk, not just navigating. Pausing to verify, seeking external confirmation, or refusing to proceed without checking something reflects usability evidence, not task failure.</p><p><strong>5. External dependency exposure.</strong> The user&#8217;s instinct is to leave the system entirely. That instinct, if widespread, may be one of the most important findings a study can yield.</p><p>When more than two of these conditions apply, the case for following the departure is usually stronger than the case for redirecting.</p><div><hr></div><blockquote><p><em>A workaround that appears to be an error under controlled conditions may be standard practice in actual conditions.</em></p></blockquote><div><hr></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WCFH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F287ab24d-b1df-4761-a259-0cdb52eba535_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WCFH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F287ab24d-b1df-4761-a259-0cdb52eba535_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!WCFH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F287ab24d-b1df-4761-a259-0cdb52eba535_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!WCFH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F287ab24d-b1df-4761-a259-0cdb52eba535_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!WCFH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F287ab24d-b1df-4761-a259-0cdb52eba535_2752x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WCFH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F287ab24d-b1df-4761-a259-0cdb52eba535_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/287ab24d-b1df-4761-a259-0cdb52eba535_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5635701,&quot;alt&quot;:&quot;Infographic&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/190650178?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F287ab24d-b1df-4761-a259-0cdb52eba535_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Infographic" title="Infographic" srcset="https://substackcdn.com/image/fetch/$s_!WCFH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F287ab24d-b1df-4761-a259-0cdb52eba535_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!WCFH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F287ab24d-b1df-4761-a259-0cdb52eba535_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!WCFH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F287ab24d-b1df-4761-a259-0cdb52eba535_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!WCFH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F287ab24d-b1df-4761-a259-0cdb52eba535_2752x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Inforgraphic created by NotebookLM</figcaption></figure></div><div><hr></div><h2>What the research tradition tells us</h2><p>This approach is not a rejection of methodological structure. It is grounded in the same research tradition that makes structured observation valuable.</p><p>Don Norman&#8217;s concept of the gulf of evaluation, developed in <em>The Design of Everyday Things</em>, describes the mismatch between how a system represents its state and how a user interprets that state. A person who takes an unexpected path is often navigating that gulf in real time. Watching that navigation is more informative than recording only whether they crossed it successfully.</p><p>Ericsson and Simon&#8217;s foundational work on think-aloud methodology establishes that concurrent verbalization, talking through a task while performing it, produces qualitatively different data than either retrospective accounts or behavioral observation alone. A user who explains why she is taking an unexpected path is generating exactly the concurrent verbalization that think-aloud protocols are designed to capture. Redirecting her too quickly closes that window.</p><p>The moderator&#8217;s job is not to protect the protocol from the participant. The moderator&#8217;s job is to protect the quality of observation. Sometimes those are the same thing. Sometimes they are not.</p><div><hr></div><h2>Measurement and discovery are not the same goal</h2><p>The tension at the center of this issue is a tension between two legitimate but distinct research purposes.</p><p>Measurement values consistency. It asks whether users can complete tasks, how long it takes to complete them, where failures cluster, and whether patterns replicate across participants. Structured protocols serve measurement well. They produce data that can be aggregated, trended, and benchmarked across studies.</p><p>Discovery values explanation. It asks what users expected, what confused them, what assumptions they brought into the session, and what the system led them to believe or do. Flexible facilitation serves discovery better. It produces insight that informs redesign, surfaces hidden assumptions, and explains the why behind the what.</p><p>Most moderated usability evaluations involve some mix of both. The challenge is not to choose one permanently. It is to know which mode applies at a given moment, and to document the shift when it occurs.</p><p>This distinction matters especially because most moderator training emphasizes measurement. The literature on standardization, inter-rater reliability, and protocol fidelity is extensive and well-developed. Hertzum and Jacobsen&#8217;s work on the evaluator effect, which documents the low agreement rate among evaluators reviewing the same interface, partly serves as a corrective to over-reliance on a single session&#8217;s data. The implication is not that structure is wrong. The value of structured sessions depends on having enough of them to detect patterns, which means that any single session shifted toward discovery mode does not meaningfully compromise the larger study, provided the shift is documented.</p><p>A moderator who follows a productive detour is not abandoning rigor. She is shifting from measurement mode to discovery mode in response to evidence. That is judgment, not permissiveness. The mistake is not allowing the shift. The mistake is failing to recognize that it happened and reporting the results as though it did not.</p><div><hr></div><h2>What a productive detour produces: a composite example</h2><p>Consider a study evaluating a prior authorization workflow in a large health system. The formal objective is to locate a pending request and update its status. Several testers complete the evaluation without incident.</p><p>One does not. He navigates to the correct screen, pauses, and says, &#8220;I would normally call the PA coordinator before updating this. If I get it wrong, the claim has to be reversed.&#8221; He reaches for his phone, catches himself, and adds: &#8220;I guess in the test I can just do it.&#8221;</p><p>A strict protocol would record a task success with elevated time. The researcher redirected him, he completed the assignment, and the data point is clean.</p><p>But what that pause revealed is more significant than whether he updated the field. He revealed that the system provides no indication of the downstream consequences of an incorrect status update. He revealed that the practical safeguard for this workflow is a phone call to a colleague, not a system confirmation. He revealed that his confidence threshold for acting without external verification is lower than the interface design assumed.</p><p>Those are not secondary observations. They are the usability findings the study should be built to surface.</p><p>Following that moment briefly, asking him to continue as he naturally would, asking what information he would want before acting, would have produced richer evidence of the system&#8217;s cognitive support failure than any completion metric. The loss of a clean time-on-task number is a small price.</p><p>More importantly, a single finding like this can redirect an entire redesign conversation. The team working from task success rates would have concluded that the workflow performs adequately. The team that followed the pause would know the workflow performs adequately only because users have invented a workaround to verify it. Those are different conclusions. They lead to different design decisions.</p><div><hr></div><h2>The edge case: letting users go fully outside the system</h2><p>There is a more demanding version of this decision, and it should be reserved for rare use.</p><p>Occasionally, a user does not drift within the interface. She stops and says, &#8220;I would just call the help desk,&#8221; or, &#8220;I would pull this up in the other system.&#8221; Standard practice is to acknowledge the comment and redirect. Usually, that is correct.</p><p>But in some cases, the instinct to leave is the finding. If the researcher allows it to continue, what emerges is evidence of dependency: what information the interface fails to provide, what external support structures users have built around it, and the real cognitive cost of the task when the system is the only available resource.</p><p>That kind of session is no longer a standardized task evaluation. It is observational fieldwork conducted under study conditions. The metrics it produces are not comparable to standardized completions, and they should not be reported as though they are. But the findings can be more revealing than anything the formal task would have produced.</p><p>The rule is straightforward: allow it when the behavior is the finding, document it clearly as an exploratory observation, and do not mix those results with standardized task data in the report.</p><div><hr></div><h2>When to redirect</h2><p>Flexibility is not the same as permissiveness. There are equally clear conditions for bringing a user back to the protocol.</p><p>If the study relies on standardized metrics and the departure would compromise the comparability of that session, redirect. The tradeoff is not worth it.</p><p>If the person has misunderstood the instructions rather than responding to the interface, redirect. The divergence is a research artifact, not a design signal.</p><p>If the deviation is outside the scope and unlikely to generalize, it may not deserve limited session time. Interesting is not the same as relevant.</p><p>If the user has shifted into enhancement brainstorming, venting about organizational problems, or speculating beyond what the study can support, redirect with appreciation and move on.</p><p>If following the detour would consume time needed for later tasks in the protocol, weigh the tradeoff explicitly. One genuinely valuable finding may justify the loss. One moderately interesting observation probably does not.</p><p>Good facilitation is not a choice between always following and always redirecting. It is making the choice deliberately, for documented reasons.</p><div><hr></div><h2>Reporting the tradeoff honestly</h2><p>The quality of this approach depends entirely on how it is documented.</p><p>When a session shifts from structured measurement to exploratory observation, that shift should be clearly described in the analysis. What happened under exploratory conditions should not be presented as equivalent to standardized task performance. The deviation, why the researcher allowed it, what it revealed, and what was given up to learn it: all of that belongs in the report.</p><p>Some teams document these observations in a brief facilitator insights section, kept separate from formal task metrics. That approach preserves the value of the finding without overstating what the data can support. It also makes the research team&#8217;s reasoning visible, thereby strengthening the analysis's credibility rather than weakening it.</p><p>Hiding the trade-off to maintain the appearance of methodological purity is worse than acknowledging it. The audience for a usability report is usually sophisticated enough to understand that not every session produces equivalent data. What they need is honesty about what each session produced and why.</p><div><hr></div><h2>Conclusion</h2><p>The protocol is a tool. It is not the purpose of the study.</p><p>A good moderator protects consistency, fairness, and methodological integrity. She also recognizes when participant behavior is producing evidence the protocol alone could not surface. The difference between those moments and ordinary deviation is not always obvious in real time, but it becomes clearer with practice and with a framework for identifying the conditions that make a detour worth following.</p><p>Rigor does not mean blind adherence to a plan. Rigor means disciplined learning. The productive detour is not the collapse of methodological discipline. It is what methodological discipline looks like when it is responsive to evidence.</p><p>The goal of a usability study is to learn something true about how people understand, approach, and experience the system. Sometimes the path to that truth is the one no one planned.</p><div><hr></div><h2>Quick Reference: Five Conditions for a Productive Detour</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Zlki!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8692f0e2-aef6-4f7a-a74f-f8419963dc7b_684x704.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Zlki!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8692f0e2-aef6-4f7a-a74f-f8419963dc7b_684x704.png 424w, https://substackcdn.com/image/fetch/$s_!Zlki!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8692f0e2-aef6-4f7a-a74f-f8419963dc7b_684x704.png 848w, https://substackcdn.com/image/fetch/$s_!Zlki!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8692f0e2-aef6-4f7a-a74f-f8419963dc7b_684x704.png 1272w, https://substackcdn.com/image/fetch/$s_!Zlki!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8692f0e2-aef6-4f7a-a74f-f8419963dc7b_684x704.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Zlki!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8692f0e2-aef6-4f7a-a74f-f8419963dc7b_684x704.png" width="684" height="704" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8692f0e2-aef6-4f7a-a74f-f8419963dc7b_684x704.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:704,&quot;width&quot;:684,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:112958,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/190650178?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8692f0e2-aef6-4f7a-a74f-f8419963dc7b_684x704.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Zlki!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8692f0e2-aef6-4f7a-a74f-f8419963dc7b_684x704.png 424w, https://substackcdn.com/image/fetch/$s_!Zlki!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8692f0e2-aef6-4f7a-a74f-f8419963dc7b_684x704.png 848w, https://substackcdn.com/image/fetch/$s_!Zlki!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8692f0e2-aef6-4f7a-a74f-f8419963dc7b_684x704.png 1272w, https://substackcdn.com/image/fetch/$s_!Zlki!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8692f0e2-aef6-4f7a-a74f-f8419963dc7b_684x704.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/p/the-productive-detour-in-usability?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/p/the-productive-detour-in-usability?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Separating Identity from Method: What Stays When the Workflow Changes]]></title><description><![CDATA[The practitioner&#8217;s guide to professional identity in an era of skill compression]]></description><link>https://johnwbrown.substack.com/p/separating-identity-from-method-what</link><guid isPermaLink="false">https://johnwbrown.substack.com/p/separating-identity-from-method-what</guid><dc:creator><![CDATA[John W Brown]]></dc:creator><pubDate>Tue, 31 Mar 2026 06:02:39 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!is_O!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21c32990-ea8a-4e5a-8ce2-b805009b47f5_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!is_O!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21c32990-ea8a-4e5a-8ce2-b805009b47f5_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!is_O!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21c32990-ea8a-4e5a-8ce2-b805009b47f5_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!is_O!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21c32990-ea8a-4e5a-8ce2-b805009b47f5_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!is_O!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21c32990-ea8a-4e5a-8ce2-b805009b47f5_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!is_O!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21c32990-ea8a-4e5a-8ce2-b805009b47f5_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!is_O!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21c32990-ea8a-4e5a-8ce2-b805009b47f5_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/21c32990-ea8a-4e5a-8ce2-b805009b47f5_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2602943,&quot;alt&quot;:&quot;Old painting blue prints&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/190122363?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21c32990-ea8a-4e5a-8ce2-b805009b47f5_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Old painting blue prints" title="Old painting blue prints" srcset="https://substackcdn.com/image/fetch/$s_!is_O!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21c32990-ea8a-4e5a-8ce2-b805009b47f5_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!is_O!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21c32990-ea8a-4e5a-8ce2-b805009b47f5_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!is_O!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21c32990-ea8a-4e5a-8ce2-b805009b47f5_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!is_O!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21c32990-ea8a-4e5a-8ce2-b805009b47f5_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Image created with generative AI</figcaption></figure></div><div><hr></div><h1>Separating Identity from Method: What Stays When the Workflow Changes</h1><p>The practitioner&#8217;s guide to professional identity in an era of skill compression</p><div><hr></div><p>Consider a scenario where a senior information architect with fourteen years of experience sits in a product review and watches an AI tool generate a site structure for a mid-complexity enterprise portal in about four minutes.</p><p>She knows the tool got several things wrong. She can see the places where it misunderstood the content model, where it optimized for discoverability at the cost of task flow, where it produced something that would test well in a first-click study and frustrate users six weeks into daily use. She could rebuild this in a day. She rebuilt things like this for a decade before she could do it in a day.</p><p>The tool took four minutes.</p><p>She fixes it in forty. And she sits with a question she does not know how to ask her manager, her professional community, or herself: if the output needed her expertise to be right, but the process no longer needs her expertise to begin, what exactly is her role?</p><p>This is not an efficiency question. Anyone can do the efficiency math. This is an identity question, and the professional development literature has not caught up to it.</p><div><hr></div><p>The previous pieces in this series have focused on AI denial as a failure of methodology, and on the framework for responsible integration. But there is a prior problem, one that sits underneath the methodological failure and explains a good deal of it.</p><p>A cultural observer I follow on X, @ConradHannon, put it precisely: &#8220;Identity anchored entirely to a method becomes brittle.&#8221; He was writing about professional culture broadly. But the sentence lands with particular weight in HSI/UX, where professional identity is often built not just on outcomes but on the specific cognitive process used to reach them. Not just &#8220;I do usability work.&#8221; But &#8220;I am someone who knows how to do usability work the right way.&#8221;</p><p>That distinction matters because it determines what feels threatened when the method changes, and whether the response to that threat is adaptation or defense.</p><div><hr></div><h2>What Is Actually Being Compressed</h2><p>It is worth being precise about what AI is and is not doing to HSI/UX practice, because vagueness here produces both unnecessary anxiety and misplaced confidence.</p><p>Generative tools are compressing the execution layer of the work: the time it takes to produce a first-pass deliverable, to generate variants, to synthesize large bodies of existing research, to write the structural scaffold of an evaluation report. These are real tasks. They have historically required real time. In many cases, they have been the primary visible output of the role.</p><p>What these systems are not compressing, at least not yet, and not in the ways that matter most for consequential decisions, is the judgment layer: the assessment of whether a first-pass deliverable is actually correct, the contextual knowledge that determines which variant fits this organization&#8217;s constraints, the synthesis that integrates user research findings with system safety requirements, the distinction between a usability problem and a workflow design problem and a training problem.</p><p>The information architect in the opening scenario did not spend fourteen years learning how to generate site structures. She spent fourteen years learning how to know when a site structure is wrong, and why, and what specifically to do about it. The tool compressed her generation time. It did not compress her judgment.</p><p>This distinction is not a consolation prize. It is the actual structure of the skill, and it has implications for how practitioners should be thinking about their development and their roles.</p><div><hr></div><h2>Where Human Judgment Remains Irreplaceable</h2><p>There are categories of professional contribution that current AI systems cannot replicate, not because the technology is immature, but because the contribution requires things that are structurally outside what AI systems do.</p><p><strong>Contextual accountability.</strong> AI tools produce outputs. They do not own them. In regulated environments, in high-stakes deployment contexts, in situations where an HSI failure will result in a documented adverse event, someone has to stand behind the work. That accountability structure is not merely procedural. It shapes the judgment process itself. The specialist who knows they will be responsible for a recommendation reasons differently about it than one reviewing a generated artifact. Accountability is not a role that can be delegated to a system, and organizations that behave as though it can are building latent failure into their processes.</p><p><strong>Organizational translation.</strong> A finding from a usability evaluation is not self-implementing. It has to be translated into something that fits the organization&#8217;s political reality, its technical constraints, its development timeline, its tolerance for change. This translation work is not documented anywhere. It lives in the specialist&#8217;s knowledge of the organization, built over time through relationships, failed recommendations, and the specific texture of how this team makes decisions. No tool has access to that knowledge. It can generate a recommendation. It cannot negotiate it into the product roadmap.</p><p><strong>Anomaly recognition.</strong> Expert practitioners notice things that do not fit the expected pattern. In usability testing, this is the participant who uses the system in a way that was not anticipated, revealing an assumption baked into the design that no one had questioned. In human factors analysis, it is the incident report that does not look like other incident reports, suggesting a failure mode that the taxonomy was not built to capture. This recognition depends on having internalized the pattern well enough to notice its absence. Current systems are trained on patterns. They are structurally less equipped to surface what lies outside them.</p><p><strong>Ethical load-bearing.</strong> The question of whether something should be built is not a usability question. But HSI/UX practitioners are often in the room when it is being decided, because they are the ones with evidence about how humans actually interact with systems under real conditions. That position carries ethical weight. Exercising that weight requires judgment about values, consequences, and what the evidence actually supports. No tool can substitute for the practitioner who chooses to say, in a product review, that the data suggests this design will cause harm.</p><div><hr></div><h2>The Brittle Identity and What Replaces It</h2><p>The challenge is not that these contributions are diminished. The challenge is that many practitioners have built their professional identity around activities that sit primarily in the execution layer, and the execution layer is precisely what is being compressed.</p><p>A usability specialist who defines herself as &#8220;someone who runs moderated usability tests&#8221; has a different vulnerability than one who defines herself as &#8220;someone who generates behavioral evidence to inform design decisions.&#8221; The first identity is anchored to a method. The second is anchored to a purpose. When the method changes, the first identity is threatened and the second one is not, or not as directly.</p><p>This reanchoring is not comfortable. It requires acknowledging that what you were proud of, the methodology you spent years developing, the specific way you did the work, was always instrumental. A means to an end. The end is still there. The means are changing.</p><p>There is real grief in that. Skill is not just instrumental. The way a practitioner learned to do close reading of a usability session, the specific cognitive texture of a well-run contextual inquiry, the judgment that accumulates through hundreds of evaluation cycles -- these have intrinsic value to the person who developed them, not just utility value to the organizations that benefited. Acknowledging the grief is not weakness. Pretending it is not grief is what produces denial.</p><p>But grief is not a career strategy.</p><div><hr></div><h2>Role Evolution: Where the Work Is Going</h2><p>For HSI/UX professionals navigating this, the question is not whether to adapt but where the adaptation goes. Three role directions are emerging in organizations that are integrating AI into their design and development processes.</p><p><strong>From executor to evaluator.</strong> The specialist who shifts from producing first-pass deliverables to evaluating generated ones is doing a different job, but not a diminished one. The evaluation role requires the same deep expertise the execution role did, and in some respects requires more of it, because the evaluator needs to understand not just what a good output looks like but what the specific failure modes of AI-generated outputs are in their domain. This role is currently underdefined in most organizations. The professionals who define it, rather than waiting for the job description, will be better positioned than the ones who wait.</p><p><strong>From individual contributor to methods owner.</strong> AI tools do not come with quality standards built in. Someone has to define what good looks like for AI-assisted HSI/UX work in a given organizational context, establish evaluation criteria, train the team on error recognition, and own the methodology for calibrating the tool over time. This is a senior practice role, and it is almost entirely vacant in most organizations right now. It requires exactly the expertise that experienced specialists have accumulated. It does not require doing the work the old way.</p><p><strong>From practitioner to integrator.</strong> The most forward-positioned role is the one that sits at the boundary between AI capability and organizational constraint, shaping how these systems are adopted rather than reacting to how they have been deployed. This is the work described in Article 3 of this series. It requires HSI/UX expertise, but it also requires comfort operating in organizational decision spaces that many in the field have historically avoided.</p><p>None of these roles requires enthusiasm about AI. They require engagement. The distinction matters.</p><div><hr></div><h2>Skill Adjacency: What to Develop Next</h2><p>For professionals who want to move intentionally rather than reactively, the question becomes which capabilities to develop that will be most durable as generative tools augment the execution layer.</p><p><strong>AI output evaluation.</strong> This is the most immediately useful adjacent skill: learning to recognize the specific failure modes of AI-generated HSI/UX work. Where do generative tools produce plausible-looking outputs that are wrong in ways that require domain expertise to detect? Building a personal taxonomy of these failure patterns, specific to your domain and your organizational context, is not an abstract exercise. It is the foundation of the evaluator role.</p><p><strong>Research synthesis at scale.</strong> Generative systems can produce summaries of existing research literature. They cannot assess the quality of that literature, weight findings by methodological rigor, or integrate conflicting evidence in ways that are accountable to the specific design question at hand. The specialist who develops genuine expertise in evidence synthesis, not just familiarity with findings, has a capability that is additive rather than substitutable.</p><p><strong>Facilitation and organizational navigation.</strong> The work of translating HSI/UX findings into organizational decisions is almost entirely human. It involves facilitation, negotiation, knowing when to push and when to hold, and the relationship-based credibility that makes a recommendation land differently than a document. These skills have always mattered. They matter more when the upstream work is increasingly automated, because translation is where the human contribution becomes most visible and most consequential.</p><p><strong>Ethical and governance expertise.</strong> As AI integrates into consequential human systems, the demand for professionals who can assess the human factors implications of AI deployment, not just for usability but for safety, equity, and accountability, is growing faster than the supply. This is not a specialization that requires leaving the field. It is an extension of what HSI/UX practitioners already know how to do, applied to a new class of systems.</p><div><hr></div><h2>Conclusion</h2><p>The fourteen-year information architect who watched the tool produce a flawed site structure in four minutes has a choice. She can define herself by the four minutes, by what the tool can now do, and feel the professional ground shifting under her. Or she can define herself by the forty minutes, by what she knew that the tool did not, and recognize that this is where she actually lives.</p><p>The tool compressed the beginning of her process. It did not touch the part that required being her.</p><p>That part is not safe from future compression. Nothing is, permanently. But it is where professional development investment pays the most durable return right now, because it is built on judgment, context, and accountability rather than on the specific sequence of steps that judgment, context, and accountability once required.</p><p>Identity anchored to a method becomes brittle when the method changes. Identity anchored to a purpose becomes a guide for finding where the purpose has moved.</p><p>The method is changing. The purpose, for anyone who chose this work seriously, is not.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/p/separating-identity-from-method-what?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/p/separating-identity-from-method-what?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p>]]></content:encoded></item><item><title><![CDATA[Adaptation with Constraint: A Practitioner’s Framework for Responsible AI Integration]]></title><description><![CDATA[The difference between setting guardrails and performing refusal]]></description><link>https://johnwbrown.substack.com/p/adaptation-with-constraint-a-practitioners</link><guid isPermaLink="false">https://johnwbrown.substack.com/p/adaptation-with-constraint-a-practitioners</guid><dc:creator><![CDATA[John W Brown]]></dc:creator><pubDate>Tue, 24 Mar 2026 06:01:25 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!nAkJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8ef1c83-1f45-4b66-bfbf-002d9340e6d2_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nAkJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8ef1c83-1f45-4b66-bfbf-002d9340e6d2_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nAkJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8ef1c83-1f45-4b66-bfbf-002d9340e6d2_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!nAkJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8ef1c83-1f45-4b66-bfbf-002d9340e6d2_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!nAkJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8ef1c83-1f45-4b66-bfbf-002d9340e6d2_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!nAkJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8ef1c83-1f45-4b66-bfbf-002d9340e6d2_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nAkJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8ef1c83-1f45-4b66-bfbf-002d9340e6d2_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a8ef1c83-1f45-4b66-bfbf-002d9340e6d2_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2844691,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/190117968?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8ef1c83-1f45-4b66-bfbf-002d9340e6d2_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!nAkJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8ef1c83-1f45-4b66-bfbf-002d9340e6d2_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!nAkJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8ef1c83-1f45-4b66-bfbf-002d9340e6d2_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!nAkJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8ef1c83-1f45-4b66-bfbf-002d9340e6d2_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!nAkJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8ef1c83-1f45-4b66-bfbf-002d9340e6d2_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h1>Adaptation with Constraint: A Practitioner&#8217;s Framework for Responsible AI Integration</h1><p>The difference between setting guardrails and performing refusal</p><div><hr></div><p>Consider a scenario in which a usability team at a regional healthcare network has 6 months to evaluate a new AI-assisted clinical documentation tool before deployment across 40 inpatient units. The vendor demo went well. Pilot data from three units looks promising. Nurse satisfaction scores are up. Documentation time is down.</p><p>And the team is paralyzed.</p><p>Not by the technology. By the question that nobody has articulated clearly: which parts of their standard evaluation process still apply, and which parts are now theater?</p><p>They have a heuristic evaluation protocol built for static interfaces. They have a cognitive walkthrough methodology designed for deterministic systems. They have usability testing scripts written before outputs changed based on prior inputs. None of it maps cleanly onto a tool that learns, adapts, and behaves differently at 2 a.m. than it did at 2 p.m.</p><p>So they do what professionals do when the framework does not fit. They argue about the framework.</p><p>Six months later, the tool deploys on schedule. A different team did the evaluation.</p><div><hr></div><p>This is the cost of the gap between &#8220;I will not uncritically adopt AI&#8221; and &#8220;here is my methodology for evaluating it.&#8221; The first is a posture. The second is a practice. HSI/UX professionals have spent two years developing the posture. The practice is still catching up.</p><p>A cultural observer I follow on X, @ConradHannon, recently framed the professional response to AI as a choice between denial and enthusiasm, and described the third path as &#8220;adaptation with constraint.&#8221; The phrase earns its keep precisely because it holds the tension. Adaptation without constraint is abdication. Constraint without adaptation is the plywood protest dressed in professional language.</p><p>This piece attempts to operationalize that phrase.</p><div><hr></div><h2>The Two Failure Modes</h2><p>The HSI/UX field is currently producing professionals who land in one of two failure modes, and both are costing organizations.</p><p>The first is categorical refusal. This is the specialist who treats any AI-assisted workflow as methodologically contaminated. They decline to evaluate AI tools because doing so would require engaging with AI. They write evaluation criteria that presuppose human-only outputs. They advocate for processes that their organizations will eventually route around, because delivery timelines do not accommodate principled abstention. The field loses their expertise not because they quit, but because they made themselves irrelevant to the decisions being made.</p><p>The second failure mode is uncritical adoption. This is the professional who accepts vendor-provided validation studies as sufficient, who treats efficiency metrics as proxies for safety and quality, who mistakes the absence of reported problems for acceptable performance. They are not skeptics who adapted. They are accommodators who skipped the adaptation.</p><p>Both failure modes share a structural problem: they substitute a stance for a methodology.</p><p>Responsible integration is not about what you are willing to accept. It is about what you are willing to verify.</p><div><hr></div><h2>What the Research Already Tells Us</h2><p>The theoretical grounding for this problem is not new. It predates the current AI moment by decades.</p><p>Lisanne Bainbridge&#8217;s 1983 paper &#8220;Ironies of Automation&#8221; identified a paradox that HSI specialists know well: the more reliable an automated system, the more dangerous the operator&#8217;s degraded vigilance becomes. We have designed systems that work well enough that operators stop practicing the skills they would need if the system failed. Bainbridge&#8217;s irony does not disappear with AI. It compounds.</p><p>Lee and See&#8217;s trust calibration model adds the second dimension. Under-trust leads to disuse. Over-trust leads to misuse. Appropriate trust is calibrated trust, and calibrated trust requires evidence, not intuition. The specialist who refuses to engage with an AI tool cannot calibrate trust in it. They have no evidence. They are operating on priors formed before the evaluation began.</p><p>Dan Kahan&#8217;s research on identity-protective cognition explains why experienced practitioners often end up on the wrong side of both failure modes. When a technology threatens professional identity, the cognitive system does not evaluate the threat neutrally. It generates motivated reasoning in the direction that protects the self-concept. Professionals with deep expertise in manual evaluation processes are not better positioned to assess AI objectively. They are, in some respects, worse positioned, because their expertise is precisely what is at stake.</p><p>None of this means experienced practitioners should defer to colleagues with less at stake. It means they should be aware of the specific cognitive pressures operating on their judgment, and build methodologies that account for those pressures rather than pretending the pressures do not exist.</p><div><hr></div><h2>The Constraint Architecture: A Four-Component Model</h2><p>What follows is a working model, not a final one. It is offered as a starting point for teams that need to move from posture to practice.</p><p><strong>Component 1: Decision Classification</strong></p><p>Not all AI-assisted decisions carry the same consequence structure. A documentation tool that pre-populates discharge summaries for physician review operates differently than a diagnostic support system surfacing recommendations under time pressure in an emergency department. Before evaluating any AI-assisted system, classify the decisions it is augmenting along two axes: reversibility (can the output be corrected before it matters?) and consequence magnitude (what is the cost of an undetected error?).</p><p>High-consequence, low-reversibility decisions require human judgment that is load-bearing. The reviewer is not there to feel useful. They are there because their assessment is the last checkpoint before an irreversible outcome.</p><p>Low-consequence, high-reversibility decisions can tolerate AI augmentation with lighter scrutiny. The human review that happens here is often performance theater, and calling it that is not cynical. It is operationally honest.</p><p>The classification matters because treating all AI review as equally important guarantees that none of it is taken seriously. When everything is critical, nothing is.</p><p><strong>Component 2: Oversight Mapping</strong></p><p>Once decisions are classified, map the verification requirement. This is not the same as mapping the verification structure that currently exists. Current structures are usually inherited, not designed. They reflect prior assumptions about where errors occur, which may no longer be accurate once an AI system is introduced.</p><p>Ask the question directly: in this workflow, which human judgments is the system designed to support, and which human judgments has the system effectively replaced? The answer is often uncomfortable. Professionals who believe they are reviewing AI outputs are sometimes ratifying them. The distinction matters enormously for accountability, for training, and for the kinds of errors that go undetected.</p><p>The practitioner who cannot answer this question about a deployed system in their organization does not yet have a methodology. They have a workflow.</p><p><strong>Component 3: Quality Gate Design</strong></p><p>Quality gates are the points in a process where human evaluation is explicitly required and explicitly documented. For AI-assisted systems, they need to be designed with the cognitive reality of AI review in mind.</p><p>The research on automation bias is not ambiguous. Parasuraman and Manzey&#8217;s review of automation-induced complacency documented consistent degradation in human detection of system errors when operators were monitoring rather than operating. If your quality gate is &#8220;a human reviews the AI output before it proceeds,&#8221; you do not have a quality gate. You have a documented handoff point that may or may not involve genuine evaluation, depending on workload, time pressure, and how often the AI has been right recently.</p><p>Effective quality gates specify what the reviewer is looking for, not just that they should look. They are calibrated to the types of errors AI systems produce, not the types of errors humans produce. And they are tested, periodically, to verify that the review is still functioning and has not drifted into ratification.</p><p><strong>Component 4: Constraint Boundaries</strong></p><p>Constraint boundaries define what the AI system is not permitted to do, regardless of its capability. These are not capability limitations. They are deliberate design choices made on grounds that efficiency alone cannot override.</p><p>There are legitimate constraint boundaries: clinical decisions that require human accountability under regulatory frameworks, outputs affecting protected information that require documented review, tasks where the organizational risk of AI error exceeds the efficiency gain. There are also illegitimate constraint boundaries: tasks where the human doing the work is not adding evaluative value, but where removing the human role would force an uncomfortable acknowledgment about what that role was actually doing.</p><p>Constraint is a design choice, not a moral stance. The specialist who cannot distinguish between these two categories is not setting guardrails. They are setting precedents they will not be able to defend when someone with budget authority asks the question.</p><div><hr></div><h2>The Framework in Practice</h2><p>Consider a scenario where a large health system is deploying an AI-assisted nursing handoff tool. The system synthesizes patient data and generates a structured handoff summary that the outgoing nurse reviews and edits before transmission.</p><p>Using the constraint architecture:</p><p>Decision classification: handoff errors are a documented contributor to adverse events in inpatient care. The category is high-consequence, moderate reversibility. Load-bearing human review is appropriate.</p><p>Oversight mapping: the tool is designed to support the outgoing nurse&#8217;s synthesis task and to surface items that might be omitted from a verbal or freeform handoff. In practice, in high-census conditions, some nurses are accepting generated summaries with minimal modification. The designed verification structure and the actual one have diverged.</p><p>Quality gate design: the gate needs to specify what the reviewing nurse is verifying, not just that they reviewed it. A checklist that takes ninety seconds to complete is more useful than a signature field. Periodic audits comparing accepted-without-edit rates by unit and shift provide early warning of drift.</p><p>Constraint boundaries: the system should not generate handoffs without nurse engagement. Full automation of the handoff document, however capable the system becomes, remains outside the constraint boundary because the handoff is also a cognitive transfer process. The nurse reading the generated summary is engaging with the patient&#8217;s status in a way that prepares them to provide care. Eliminating that engagement has costs that efficiency metrics do not capture.</p><p>This is what adaptation with constraint looks like in practice. It is not aesthetically satisfying. It does not resolve cleanly into a policy memo. It requires ongoing calibration, because the gap between designed verification and actual verification tends to widen over time unless someone is specifically watching it.</p><div><hr></div><h2>A Staged Adoption Model</h2><p>Teams moving from current state to this model benefit from a staged approach rather than a full restructuring of their evaluation methodology at once.</p><p><strong>Stage 1: Observe and document.</strong> Use the AI tool in low-stakes applications while systematically documenting its error patterns, edge cases, and performance in your specific organizational context. Do not rely on vendor validation data as your primary evidence base. Build your own. This stage typically runs six to twelve weeks, depending on use case volume.</p><p><strong>Stage 2: Calibrate trust.</strong> Using your documented error taxonomy, develop explicit criteria for what kinds of AI outputs require high-scrutiny review versus lighter verification. Share this taxonomy across the team. The goal is to replace implicit, individual trust calibration with an explicit, shared one. Individual team members will trust the system differently based on prior experience. That variance is a risk.</p><p><strong>Stage 3: Integrate with defined constraint.</strong> Deploy the tool into standard workflow with documented constraint boundaries and quality gates as designed in the model above. Assign someone specific responsibility for monitoring the gap between designed and actual verification. This is not a quality assurance role. It is an active systems monitoring role.</p><p><strong>Stage 4: Reassess quarterly.</strong> The constraint boundaries and quality gates appropriate at deployment will not necessarily remain appropriate at six months or two years. The system changes. The users change. The organizational context changes. Reassessment is not an admission that the original design was wrong. It is the mechanism by which a working design stays working.</p><div><hr></div><h2>Limitations of This Model</h2><p>This approach does not resolve the governance questions that sit upstream of any individual team&#8217;s decisions. If your organization has made AI adoption decisions without adequate HSI input, the model gives you a way to operate responsibly within a decision you did not make. It does not give you a way to retroactively change the decision.</p><p>It also does not address the workforce implications of AI augmentation in HSI/UX practice itself. That is the subject of the next piece in this series.</p><p>And it requires a level of organizational access that not every specialist has. Oversight mapping and quality gate design require visibility into actual workflow, not documented workflow. In organizations where practitioners are downstream of implementation decisions, the model&#8217;s ambition may exceed what the role allows. In those cases, the most useful contribution is often a documented gap analysis: here is what responsible integration would require, here is what we currently have, here is the distance between them.</p><p>Gap analysis is not as satisfying as a structured model. But it is what organizations can act on when the practitioner cannot act directly.</p><div><hr></div><h2>Conclusion</h2><p>The professionals best positioned to shape how AI integrates into consequential human systems are not the ones who refused to engage, and not the ones who adopted without scrutiny. They are the ones who brought a methodology to an environment that was otherwise relying on vendor assurances and deployment timelines.</p><p>That methodology does not need to be elegant. It needs to be explicit, shared, and subject to revision when the evidence changes.</p><p>Adaptation without constraint is not adaptation. It is drift. Constraint without adaptation is not rigor. It is obstruction.</p><p>The work is holding both at once, which is harder than either stance and more useful than either of them.</p><p>The tool is already in the building. The question that remains is whether HSI/UX professionals will be the ones who designed the integration, or the ones who were consulted afterward.</p><div><hr></div><h2>Quick Reference: Adaptation with Constraint</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Uz8m!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F284c69f3-f916-4c68-a864-5a49d31f312f_684x704.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Uz8m!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F284c69f3-f916-4c68-a864-5a49d31f312f_684x704.png 424w, https://substackcdn.com/image/fetch/$s_!Uz8m!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F284c69f3-f916-4c68-a864-5a49d31f312f_684x704.png 848w, https://substackcdn.com/image/fetch/$s_!Uz8m!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F284c69f3-f916-4c68-a864-5a49d31f312f_684x704.png 1272w, https://substackcdn.com/image/fetch/$s_!Uz8m!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F284c69f3-f916-4c68-a864-5a49d31f312f_684x704.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Uz8m!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F284c69f3-f916-4c68-a864-5a49d31f312f_684x704.png" width="684" height="704" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/284c69f3-f916-4c68-a864-5a49d31f312f_684x704.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:704,&quot;width&quot;:684,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:112958,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/190117968?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F284c69f3-f916-4c68-a864-5a49d31f312f_684x704.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Uz8m!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F284c69f3-f916-4c68-a864-5a49d31f312f_684x704.png 424w, https://substackcdn.com/image/fetch/$s_!Uz8m!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F284c69f3-f916-4c68-a864-5a49d31f312f_684x704.png 848w, https://substackcdn.com/image/fetch/$s_!Uz8m!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F284c69f3-f916-4c68-a864-5a49d31f312f_684x704.png 1272w, https://substackcdn.com/image/fetch/$s_!Uz8m!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F284c69f3-f916-4c68-a864-5a49d31f312f_684x704.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><p><strong>When human review is load-bearing:</strong> high-consequence, low-reversibility decisions; regulatory accountability requirements; tasks where the cognitive engagement matters, not just the signature.</p><p><strong>When review may be theater:</strong> low-consequence, high-reversibility tasks; high-acceptance-rate queues with minimal modification; review steps added for comfort rather than error-catching.</p><p><strong>The test:</strong> If the AI output was wrong and went undetected, how would you know? If the answer involves waiting for a downstream consequence rather than a detection mechanism you designed, the quality gate is not functioning as a gate.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/p/adaptation-with-constraint-a-practitioners?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/p/adaptation-with-constraint-a-practitioners?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p>]]></content:encoded></item><item><title><![CDATA[The Human Factors of AI Resistance]]></title><description><![CDATA[Why Smart Practitioners Make Irrational Tool Adoption Decisions]]></description><link>https://johnwbrown.substack.com/p/the-human-factors-of-ai-resistance</link><guid isPermaLink="false">https://johnwbrown.substack.com/p/the-human-factors-of-ai-resistance</guid><dc:creator><![CDATA[John W Brown]]></dc:creator><pubDate>Tue, 17 Mar 2026 06:01:53 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!FlAE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f888f6b-d478-4713-9d90-700b677c272b_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!FlAE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f888f6b-d478-4713-9d90-700b677c272b_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!FlAE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f888f6b-d478-4713-9d90-700b677c272b_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!FlAE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f888f6b-d478-4713-9d90-700b677c272b_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!FlAE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f888f6b-d478-4713-9d90-700b677c272b_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!FlAE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f888f6b-d478-4713-9d90-700b677c272b_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!FlAE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f888f6b-d478-4713-9d90-700b677c272b_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7f888f6b-d478-4713-9d90-700b677c272b_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2623415,&quot;alt&quot;:&quot;The Human Factors of AI Resistance&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://johnwbrown.substack.com/i/190113739?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f888f6b-d478-4713-9d90-700b677c272b_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The Human Factors of AI Resistance" title="The Human Factors of AI Resistance" srcset="https://substackcdn.com/image/fetch/$s_!FlAE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f888f6b-d478-4713-9d90-700b677c272b_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!FlAE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f888f6b-d478-4713-9d90-700b677c272b_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!FlAE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f888f6b-d478-4713-9d90-700b677c272b_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!FlAE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f888f6b-d478-4713-9d90-700b677c272b_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Image created with generative AI</figcaption></figure></div><div><hr></div><h1>The Human Factors of AI Resistance</h1><h2>Why Smart Practitioners Make Irrational Tool Adoption Decisions</h2><p><em>After the Plywood, Article 2 of 4 | Human Systems Integration Brief</em></p><div><hr></div><p>Consider a scenario that illustrates a specific kind of professional paradox.</p><p>A human factors engineer at a medical device firm has spent twelve years documenting how clinicians develop automation bias: the tendency to over-rely on automated systems, to accept machine outputs without sufficient critical evaluation, to allow the algorithm to substitute for clinical judgment in situations where judgment is exactly what the situation requires. She has published on this. She has built evaluation frameworks around it. She has sat in design reviews and told engineering teams, with evidence, that if you do not design for appropriate trust calibration, you will train users to stop thinking.</p><p>When her organization begins piloting an AI-assisted documentation tool, she refuses to participate in the evaluation. The tool is undertested, she says. The validation methodology is insufficient. She does not trust what she cannot fully explain.</p><p>Her colleagues find this puzzling. She does not.</p><p>From the inside, her position is consistent: she is applying rigorous standards. From the outside, something else is visible. A practitioner who has spent a career analyzing miscalibrated trust in human-system relationships has developed a calibration problem of her own. She has diagnosed automation complacency in others for twelve years. She is exhibiting its mirror.</p><p>The field has a name for what she is experiencing. It just has not applied that name to this particular situation yet.</p><div><hr></div><p>The previous article in this series identified six patterns of AI denial visible in HSI and UX teams: the methodology purist, the scope deflector, the credential protector, the standards invoker, the principled abstainer, and the organization that opted out. Naming the patterns is useful. Understanding why they persist in a field specifically equipped to analyze them is more useful still.</p><p>That is what this article takes up. The frameworks required are not new. They are the field&#8217;s own.</p><div><hr></div><h2>Miscalibrated Distrust: The Mirror Nobody Wants to Look At</h2><p>In 2004, Jens Rasmussen&#8217;s successor work on trust in automation was extended by John Lee and Katrina See into what became the foundational trust calibration framework in human factors research. Their analysis distinguished between trust that is appropriately calibrated to a system&#8217;s actual reliability and trust that is miscalibrated in either direction. Automation complacency, the over-trust failure mode, received the bulk of subsequent research attention. It is dramatic: planes crash, diagnoses are missed, alarms are ignored. The under-trust failure mode is quieter, and the literature treated it as the lesser problem.</p><p>Lee and See were careful to note that both represent failures of calibration. Undertrust produces its own category of error: capable systems go unused, human workload remains unnecessarily high, the performance benefits of automation are forfeited, and the humans who should be supervising automated systems at the appropriate level of oversight instead reject them entirely.</p><p>What the trust calibration literature did not anticipate, because it was largely focused on operators interacting with deployed systems, is what happens when the humans exhibiting miscalibrated distrust are not end users but domain experts. Practitioners with deep knowledge of a technology area, or of the human factors implications of technology adoption generally, are not immune to miscalibration. They are, in certain conditions, more susceptible to systematic undertrust than naive users.</p><p>The mechanism is worth understanding. A practitioner who knows what can go wrong with automated systems, who has documented failure modes, who has built their expertise around the ways humans and machines interact badly, has a well-developed mental model of risk. That mental model is an asset in most contexts. When a genuinely new capability is introduced, one that does not fit cleanly into existing failure taxonomies, the same mental model can generate distrust that outruns the evidence. The expert knows too specifically what to worry about, and that knowledge becomes a filter that amplifies perceived risk beyond what the system&#8217;s actual reliability warrants.</p><p>This is not irrationality in the colloquial sense. It is a predictable consequence of expertise applied outside its calibration range. The practitioner is running valid heuristics. The heuristics are producing systematically biased outputs because the input is outside the distribution they were developed on.</p><p>Stated plainly: deep domain knowledge about how automation fails does not protect practitioners from miscalibrated distrust. It can produce it.</p><div><hr></div><h2>Bainbridge&#8217;s Irony, Reversed</h2><p>In 1983, Lisanne Bainbridge published what became one of the most cited papers in human factors: &#8220;Ironies of Automation.&#8221; Her central argument was that automation does not eliminate the need for human skill. It changes its character in ways that tend to undermine the very skills it most depends on. Highly automated systems require operators to intervene in novel failure states, but the automation has simultaneously reduced the practice opportunities that would have maintained the skills required for those interventions. The system needs human expertise exactly when the system design has been most effective at eroding it.</p><p>The paper was written about operators. It applies, with uncomfortable precision, to the practitioners who study operators.</p><p>The irony Bainbridge identified runs in two directions. The first, her original formulation, concerns what automation does to the skills of users over time. The second, which the field has been slower to name, concerns what expertise about automation does to the judgment of experts when they become the relevant actors.</p><p>Here is the reversed irony: the practitioners who have spent careers documenting how automation changes human performance are now encountering a technology that changes their performance. They have the most refined conceptual tools for analyzing this situation. They are also, for reasons that the trust calibration literature helps explain, among the practitioners most likely to apply those tools selectively, documenting the risks with precision while declining to evaluate whether those risks are calibrated to actual system reliability.</p><p>Consider what this looks like in practice. A senior usability specialist who has spent years evaluating enterprise software knows, with genuine expertise, that AI-generated interface recommendations frequently optimize for engagement metrics at the cost of user control. That knowledge is accurate. Applied as a prior to every AI tool encounter, it produces a stance that is confident, articulately justified, and systematically resistant to evidence that any particular tool behaves differently than the class. The expertise generates the distrust. The distrust prevents the evaluation that would update the expertise. The loop closes.</p><p>The evaluator effect, documented extensively in usability research by Hertzum and Jacobsen, shows that expert evaluators find different problems than naive users, and that expertise does not reliably produce more accurate overall assessments. It produces more confident ones. The mechanism is not limited to usability evaluation. Expert confidence in a risk assessment does not make the assessment accurate. It makes it feel accurate.</p><p>The field built its credibility on the insight that humans, when interacting with complex systems, behave in ways that differ predictably from how they believe they behave. That insight does not have a practitioner exemption.</p><div><hr></div><h2>Identity-Threat as Cognitive Bias</h2><p>The trust calibration and Bainbridge analyses explain the epistemic dimension of AI resistance. They do not fully explain the intensity. Miscalibrated distrust, taken alone, should produce skepticism and calls for better evidence. What the field is actually seeing in many cases is categorical refusal, professional identity construction around non-use, and significant emotional investment in positions that resist updating.</p><p>That affective intensity points to a different mechanism, one the field knows under the heading of identity-protective cognition.</p><p>The research tradition associated with Dan Kahan and colleagues at Yale demonstrated that on topics where factual beliefs intersect with group identity or professional self-concept, motivated reasoning produces systematically different information processing. People do not simply weigh evidence. They weigh evidence while simultaneously managing what the evidence implies about who they are. When the two processes conflict, identity protection tends to win.</p><p>For a practitioner whose expertise was built on mastering a specific form of difficulty, AI tools create an identity-threatening inference. If a system can approximate the output that fifteen years of practice produced, what does that say about the fifteen years? The rational response, updating beliefs about the tool&#8217;s capabilities while maintaining the professional identity grounded in judgment rather than method, requires a distinction the emotional response does not readily support. When effort and identity are fused, evidence of compression threatens both.</p><p>What looks like a principled methodological position is often, in part, identity protection. This is not a character flaw. It is a predictable feature of how expertise gets constructed and how humans respond when that construction is challenged. Understanding it as a cognitive mechanism rather than a moral failure changes what interventions are likely to work.</p><p>The credential protector described in the previous article is not being dishonest about their values. They are experiencing a genuine fusion of methodological preference and self-concept, and processing the challenge to one as a challenge to the other. Argument about methodology will not resolve this, because methodology is not what is actually at stake.</p><div><hr></div><h2>How Individual Denial Compounds Organizationally</h2><p>Each of the mechanisms above operates at the individual level. Organizations are not simply collections of individuals, but organizational denial has to start somewhere, and it tends to start with the practitioners who carry the most disciplinary authority.</p><p>When a senior human factors specialist adopts a categorical refusal posture, the organizational effects compound in two directions. Downward: junior practitioners calibrate their own positions against visible senior leadership. A team lead who has made non-use a professional identity provides both permission and pressure for the team to adopt similar positions. The message, whether or not it is stated explicitly, is that engagement with AI tools is a signal of insufficient rigor. Upward: the practitioner with disciplinary authority is often the one whose input shapes procurement decisions, evaluation criteria, and integration requirements. When that voice is absent from AI-related conversations, not because it was excluded but because it declined to engage, the decisions get made without human factors input.</p><p>The organizational outcome is a version of what James Reason called latent failures in his analysis of complex system accidents: conditions that do not immediately cause harm but that degrade the system&#8217;s capacity to respond effectively when stress arrives. The HSI team that has no seat at the AI procurement table, the organization that has routed all AI decisions through IT, the evaluation framework that has no methodology for AI-mediated interfaces, these are latent failures. They are not crises yet. They are the conditions under which crises will be harder to prevent and harder to recover from.</p><p>Reason&#8217;s Swiss cheese model was built around the observation that accidents rarely have single causes. They result from the alignment of multiple independent failure conditions, each individually insufficient to cause harm, but collectively creating an unobstructed path to a bad outcome. The denial patterns described in this series are not independent. They interact. The methodology purist influences the standards invoker. The principled abstainer gives cover to the scope deflector. The organization that opted out did so because the practitioners with the most relevant expertise had already opted out individually.</p><p>The holes line up.</p><p>The point is not that HSI and UX teams are uniquely culpable in AI governance failures. It is that the discipline&#8217;s own analytical frameworks predict exactly this kind of compounding effect when expert practitioners disengage from novel technology domains. Predicting it in others while exempting oneself from the prediction is not rigorous. It is selective application of the literature.</p><div><hr></div><h2>What the Literature Is Actually Saying</h2><p>Taken together, these frameworks are saying something that practitioners in this field are professionally equipped to understand and personally resistant to applying to themselves.</p><p>Miscalibrated distrust is a known failure mode with known consequences. Expertise does not prevent it; in certain conditions, it facilitates it. The ironies of automation run in both directions: practitioners who study how systems change human performance are not exempt from having their own performance changed by systems. Identity-protective cognition explains why rational argument about methodology often fails to shift categorical positions. And individual denial compounds into organizational vulnerability through mechanisms the field has been documenting in other domains for decades.</p><p>None of this is new knowledge. All of it is applicable. The discipline has simply not turned the lens on itself.</p><p>That turn is not comfortable. But discomfort has never been a valid reason to skip an evaluation. The field knows this. It says so in its own literature, in almost exactly those words, in the context of the users it studies.</p><p>The next article in this series moves from diagnosis to framework. What does responsible AI integration actually look like for an HSI or UX practitioner? Not enthusiasm, not uncritical adoption, but the specific guidance that &#8220;adaptation with constraint&#8221; requires when the practitioners doing the adapting are the ones who understand the constraints best.</p><div><hr></div><h2>The Instruments Are Available</h2><p>The scenario at the opening of this article describes a human factors engineer with a calibration problem she cannot see from the inside. That is not an unusual situation. It is the situation the field was built around.</p><p>What would she do if a clinician described the same dynamic? She would reach for the evaluation tools. She would design a study. She would recommend staged implementation with monitored outcomes and explicit criteria for adjusting trust levels based on observed reliability. She would distinguish between the initial distrust that appropriate validation requires and the categorical distrust that forecloses validation entirely.</p><p>She has the instruments. The question is whether she will use them on herself.</p><p>That question is not rhetorical. It has a practical answer, and that answer is what the next article is for.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://johnwbrown.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://johnwbrown.substack.com/subscribe?"><span>Subscribe now</span></a></p>]]></content:encoded></item></channel></rss>