AI Does Not Remove Cognitive Load, It Moves It
What a decade of mammography automation should have taught us about AI at work
The views expressed are my own and do not represent any organization I’m affiliated with.
AI tools are often introduced with a simple promise: they will reduce work.
Sometimes they do. They can summarize documents, draft text, classify records, generate code, retrieve information, identify patterns, and reduce the manual effort required to produce an initial output.
But in complex work environments, the story is rarely that simple.
AI does not always remove cognitive load. Often, it moves cognitive load from one part of the task to another.
A person who once wrote a document now reviews an AI-generated draft. A person who once searched for information now evaluates a generated summary. A person who once compared data points now decides whether a recommendation is reasonable. The work did not vanish. It changed shape.
That may sound like a reduction in burden, but supervision is not effortless. Verification is not passive. Trust calibration is work.
This matters for human factors, usability, and human systems integration because AI changes the shape of the task. If teams only measure whether AI reduces production time, they may miss the new work created around review, interpretation, accountability, and recovery.
The real question is not simply, did AI make the task faster? The better question is, where did the human work go?
We are not the first profession to ask it. Radiology asked it twenty years ago, ran the experiment at national scale, and got an answer worth studying before the rest of us repeat the mistake.
The experiment we already ran
Computer-aided detection for screening mammography is the closest thing we have to a controlled, decade-long, population-scale trial of putting a pattern-recognition machine next to a trained expert and asking the expert to verify its output.
The setup was exactly the one most organizations are now building for knowledge work. The FDA approved CAD for mammography in 1998. The Centers for Medicare and Medicaid Services increased reimbursement for it in 2002. The tool spread quickly, until it was used for most screening mammograms in the United States, at a cost of more than four hundred million dollars a year. The machine flagged suspicious regions on the image. The radiologist remained responsible for the final read. A human was in the loop, by design and by regulation.
Then someone measured the whole system rather than the tool.
In 2015, Constance Lehman and colleagues at the Breast Cancer Surveillance Consortium published a study in JAMA Internal Medicine comparing the accuracy of digital screening mammography interpreted with CAD against mammography interpreted without it. The dataset was not a laboratory sample. It covered 495,818 mammograms read with CAD and 129,807 read without it, across 323,973 women. The finding was blunt: screening performance was not improved with CAD on any metric the study assessed, and CAD did not improve individual radiologist accuracy [1].
A tool that worked, in the narrow sense, for over a decade. A tool that reliably produced output. A tool that regulators approved and insurers paid for. And the human-AI system as a whole detected no more cancer than the human alone.
The reason is the entire subject of this essay. The output was nearly free. The work the output created was not, and the system did not support that work well.
Earlier observer studies show the mechanism in close detail. When CAD correctly marked a cancer that radiologists had missed, the radiologists still failed to act on the correct prompt in the large majority of those cases [2]. The mark was there. The information was present on the screen. But noticing a mark and integrating it into a confident clinical judgment are different cognitive acts, and the second one is the hard one. Other work on CAD found something more uncomfortable still: when the system was wrong, readers missed more than they would have caught working alone, because the absence of a mark was quietly read as reassurance [2].
This is the pattern. The tool shifted the radiologist’s task from searching to verifying. Verifying a confident machine is its own skill, performed under its own pressures, and it is not the skill the tool was marketed to replace. Nobody designed for the new task because everybody assumed the tool had removed work rather than moved it.
We are now installing the same arrangement across law, finance, medicine, software, and administration, and we are measuring it the way the early CAD adopters did: by whether the tool produces output, not by whether the human-AI system produces better work.
The visible task and the hidden task
Many AI tools reduce the visible task. A draft appears faster. A summary arrives sooner. A recommendation is generated automatically. A table is populated. A message is categorized.
This creates an immediate impression of efficiency. The user no longer has to start from a blank page, read every document, or manually perform every classification.
But the visible task is not the whole task.
The hidden task may now include checking the output for accuracy, identifying missing context, detecting inappropriate assumptions, deciding whether the system’s confidence is justified, reconciling the output against other information, documenting the basis for acceptance or rejection, and correcting the result when it fails.
In some cases, the human does less typing but more judging. Less searching but more validating. Less generating but more monitoring.
That is not automatically bad. Shifting work can help when the new task is easier, safer, more consistent, and better supported. But shifting work without recognizing it creates risk. A poorly designed workflow may save time at the front end and add burden downstream, make routine cases faster while making exceptions harder to detect, or reduce clerical effort while increasing cognitive responsibility.
This is why “AI saves time” is too blunt as a usability claim. The more useful claim is conditional: AI saves time when the system reduces total work across the task, including verification, correction, coordination, and recovery. CAD failed that conditional test for a decade before anyone added up the total.
Cognitive load does not disappear
To see why the radiologists struggled, it helps to be precise about what was being moved. Cognitive load refers to the mental effort required to process information, solve problems, decide, remember, and act. The basic concern is practical: people have limited attention and working memory, and poorly designed systems waste those resources [3].
AI can reduce some forms of cognitive load. It can organize large amounts of information, convert unstructured material into usable form, suggest next steps, and help users get past a blank page. In the right context, that is valuable.
But AI can also create new cognitive demands. The user may need to understand what the system did, what it did not do, what data it relied on, what it may have missed, whether the output is complete, whether it fits the current case, and whether it can be trusted for the decision at hand.
This is especially important when the output appears polished. A confident draft, summary, or recommendation can look finished even when it contains omissions or errors. Fluency conceals uncertainty. The user must then perform a difficult kind of review: not just editing what is present, but noticing what is absent. The radiologist who failed to act on a correct mark was not lazy. Noticing what a confident system left out, or got wrong, is the hardest perceptual and cognitive work in the whole task.
It is not a minor task. It is the work.
From production burden to verification burden
One of the most common AI shifts is from production burden to verification burden.
Before AI, a user often produced an artifact directly. That could mean writing a note, preparing a report, searching for policy language, or reviewing a case manually. With AI, the user receives a first draft or recommendation. The production burden drops. But now the user must verify the output, and verification is not a single act. The user has to judge whether the output is factually correct, whether important context is missing, whether it fits the intended purpose, whether the system is overgeneralizing from incomplete information, whether a plausible-sounding recommendation is actually wrong, and whether accepting it creates downstream risk.
This burden may be manageable for expert users. It may be much harder for novice users, overloaded users, or users working under time pressure. It may also be harder when the system does not explain its basis, show source material, indicate uncertainty, or support easy comparison with the original information.
In that case, the tool has not eliminated work. It has made a difficult review task appear simple. That is a human factors problem.
A second case, closer to the average desk
Mammography is a high-stakes example with a clean dataset. Most knowledge work is lower-stakes and messier, but the mechanism is identical. Consider a composite case, drawn from common patterns rather than any specific individual or organization.
A mid-sized insurer introduces an AI tool that drafts first-pass responses to customer coverage questions. A claims specialist who used to research and write each response now receives a generated draft and approves, edits, or rejects it. The vendor demonstration measures one thing: average handling time per inquiry, which drops by roughly forty percent. On that metric, the rollout succeeds.
What the metric does not capture is where the specialist’s work went. The draft is fluent and formatted like the specialist’s own writing, so its errors do not announce themselves. Most drafts are correct, which is the problem, because a long run of correct drafts trains the specialist to skim. The few drafts that cite a superseded policy version, or quietly omit an exclusion that applies to this customer’s plan, look exactly like the correct ones. The specialist is now doing low-prevalence error detection, the same task that defeated the radiologists, under a quota that assumes the tool made the job easier.
Six months in, handling time is still down, the specialist reports that the tool is helpful, and a small but rising number of responses are going out with confident, well-formatted, wrong answers that nobody upstream is positioned to catch. The tool did exactly what it promised. The system around it was never redesigned for the task the tool created.
This is a hypothetical, and it is deliberately ordinary. You do not need a cancer diagnosis on the line for the dynamic to bite. You need only a fluent machine, a human held responsible for its output, a workflow that rewards speed, and an evaluation that measures production instead of total burden. Those four conditions describe a large share of the AI deployments now underway.
The problem of polished uncertainty
Human beings are sensitive to presentation. Output that is organized, fluent, and confident tends to feel more credible than output that is hesitant or messy.
Generative AI creates a special version of this problem. It can produce highly readable material even when the underlying answer is incomplete, poorly grounded, or wrong. A human reviewer must separate fluency from reliability, and that is not always easy. A rough draft written by a person usually carries visible signs of incompleteness: notes, gaps, questions, awkward transitions, uncertain phrasing. A generated draft hides those seams. It can look complete before it has earned that confidence.
This creates what might be called polished uncertainty. The uncertainty has not disappeared. It has been wrapped in professional-looking output.
For users, this changes the review task. They cannot only ask whether the text reads well. They have to ask whether it is true, whether it is complete, whether it is appropriate, what evidence supports it, and what the system left out. Those questions take effort. If the interface does not support them, the user supplies that effort alone, or skips it.
Automation changes attention
Automation does not merely perform tasks. It changes what people attend to.
When a person performs a task manually, attention is distributed across the steps of the work. The user sees the material, makes small judgments, notices anomalies, and builds a sense of the case through direct interaction. When AI performs part of the task, the human may enter later in the process, receiving a result rather than constructing it. That can be efficient, but it can also reduce situation awareness. The user may know what the system recommends without understanding how the recommendation emerged, see a summary without knowing what was excluded, or approve a draft without having engaged the underlying details.
This is not a new concern. Automation research has long recognized that people become overreliant on automated aids and less prepared to intervene when those aids fail [4], [5]. The out-of-the-loop problem was named in aviation and process control decades ago. The specific tools are new. The design challenge is familiar. The human must remain meaningfully engaged with the work, not placed at the end of the pipeline as a formal approver.
Human-in-the-loop is not enough
Many AI systems are defended by saying there is a human in the loop. The phrase can be useful, but it can also obscure more than it explains.
A human in the loop may be an active decision-maker. They may also be a rushed reviewer, a rubber stamp, a downstream recipient, an exception handler, or a person held accountable for an output they had little practical ability to evaluate. The radiologists in the CAD studies were unambiguously in the loop. Regulation required it. It did not help, because being in the loop and being equipped to catch the machine’s errors are not the same condition.
The key question is not whether a human appears somewhere in the workflow. It is whether the human has enough information, time, authority, skill, and interface support to perform the role assigned to them. If the system expects the user to catch errors, the interface has to make error detection realistic. If the user is expected to calibrate trust, the system has to communicate uncertainty and limits. If the user is expected to override the system, the workflow has to make override feasible. If the user is accountable for the final action, the system has to support traceability and review.
Otherwise, human-in-the-loop becomes a label for accountability without control.
The new work of trust calibration
Trust is not a simple target. The goal is not maximum trust. The goal is appropriate trust.
Overtrust leads users to accept poor outputs, ignore contradictions, or stop checking, which is the CAD failure exactly. Undertrust leads users to reject useful support, duplicate effort, or abandon the tool. Both are failures.
AI systems need to help users calibrate. Users need to understand what the system is good at, where it is limited, when it is uncertain, what evidence it used, and which cases demand extra caution. This is not only an ethics issue. It is a usability issue. A system that delivers every answer in the same confident tone makes calibration harder, and one that hides its source material pushes users toward either blind acceptance or blanket skepticism.
Good design helps users ask better questions of the output: what is this based on, what evidence supports it, what alternatives were considered, what information was unavailable, and what would make this recommendation wrong. Those questions should not depend entirely on the user’s memory or skepticism. The system should support them.
AI review fatigue
As AI systems multiply, another issue will grow: review fatigue.
If every tool generates drafts, summaries, alerts, classifications, and recommendations, users may spend an increasing share of the day reviewing machine output. Each individual review seems manageable. In aggregate, the burden becomes substantial, especially in environments already crowded with alerts, dashboards, messages, and forms. The user may not experience AI as one helpful assistant. They may experience it as another stream of things to check.
Review fatigue leads to shallow checking, missed errors, overreliance, irritation, avoidance, and informal workarounds. Some users accept outputs because the cost of careful review is too high. Others stop using the tool because the review burden cancels the promised efficiency.
The practical lesson is simple. AI-generated output should be treated as workload until proven otherwise. It is not free just because the system produced it automatically.
The importance of role clarity
AI changes roles, and that change should be explicit. Is the user asking the system for help? Is the system making a recommendation, drafting something for human revision, or making a classification that drives downstream action? Is the human expected to approve, edit, override, monitor, or explain the output?
These roles are different, and they need different interface support. A reviewer needs access to source material. An approver needs confidence information and traceability. An editor needs clear boundaries between human and machine-generated content. A supervisor needs exception visibility. A downstream recipient needs to know what role AI played in the information they are now relying on.
If the design does not clarify the human role, the organization will still assign responsibility after something goes wrong. That mismatch creates risk for both the user and the people affected by the system. Role clarity belongs in AI workflow design from the beginning, not in the incident report afterward.
What usability testing should measure
If AI shifts cognitive load, usability testing has to measure more than speed and satisfaction.
A test that only asks whether users completed the task can miss the central issue. Users may finish quickly while failing to notice an error, call the tool helpful while misunderstanding its limits, or accept a recommendation because it sounds reasonable rather than because they verified it. The CAD rollout would have passed a naive usability test for years.
AI usability testing should therefore observe verification behavior, error detection, trust calibration, and recovery. Did users notice incorrect or incomplete output? Did they recognize when the system was uncertain? Could they trace the output back to supporting information? Did they know when to accept, reject, edit, or escalate? Did the output shift their confidence appropriately, or just shift it? Did they over-rely under time pressure? Did the workflow support correction? Did total task burden actually fall, or did it just move to a later step?
These questions move evaluation from surface usability to operational usability. That is the level at which AI systems should be judged.
Design principles for reducing transferred burden
If AI shifts work, designers should make the new work visible and supportable. None of what follows is novel; it converges with existing human-AI design guidance [6], [7], [8]. Several principles follow.
Show the basis for the output, so users can inspect source material, inputs, or evidence in a form appropriate to the task. Communicate uncertainty, because not every output deserves the same confidence, and the system should help users separate routine, well-supported results from uncertain or high-risk ones. Design for comparison, so that verifying a summary, recommendation, or classification against the original is efficient rather than a separate research project. Support exception handling, because AI tends to perform well on common cases and worse on edge cases, and the design should make exceptions visible with clear paths for escalation, correction, or override. Preserve situation awareness, so users are not reduced to final-stage approvers who see only the answer. Measure downstream burden, since a tool that saves time for one role may create work for another, and evaluation should follow the output through the whole workflow. Finally, define accountability honestly: if a human is responsible for an AI-supported decision, the system has to give that human meaningful control.
These principles are not exotic. Most of them are what the mammography workflow lacked. CAD marked the image but never helped the radiologist weigh the mark, compare it against their own read on equal footing, or treat a missing mark as anything other than reassurance. The design assumed the hard problem was detection. The hard problem was integration.
What the radiologists should have told us
AI can be useful. It can reduce effort, improve access to information, support drafting, accelerate review, and help people manage complexity. But it is not automatically a cognitive load reducer. In many real systems it changes the user’s task from doing to checking, from searching to judging, from producing to supervising, and from acting directly to managing uncertainty. That shift may be valuable. It must be designed.
This is why the first step is never simply adding AI to a workflow. It is understanding the workflow well enough to know what burden the tool will shift, who will inherit it, and whether anything supports them. That means treating the user as an operator inside a system, not a satisfied end user of a feature. For human factors and HSI practice, the essential question is not whether AI can perform a function. It is whether the human-AI system supports safe, effective, understandable, and accountable work. A useful AI system should not simply generate output. It should support the human work required to evaluate that output. It should make uncertainty visible, make verification feasible, preserve situation awareness, clarify roles, and reduce total burden rather than relocate it.
The promise of AI is not that humans will stop thinking. In high-stakes systems, the promise should be that humans think better, with better support, better context, and better tools for knowing when the system is wrong.
So the practical question is worth asking before the rollout, not after the audit. When the tool produces its output, where does the cognitive work go? If the answer is “to the user,” the next question is whether anyone has designed for that user: whether they can verify the output, see its uncertainty, recover from its errors, override it, explain the final decision, and stay meaningfully in control.
Computer-aided detection answered those questions by default for more than a decade, at a cost of four hundred million dollars a year, and the answer was no. The technology was never the point. AI should not be evaluated only by what it produces. It should be evaluated by the work it leaves behind.
References
[1] Constance D. Lehman, Robert D. Wellman, Diana S. M. Buist, Karla Kerlikowske, Anna N. A. Tosteson, and Diana L. Miglioretti, “Diagnostic Accuracy of Digital Screening Mammography With and Without Computer-Aided Detection,” JAMA Internal Medicine 175, no. 11 (2015): 1828–1837. https://doi.org/10.1001/jamainternmed.2015.5231.
[2] Robert M. Nishikawa, Robert A. Schmidt, Michael N. Linver, Alan V. Edwards, John Papaioannou, and Michael A. Stull, “Clinically Missed Cancer: How Effectively Can Radiologists Use Computer-Aided Detection?,” American Journal of Roentgenology 198, no. 3 (2012): 708–716, reporting that radiologists failed to recognize a correct CAD prompt in 71 percent of missed-cancer cases. On over-reliance when CAD is absent or incorrect, see Melina A. Kunar et al. on binary CAD and the low-prevalence effect.
[3] John Sweller, “Cognitive Load During Problem Solving: Effects on Learning,” Cognitive Science 12, no. 2 (1988): 257–285.
[4] Raja Parasuraman and Dietrich H. Manzey, “Complacency and Bias in Human Use of Automation: An Attentional Integration,” Human Factors 52, no. 3 (2010): 381–410.
[5] Mica R. Endsley and Esin O. Kiris, “The Out-of-the-Loop Performance Problem and Level of Control in Automation,” Human Factors 37, no. 2 (1995): 381–394.
[6] Saleema Amershi et al., “Guidelines for Human-AI Interaction,” in Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (CHI ‘19), Glasgow, Scotland UK, May 4–9, 2019 (New York: ACM, 2019), 1–13. https://doi.org/10.1145/3290605.3300233.
[7] National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (Gaithersburg, MD: U.S. Department of Commerce, January 2023).
[8] Google, People + AI Guidebook (People + AI Research, PAIR), accessed via pair.withgoogle.com.
#UXDesign #HumanFactors #HSI #CognitiveUX #AIUX #HumanCenteredAI #UsabilityTesting



