The views expressed are my own and do not represent any organization I am affiliated with.
Consider the last hour of a usability study, the part that rarely makes it into the report: the finding-review meeting.
Two evaluators have watched the same recorded session. A participant using a clinical scheduling tool reached the final confirmation screen, paused for roughly forty seconds, backed out to the previous step, returned, and completed the booking. She did not comment on the pause. She did not repeat it across the next three tasks. Afterward she rated her confidence as high.
The first evaluator logs this as a high-severity finding. The confirmation flow, in this reading, is confusing enough to make a user retreat at the moment of commitment. The second evaluator logs almost nothing: a moment of orientation on first contact, gone by the second task, not worth a redesign. Same clip. Same forty seconds. Opposite severity.
The meeting resolves the disagreement the way most meetings do, by seniority, by who speaks with more conviction, or by splitting the difference and calling it medium. What it does not do is resolve the disagreement with evidence. And the reason is worth sitting with: the two evaluators are not actually disagreeing about what happened. They are disagreeing about where to set the line between a problem and a non-problem, and neither of them has said so out loud.
That unstated line is the subject of this piece. Usability testing is usually described as the work of finding problems. It is at least as much the work of deciding which observed behaviors count as problems at all, and that second task is harder, less examined, and more consequential than the first.
The finding that looked worse than it was
A false positive, borrowing the term from diagnostic testing, is a finding that looks serious under study conditions but does not produce real harm, inefficiency, or risk once the system is in use. The participant stumbles, asks a clarifying question, frowns at a label, then recovers and never encounters the issue again. In the room, it is vivid. The behavior is visible, the reaction is real, the note gets written. At scale, the predicted problem never arrives.
These behaviors are worth noticing. The cost comes from what happens after they are escalated. Engineering time goes to a confirmation screen that was never going to slow anyone down. The genuine issue two rows down the priority list waits another quarter. And there is a slower, more corrosive cost. When stakeholders watch three findings marked critical get fixed with no measurable change in any metric they care about, they begin to discount the fourth report before they have read it. Credibility, once spent this way, is expensive to rebuild.
So far this is the familiar argument, and it is correct as far as it goes. But a study that worries only about false positives has solved half a problem and opened another. Every decision to dismiss a finding is also a decision that can be wrong. The behavior you wave off as orientation may be the first visible edge of a defect that will cost thousands of real users real time once the system ships. A practice that takes pride in escalating less is not automatically more disciplined. It may be failing in the opposite direction, and failing quietly, because no one writes a report about the problem they decided not to report.
Four outcomes, one threshold
It helps to borrow a frame from a field that has spent decades on this exact decision. Signal detection theory, formalized for perception research and radar operators in the middle of the last century, describes what happens whenever someone tries to detect a faint signal against background noise.[^1] There are four possible outcomes. You can catch a real signal, which is a hit. You can fail to catch one that was present, which is a miss. You can raise an alarm at noise that meant nothing, which is a false alarm. Or you can stay quiet when there was in fact nothing to report, which is a correct rejection.
Map this onto a usability study and the picture sharpens. A real problem you identify is a hit. A false positive is a false alarm. A real problem you dismiss, or never surface at all, is a miss. And the orientation behavior you correctly decline to escalate is a correct rejection, the quiet, unglamorous outcome that good judgment produces all day and no one ever praises.
The reason this frame earns its place is the part practitioners tend to skip. You cannot drive down false alarms and misses at the same time by simply trying harder or caring more. They trade against each other across a threshold. Lower the threshold, escalate at the faintest signal, and you catch more real problems while also flagging more noise. Raise it, escalate only the unmistakable, and you cut the noise while letting more real problems slip past unreported. The threshold is not a flaw in the method. It is the central decision of the method, and most studies make it by accident.
Why two careful evaluators see different things
The disagreement in that review meeting reflects the normal condition of the work rather than carelessness on anyone’s part, and it has been measured repeatedly.
When Jacobsen, Hertzum, and John had four trained evaluators independently analyze the same four recorded usability sessions, the four together identified ninety-three problems. Only about a fifth of those were caught by all four evaluators, and nearly half were caught by a single evaluator working alone.[^2] When each evaluator then selected the ten problems they considered most severe, the top-ten lists did not converge on a shared core. A later review of eleven studies of this kind found agreement between evaluators ranging from five percent to sixty-five percent, and the effect held for severity judgments specifically, not only for whether a problem got noticed in the first place.[^3]
This has a direct consequence for how the word “severe” should be read in any report. A severity rating is partly a statement about the person who assigned it, not a measurement taken off the system. Which means the inflation of findings does not usually come from bad faith or sloppiness. It comes from a low-reliability instrument being treated as if it were a precise one, and from a threshold that nobody set deliberately doing its work in the background.
The half-true story about the lab
There is a comfortable explanation that makes the whole problem disappear. The lab is artificial, the story goes, so it manufactures problems the field would never produce, and the remedy is to trust real-world use over the test. The trouble is that this story is half right, which is the most dangerous kind of right.
Testing environment does shape what a study surfaces. That much is established.[^4] But the further assumption, that more realism reliably means fewer problems, does not survive contact with the evidence. When researchers evaluated a clinical device across lower and higher fidelity conditions, both conditions surfaced the same types of use error, and increasing ecological validity did not dependably shrink the count or reveal that the lab results had been phantoms.[^5] The lab does not conjure problems from nothing; it changes which problems are salient and how often they appear. That is a reason to interpret findings with care, not a license to assume the field will quietly absolve whatever the lab turned up.
So the false positive is real, and it is worth managing. It is just not the simple artifact of an unrealistic room that the convenient story makes it out to be.
A way to triage without pretending
If the threshold is the real decision, the practical question is how to set it on purpose. Three diagnostic questions do most of the work, and a fourth step ties them to the stakes.
First, recurrence. Did the behavior repeat across participants and across sessions, or did it appear once and vanish? A pattern that shows up in four of eight participants is a different object than a single startled pause. Recurrence is the closest thing a study has to a signal-strength reading, and it gets more trustworthy the more participants you ran.
Second, consequence. Trace what the behavior actually led to. Did it produce a wrong decision, unnecessary rework, an abandoned task, or an unsafe action? Or did the user absorb a few seconds of friction and arrive in the right place with no downstream effect? A finding with no consequence attached to it is a finding waiting for a justification.
Third, trajectory. Distinguish learning from breakdown. New users orient themselves. They pause, test a control, confirm an assumption, and then proceed, and on the next encounter the pause is gone. Breakdown is the opposite shape. It does not resolve with exposure, it recurs or worsens, and it tends to leave a residue of workarounds behind it. The question is whether the struggle is on its way out or on its way in, not whether the user stumbled once.
Those three questions estimate how confident you should be that a finding represents a real-use signal. The fourth step decides what to do with that confidence, and it is the one most studies omit. Calibrate the threshold to the cost structure of the system in front of you. In a safety-critical tool, a missed problem can injure someone, so a miss costs far more than a false alarm, and the rational move is to escalate on thinner evidence and accept that you will chase some noise. In a low-stakes, high-volume consumer flow, the arithmetic inverts, and a threshold that escalates everything will bury the team in redesigns that no user needed. Same method, deliberately different line, set before the findings arrive rather than argued about after.
Two findings from the same study
Picture two findings from a single hypothetical study of that clinical scheduling tool, and run each through the triage.
The first is the forty-second pause from the opening. One participant, one occurrence, resolved by the second task, no downstream error, confidence high afterward. Recurrence is absent, consequence is none, trajectory points toward learning. The triage answer is to document it honestly and decline to escalate it, with the reasoning written down so that the next evaluator who sees the clip does not relitigate it from scratch. This is a correct rejection, and getting it right is a skill, not a failure of diligence.
The second finding is quieter, and worse. Three of eight participants accepted an incorrectly pre-populated field, a default value carried over from a prior screen, without hesitation. No frown, no pause, no comment. By the task-completion metric, all three succeeded. In the room, nothing appeared to go wrong. Recurrence is present, consequence is high, and trajectory is irrelevant because the users never registered a problem to learn their way out of. The triage flags this as severe precisely because visible reaction is not one of its criteria.
That contrast is the whole argument in miniature. The most dangerous finding in a usability study is often the one where nothing appears to go wrong. An evaluation practice tuned to suppress false positives by looking for drama will reliably catch the harmless pause and reliably miss the silent acceptance of a wrong value, which is exactly backward. Friction is easy to see and frequently cheap. The calm error is hard to see and sometimes the one that ships.
Putting it to work
A few moves turn this from a frame into a practice.
Decide the severity rubric and the expected error types before the sessions, not during the debrief. Pre-commitment will not eliminate the evaluator effect, but it shrinks the space in which conviction and seniority substitute for evidence after the fact.
Record two things separately for every finding rather than collapsing them into one number. One is how confident you are that the behavior reflects real use. The other is how often it recurred. A single severity score hides both, and the hidden parts are where the disagreements live.
For any finding you mean to call severe, get a second evaluator on the tape independently before the recommendation leaves the room. The research is unambiguous that agreement is lowest exactly where it matters most, on the severe calls, so that is where a second pass buys the most.
Keep the honest annotations the discipline already knows how to write. “Observed once.” “May be orientation.” “Did not recur.” Far from hedging, these annotations are the confidence half of the record, and a report that omits them overstates what it knows.
Finally, separate what you observed from what you recommend. The observation can be reported with certainty while the recommendation stays provisional. Conflating the two is how a forty-second pause becomes a redesign mandate.
Where this can go wrong
A triage framework is abusable, and pretending otherwise would undercut the point. “It is probably just learning” sits one motivated step away from “we would rather not fix this,” and a team under deadline pressure will find the criteria remarkably accommodating. The protection is sequence. The threshold has to be set before the findings are known, not reverse-engineered afterward to fit the roadmap that already existed.
Recurrence also depends on having enough participants to mean anything. In a five-person study, one occurrence cannot be reliably told apart from a problem that would surface in six of thirty. Small studies should lean harder on consequence and trajectory and treat a single observation as weak evidence rather than a settled non-problem.
The lab-and-field relationship is genuinely conditional, not directional. Studies have found the two equivalent under favorable conditions and divergent under harder ones, which means you cannot promise a stakeholder that the field will confirm or overturn a given finding. You do not know in advance which way any particular finding will move, and claiming otherwise trades one false certainty for its mirror image.
And signal detection is a lens here, not a measurement. I am not proposing that anyone compute a sensitivity index from a think-aloud session. The value is in the shape of the tradeoff and the discipline of naming the threshold, not in a number that would imply more precision than the data can carry.
Back to the meeting
Return to the two evaluators and the forty seconds. They were never really disagreeing about the pause. Each was applying a different threshold and reporting the result as severity, and because neither threshold was visible, the conversation had nowhere to go but conviction. Make the threshold explicit, tie it to what a miss costs against what a false alarm costs in this particular system, and the meeting turns from a negotiation into a decision that can be defended later.
Strong usability practice depends as much on deciding what not to escalate as on identifying genuine problems. But the decision only holds up when both errors are kept in view at once. The loud finding that wastes a sprint and the silent one that ships are failures of the same instrument, set to the wrong threshold in opposite directions. The findings most worth acting on are usually the quiet ones, recurring and consequential while making no noise in the room.
Quick Reference: Triaging a Usability Finding
Before the sessions: set the severity rubric, name the error types you expect, and decide the threshold based on the cost of a miss versus a false alarm in this system.
For each finding, ask:
Recurrence. Did it repeat across participants and sessions, or appear once and vanish?
Consequence. Did it cause a wrong decision, rework, abandonment, or an unsafe action, or did it resolve with no downstream effect?
Trajectory. Is the difficulty resolving with exposure (learning) or recurring and breeding workarounds (breakdown)?
For each finding, log two values separately: confidence that it reflects real use, and frequency of recurrence. Do not collapse them into one severity score.
For any severe call, put a second evaluator on the recording independently before recommending.
Watch the asymmetry. A practice that never produces a false positive is escalating too little, not succeeding. The quiet, consequential, recurring finding is the one most worth protecting from dismissal.
Notes
[^1]: David M. Green and John A. Swets, Signal Detection Theory and Psychophysics (New York: Wiley, 1966).
[^2]: Niels Ebbe Jacobsen, Morten Hertzum, and Bonnie E. John, “The Evaluator Effect in Usability Studies: Problem Detection and Severity Judgments,” in Proceedings of the Human Factors and Ergonomics Society 42nd Annual Meeting (Santa Monica, CA: HFES, 1998), 1336–40.
[^3]: Morten Hertzum and Niels Ebbe Jacobsen, “The Evaluator Effect: A Chilling Fact About Usability Evaluation Methods,” International Journal of Human-Computer Interaction 15, no. 1 (2003): 183–204.
[^4]: Juergen Sauer and Andreas Sonderegger, “Methodological Issues in Product Evaluation: The Influence of Testing Environment and Task Scenario,” Applied Ergonomics 42, no. 3 (2011): 487–94.
[^5]: Romaric Marcilly, Helen Monkman, Sylvia Pelayo, and Blake J. Lesselroth, “Usability Evaluation Ecological Validity: Is More Always Better?” Healthcare 12, no. 14 (2024): 1417.
Bibliography
Green, David M., and John A. Swets. Signal Detection Theory and Psychophysics. New York: Wiley, 1966.
Hertzum, Morten, and Niels Ebbe Jacobsen. “The Evaluator Effect: A Chilling Fact About Usability Evaluation Methods.” International Journal of Human-Computer Interaction 15, no. 1 (2003): 183–204.
Jacobsen, Niels Ebbe, Morten Hertzum, and Bonnie E. John. “The Evaluator Effect in Usability Studies: Problem Detection and Severity Judgments.” In Proceedings of the Human Factors and Ergonomics Society 42nd Annual Meeting, 1336–40. Santa Monica, CA: HFES, 1998.
Marcilly, Romaric, Helen Monkman, Sylvia Pelayo, and Blake J. Lesselroth. “Usability Evaluation Ecological Validity: Is More Always Better?” Healthcare 12, no. 14 (2024): 1417.
Sauer, Juergen, and Andreas Sonderegger. “Methodological Issues in Product Evaluation: The Influence of Testing Environment and Task Scenario.” Applied Ergonomics 42, no. 3 (2011): 487–94.



