UX & Usability Testing: The Complete Guide
Watch real users attempt real tasks, find what breaks, and prove what got fixed, all on one evidence chain.
Last updated July 202601What this is
Moderated usability testing is the practice of watching real users attempt real tasks on your product or prototype, live and one to one, so you can see exactly where the design helps them and exactly where it fails them.
You give a participant a realistic goal, share their screen, and watch what happens. You do not demo, you do not coach, and you do not ask people whether they like the design. You observe behavior. When a participant hesitates, backtracks, misreads a label, or gives up, you have found something no survey or analytics dashboard can show you: the moment of failure and the reason for it, in the user's own words and actions. Then, after your team ships a fix, you run the same tasks again and prove the fix worked.
Here is the BEYOND difference in one paragraph. On the BEYOND Research Platform, your usability work lives on a single evidence chain: every participant is linked to the tasks they attempted, every task attempt to the clips recorded during it, every clip to the issues it supports, every issue to the retest that verified or refuted the fix, and every report back to all of it. Nothing in a readout floats free of its source. And because Qual Studios sits beside your survey work, you can recruit usability participants directly from your own survey respondents, with their segment membership and prior answers attached, so a session with "a heavy user in the value-seeker segment" is not a screener claim but a verified fact from your own data.
02When to use which method
Qual Studios supports two lanes, moderated and unmoderated, and several methods within them. Pick the one that answers your actual question; do not default to the most familiar. Each row below is a distinct study type in the platform, with its own setup flow and its own reporting rules. Section 11 covers the unmoderated lane in depth.
| Method | What it answers | Stimulus | Output |
|---|---|---|---|
| Moderated usability session | Can users do X, and when they fail, why do they fail? Run with 5 to 8 participants per user group. | A prototype or the live product, shared on screen in a live 1:1 session. | Task outcomes, timed observations, evidence clips, a prioritized issue list. |
| Unmoderated recorded testing | Can users do X on their own, at larger numbers, without a moderator in the room? Send a link and let behavior answer. | An in-app prototype worked solo from a survey link; behavior instrumented, with opt-in camera and voice. | Success and first-click, paths, misclicks and heatmaps, per-task ease and confidence, opt-in recordings, a session replay. |
| Rapid design and IA tests | What does a screen convey at a glance, how do people group concepts, and can they find things in your navigation? | A screen, a set of cards, or an information tree, delivered through the same survey infrastructure. | Five-second recall, a card-sort grouping map, and tree-test findability with the path taken. |
| UX diary study | How does adoption and habit form over days or weeks of real use, in the user's own environment? | The live product, used naturally; participants report on an async board with prompts from the exercise library. | Longitudinal entries, moments of friction and delight over time, themes across the cohort. |
| Expert heuristic review | What does a fast inspection against established usability principles reveal, before any participant is recruited? | The product or prototype, inspected by a trained reviewer. No participants. | A findings list, always labeled as expert findings, never mixed with participant evidence. |
| Accessibility session | Is the product usable with assistive technology, by people who rely on it every day? | The product, operated by the participant with their own tools (screen reader, magnification, switch input) in a moderated session. | Task outcomes and barriers specific to assistive-technology use, with evidence clips. |
A useful habit: if your question starts with "can users" or "why do users fail to", you want a moderated session for depth, or the unmoderated lane when you want the same tasks at larger numbers without a moderator. If it starts with "what does this screen say at a glance", "how do people group these", or "can they find it in the menu", reach for a five-second test, a card sort, or a tree test. If it starts with "do users keep", you want a diary study. If you have no participants and need answers this week, commission an expert review and label it as such. If your question is about assistive technology, nothing substitutes for an accessibility session with a participant using their own setup.
03Before you start: the study contract
Every usability study on the platform begins with the study wizard, a six-step flow that turns your intent into a contract the whole study honors. Take it seriously; everything downstream inherits from it.
The six steps
- Decision. State the decision this study will inform, in one sentence. "Should we ship the redesigned checkout, or hold it for another iteration?" is a decision. "Learn about the checkout" is not.
- Artifact and version. Name exactly what is being tested: which prototype file or which build of the live product, down to the version identifier.
- Mode and protocol. Choose formative or summative, and choose the think-aloud protocol (see below).
- Tasks. Write the tasks participants will attempt, with success criteria and test data. Section 4 covers this in depth.
- Instruments and comparison. Select the measures you will collect (task outcomes, SEQ, SUS) and, if this is a retest or a comparison round, the baseline you are comparing against.
- Read-back and confirm. The wizard reads the whole plan back to you as a single statement: the decision, the artifact and version, the mode, the tasks, the instruments. You confirm it, and the plan freezes.
No study fields without a stated decision. The wizard will not let you configure tasks, instruments, or anything else until step one is complete. This is deliberate: a study that cannot name its decision cannot be scoped, and its findings cannot be judged relevant or irrelevant.
The exact product version is recorded, and findings from different versions never mix. If the build changes mid-study, the platform treats subsequent sessions as a separate version cohort. You cannot average a task's success across two different designs and call it one number.
Formative or summative
A formative study exists to find and diagnose problems while the design can still change. It favors rich think-aloud, generous probing, and flexible task order. A summative study exists to measure how a finished or near-finished design performs against defined criteria. It favors strict consistency: same tasks, same order, same wording, minimal intervention. Most studies you run will be formative. Choose summative when you need numbers you can defend, such as a pre-launch gate or a wave-over-wave comparison.
Protocol: how participants think aloud
The default is thinking aloud during tasks (concurrent think-aloud): participants narrate what they are doing, expecting, and reacting to while they work. It produces the richest diagnostic material and it is right for nearly all formative work. Choose retrospective think-aloud (participants work silently, then walk you back through a replay) or a silent protocol when clean timing matters, because narration slows people down and contaminates any time measure. The wizard asks you to record the reason whenever you leave the default, and that reason appears in the method appendix of your report.
The freeze
Once you confirm the read-back, the study contract freezes. Tasks, criteria, and instruments can no longer be silently edited; any change after freeze creates a visible amendment with a timestamp and a reason. This is what lets a client, or a future you, trust that the plan reported is the plan that ran.
04Writing tasks
A task is a scenario plus a goal. It is never a set of instructions. The moment your task tells the participant where to click, you are testing their ability to follow orders, not your product's ability to guide them.
The rules, with examples
Never use the interface's own words. If the task contains the label the participant needs to find, you have handed them the answer.
No priming words. Words like "easily", "simply", and "just" tell the participant the task is supposed to be easy, which suppresses their willingness to admit struggle.
One goal per task. A task with two goals produces an outcome you cannot score. Split it.
Actionable, not speculative. Ask people to do things, not to imagine whether they would.
Realistic. The scenario should be something this participant plausibly does. If your recruits are small-business owners, do not hand them a consumer scenario, and vice versa.
Include test data. Give participants everything the scenario requires: a test account, a card number, an address, dates, names. A participant improvising fake data mid-task is a participant who has left the task.
How many tasks
Four to eight tasks fit a 60-minute session. Fewer than four and you are underusing the hour; more than eight and either the tasks are trivial or the session will run long and the late tasks will be rushed and unusable. Order them so an early failure does not block everything after it, and put your highest-priority task early, while the participant is fresh.
The platform lints your tasks as you write them. It flags interface vocabulary, priming words, compound goals, and speculative phrasing, and it blocks route-giving language outright. Treat a lint block as a gift: it is far cheaper to fix a task in the wizard than to discover in session three that every participant has been reading the answer off the task card.
05Running the session
A good session feels unremarkable from the outside: a calm moderator, a participant working and talking, and a quiet stream of structured evidence accumulating in the background. Here is how to produce that.
The 60-minute skeleton
- Minutes 0 to 5: setup and consent. Confirm recording consent on the record, start screen share, and run a short think-aloud practice on something neutral so the participant learns to narrate before it counts.
- Minutes 5 to 10: warm-up. A few easy questions about the participant's context and habits. This settles nerves and gives you interpretation material for later.
- Minutes 10 to 50: task work. The core of the session: 35 to 40 minutes of tasks, pushed one at a time, scored as you go.
- Minutes 50 to 58: debrief. Overall impressions, the SUS questionnaire, and your held-back probes.
- Final minutes: close. Thanks, incentive logistics, stop recording.
Before any screen share begins, walk the participant through the privacy routine every time, without exception: close sensitive windows and tabs, silence notifications, and use the test account only, never their real one. Remind them they may pause sharing at any moment, for any reason, and show them where the pause control is. A participant's stray email on screen is a governance incident, not an anecdote.
Moderation rules
Your job during task work is to be a neutral, curious presence that changes the participant's behavior as little as possible.
- Neutral face, neutral voice. No wincing at failures, no delight at successes. Participants read you constantly.
- Answer questions with questions. "Should I click this?" gets "What would you do if I were not here?" You are not part of the product.
- Count to five before rescuing. Silence and struggle are data. Most moderators intervene too early; the recovery a participant finds on their own is often the most valuable observation of the session.
- Announce and log every hint or assist. If you must help, say so out loud ("I am going to give you a hint now") and log it in the cockpit. The cockpit enforces this: the moment a hint is logged, unassisted-success scoring is locked for that task. There is no way to quietly help a participant and still record a clean success.
- Probe behavior, not hypotheticals. "What were you expecting when you clicked that?" is a good probe. "Would you use this feature in real life?" is not; it invites polite fiction.
The moderator cockpit
The cockpit is the live-session workspace, and it is built so that running the session and capturing the evidence are the same motion. The task rail holds your frozen task list with each task's success criteria in view. Push-to-participant sends the current task's scenario text to the participant's screen, so tasks are always delivered in identical wording. The outcome buttons let you score each task from the full outcome taxonomy the moment it ends, while your memory is exact. Quick tags drop structured markers (hesitation, error, backtrack, quote) on the timeline as you watch. Clip markers flag the start and end of moments worth keeping, so the evidence reel builds itself during the session instead of in a late-night editing pass. The backroom queue lets observers submit questions that only you see, and you decide what reaches the participant in the debrief. And recording health keeps the capture status in your peripheral vision, because discovering after the session that nothing recorded is a mistake you make exactly once.
06Scoring tasks honestly
A task outcome is a judgment, and judgments drift unless they have names. The platform uses a fixed outcome taxonomy so that "success" means the same thing in session one and session eight, and in your study and your colleague's.
The outcome taxonomy
| Outcome | Definition |
|---|---|
| Success, direct | Reached the correct end state by a reasonable path, without notable difficulty. |
| Success, with struggle | Reached the correct end state, but with clear hesitation, errors, or detours along the way. |
| Success, self-corrected | Went wrong, recognized it without help, recovered, and reached the correct end state. |
| Success, after hint | Reached the correct end state, but only after a logged moderator hint. |
| Success, with assistance | Reached the correct end state only because the moderator directly assisted. |
| Failed, wrong end state | Believed the task was done, but the end state was incorrect. Often the most dangerous outcome, because the user would never know. |
| Abandoned | Gave up before reaching an end state. |
| Timed out | The task's time budget expired before an end state was reached. |
| Blocked by environment | A crash, network failure, prototype dead end, or other external problem prevented a fair attempt. |
| Not attempted | The session ended, or was redirected, before this task was presented. |
The three reporting tiers
For reporting, outcomes roll up into three tiers plus exclusions. Unassisted success covers direct, with-struggle, and self-corrected successes: the participant got there on their own. Assisted completion covers success after hint and success with assistance. Failure covers wrong end state, abandoned, and timed out. Blocked-by-environment and not-attempted are exclusions: they are removed from the denominator and reported separately, because they say nothing about the design.
Assisted is not failure, and it is not clean success. It is its own diagnostic tier. Folding assisted completions into success inflates your numbers; folding them into failure hides the fact that a small nudge was enough. Report all three tiers, always.
The standard instruments
After every task, ask the SEQ (Single Ease Question): "Overall, how difficult or easy was this task?", answered 1 to 7 where 7 is very easy. Whenever the answer is 4 or below, the platform prompts the why-probe: one open question about what made it hard, asked immediately, while the difficulty is fresh. The SEQ's value is not the number itself but the mismatches it exposes, such as a participant who struggled visibly and still answers 6, which tells you their expectations, not your interface, absorbed the cost.
At the close of the session, administer the SUS (System Usability Scale): ten standard statements, answered on five-point scales, scored 0 to 100. Two things to know before you ever put a SUS score on a slide: the published average across studies is 68, so 68 is "typical", not "two points off perfect", and the score is not a percentage. A SUS of 72 does not mean 72 percent of anything.
Task times collected under concurrent think-aloud are exploratory, never benchmarks. Narrating slows people down, unevenly. Use those times to compare tasks within the session ("task 3 took three times as long as anything else"), not to claim how long the task takes in the world. If you need defensible times, run a silent or retrospective protocol and say so in the method appendix.
07From observations to issues
Sessions produce raw material. Analysis turns that material into a small number of well-evidenced issues a product team can act on. The platform separates the two so you always know which is which.
Three note types
- Observation records what happened, neutrally. "P3 scrolled past the delivery options twice before noticing the expander."
- Quote records what was said, verbatim and attributed. "P3: 'I assumed that grey bar was a footer.'"
- Issue is an analytical claim that something in the design causes a problem. Issues are built from observations and quotes; they are never a substitute for them.
The anatomy of an issue
Every issue on the platform carries the same eight parts, and the workbench will not consider it complete without them:
- Title in user terms. "Users read the delivery expander as a footer and never open it," not "Expander affordance suboptimal."
- Location. The screen, state, and element where it happens.
- Observed vs expected. What participants did, next to what the design intended.
- Who hit it. Incidence as n of N: "4 of 6 participants."
- Severity with rationale. The 0 to 4 rating and the reasoning behind it, not the number alone.
- Evidence. Linked clips and quotes, drawn straight from the sessions.
- Cause hypothesis. Your best account of why the design produces this behavior, labeled as a hypothesis.
- Recommendation. A concrete, testable change, scoped to the cause.
Severity, 0 to 4
| Level | Name | Meaning |
|---|---|---|
| 0 | Cosmetic | Noticed, mildly untidy, no effect on task outcomes. |
| 1 | Minor | Causes brief hesitation or irritation; users recover quickly on their own. |
| 2 | Moderate | Causes real delay, errors, or repeated confusion; users usually still complete the task. |
| 3 | Severe | Causes task failure or requires assistance for a meaningful share of users. |
| 4 | Catastrophe | Blocks the task outright, causes data loss or a wrong end state users cannot detect, or is certain to drive abandonment. Fix before release. |
Rate severity by weighing three factors together: frequency (how many participants hit it, and how often the situation arises in real use), impact (how much damage it does when it strikes), and persistence (whether users get past it once and are done, or hit it every single time). A rare issue with catastrophic impact can outrank a universal cosmetic one. For any rating of 3 or above, the standard practice is a second reviewer: another researcher reads the evidence and independently confirms or contests the rating before it stands.
Expert findings are never presented as participant incidence. If an issue came from a heuristic review, it is labeled as an expert finding and it carries no "n of N". Mixing inspection findings into participant counts fabricates evidence, and the platform will not let a report do it.
The issue workbench
Issues live in the workbench, where each one moves through a lifecycle from observed (drafted, evidence attached) to verified (evidence reviewed, severity confirmed). The workbench proposes merge suggestions when two draft issues appear to describe the same underlying problem in different words, and you decide whether they are one issue or two. Nothing reaches a report without explicit researcher approval: an unapproved issue simply does not exist as far as deliverables are concerned.
08Honest numbers at small n
Usability testing works at small sample sizes, but only if you report what small samples can actually support. This section is the difference between findings a skeptic can trust and findings a skeptic can dismantle.
Five participants per homogeneous user group will surface most of the prominent problems in a design. This is why the method is efficient: the big problems are big precisely because many people hit them, so a handful of sessions finds them. It follows that three rounds of five, with fixes between rounds, beats one round of fifteen. Fifteen sessions on the same broken design mostly rediscover the same top issues fourteen more times; three rounds of five find the top issues, verify the fixes, and then find the next layer that the first layer was hiding.
"Per homogeneous user group" is doing real work in that sentence. If your product serves distinct segments whose behavior differs, such as administrators and end users, each segment needs its own three to five participants. Five admins tell you nothing reliable about end users.
Counts, not percentages
Below n=20, report counts: "4 of 6 completed without assistance." Never "67 percent." A percentage at n=6 dresses a handful of people in the costume of a population estimate, and every reader who knows better will discount everything else you say. The platform's scorecards follow this rule automatically.
The platform also computes confidence intervals on your success counts, and at usability sample sizes they are wide. That is not a flaw to hide; the width is the finding. "4 of 6, with a plausible range of roughly 30 to 90 percent" tells a decision-maker precisely how much certainty six sessions bought, and precisely what a larger follow-up would be for.
External anchors
Published reference points exist: roughly 78 percent average task completion across published usability studies, an average SEQ around 5.5, and the SUS average of 68. Use them as context, cited with their sources, never as pass or fail lines. They are averages over other people's products, tasks, and users. The comparisons that actually mean something, in order: your own stated target from the study contract, your own prior round on the same tasks, and a competitor tested with the same protocol. Reach for external anchors only when you have none of those, and say so.
09The retest loop
Finding problems is half the method. The other half is proving they got fixed, and this is where the platform's evidence chain pays for itself.
When a new version ships, clone the study onto it. The clone carries the same tasks, the same success criteria, and the same instruments, pointed at the new version identifier. Because the contract froze in round one, "same tasks" is a guarantee, not a hope.
In the retest, every issue from the prior round gets a fate, each with fresh evidence from the new sessions:
- Resolved: participants no longer hit it, demonstrated on the new version.
- Persistent: still occurring, with new clips to prove it.
- Regressed: previously resolved, now back.
- New: introduced by the changes, or newly visible now that a larger issue stopped masking it.
A fix is only verified with retest evidence. "The team changed the label" is a change record, not a verification. Until participants on the new version demonstrate the problem is gone, the issue stays open, and no report may claim otherwise.
Run this loop wave over wave and something valuable accumulates: your own norms. After a few rounds you know what your product's completion rates, SEQ scores, and severity profiles look like, on your tasks, with your users, with full provenance behind every number. At that point external anchors become trivia, because you are comparing against the only baseline that matters: your own product, one round ago.
10Deliverables
The standard readout has six parts, in this order. Resist the urge to reorganize it around the session chronology; nobody who matters cares what happened in what order. They care what you found and what to do.
- Executive summary, answer first. The decision from the study contract, answered in the first sentence, with the two or three findings that drive the answer.
- Task scorecard. Every task with its three-tier outcome counts, SEQ, and exclusions, on the exact version tested.
- Prioritized issue table. Approved issues, ordered by severity, each linked to its clip evidence. A reader who doubts any row clicks through and watches.
- Highlight reel. A short cut of the moments that matter most. Nothing persuades a product team like watching a customer struggle.
- Positive findings. What worked, explicitly, so the team does not "fix" it in the next iteration.
- Method appendix, with limitations as consequences. Who participated, what version, which protocol, and what the limitations mean in practice: not "small sample" but "n=6 in one segment, so incidence figures are directional and no claim in this report extends to the admin segment."
Every artifact in the readout is generated from the evidence chain, so numbers in the scorecard, rows in the issue table, and clips in the reel can never disagree with each other or with the sessions they came from.
11Unmoderated and rapid testing
Not every question needs a moderator. When you want the same task discipline at larger numbers, or a fast read on a first impression, an information architecture, or a single screen, the unmoderated lane runs the study from a link and rolls the results into the same evidence chain and scorecards.
How a self-run study runs
You send a branded link. The participant completes the tasks on their own, on your prototype, with no researcher present. Because the prototype runs inside the platform, the behavior is captured directly: every screen entered, every click, the path taken, the first click on each task, misclicks, and elapsed time. From that the platform builds the aggregate views a moderator never has time to compute by hand: a first-click accuracy figure, a screen-flow and drop-off map, click heatmaps, a misclick rate, and a blended usability index across tasks. These behavioral measures are available for in-app prototypes today.
Recording, without a screen-share fumble
A self-run participant can opt in to record their think-aloud. The platform records their camera and voice, and shows them a live self-view so they can see they are on camera; consent is explicit and the choice is theirs. There is deliberately no screen-share prompt, so no participant has to decide which window to share or risks sharing the wrong one. The on-screen journey is instead rebuilt from the captured behavior, and the cockpit replays the session as one picture-in-picture: the participant's face and voice over the exact screens they moved through, with a marker where each click landed. Transcripts of the think-aloud are a near-term addition.
What each task asks
The per-task questions mirror the moderated instruments, with a few additions suited to a participant working alone. After a completion, the platform asks the SEQ and a short confidence rating, since reaching the right screen and being sure you reached it are different things. When a participant gives up, it first asks what stopped them, separating a design problem ("I could not figure out how to do it") from a broken prototype ("the prototype seemed unfinished"), so a test-environment failure is never mistaken for a usability failure. On a give-up or a low ease score, it asks what they expected to happen instead, which surfaces the mental model rather than only the complaint. A task, once completed, is sealed: a participant cannot return and redo it, so the record of a first attempt stays a first attempt.
Rapid design and IA tests
The same survey infrastructure carries the fast, focused methods, each joining the same reporting chain: a five-second test for first impressions, showing a screen briefly and then measuring what registered; a card sort for how people group concepts; and a tree test for findability against an information architecture, which records the full path and whether it ended somewhere correct.
The behavioral analytics above, success and first click, paths, misclicks, and heatmaps, are captured for in-app prototypes. Testing an external Figma prototype or a live site is on the roadmap; until it ships, do not promise clickstream or heatmap analytics for those targets. Where behavior is not instrumented, report the recorded attempt, the participant's self-reported outcome, elapsed time, and the ease and confidence answers, and say so plainly.
12Governance, privacy, and the BI Agent
Recording and consent
Recording is always explicit and opt-in. A moderated session records the shared screen and audio, with consent captured on the record at the start; the participant can pause sharing at any time, and the pause is theirs to use without explanation. An unmoderated participant records their camera and voice, with a live self-view so they can see they are on camera, and no screen share is requested; consent is theirs to give or decline, and declining does not affect participation. Permission to use a short excerpt as a demonstration clip is asked separately from permission to participate, so one never rides on the other. Recordings and clips follow your project's retention policy: they are retained for the period the project defines and deleted on its schedule, and deletion cascades through the evidence chain, so a deleted recording does not survive as an orphaned clip in a report. Transcripts of the think-aloud are a near-term addition, not yet part of the record.
What the BI Agent does, and what it never does
The BI Agent works alongside you throughout a study. It may draft task plans from your stated decision, flag leading wording in tasks and probes, summarize sessions minutes after they end, propose issue clusters from the accumulated observations, and find supporting and contradicting evidence for any claim you are considering, including the evidence that undercuts your favorite hypothesis.
Four things it never does. It never decides ambiguous outcomes: when a task attempt could be scored two ways, a researcher makes the call. It never assigns final severity: it may suggest, but the rating that stands is a researcher's, with the second-reviewer practice for 3 and above. It never publishes an issue without researcher approval: an unapproved issue cannot reach a deliverable, whoever drafted it. And it never claims a fix worked without retest evidence: the verification rule in section 9 binds the BI Agent exactly as it binds you.
13What is coming
The following are planned capabilities, not current ones, and nothing in this section should shape a study you are scoping today. On the roadmap for Qual Studios: think-aloud transcripts, transcribing the recorded audio so verbatims are searchable and clip-linkable on the evidence timeline; a side-by-side retest comparison view, showing a version against its predecessor issue by issue; external-target testing for Figma prototypes and live sites, with the instrumentation the in-app lane already has; AI theme synthesis of open-ended feedback; and a built-in recruitment panel. Each is designed to join the same evidence chain described in this guide.
Questions about any of this, or about a study you are planning now, go to your BEYOND account team.
AAccessibility session guides
Two session templates for accessibility mode. Both assume a moderated session with the participant using their own assistive technology on their own device. Run them with the same contract, tasks, and honest scoring as any other session; what changes is the setup, what you watch for, and the barriers you record. Capture the participant's assistive-technology details in the participant record (tool and version, magnification level, input method) so a finding can be read in context. Accessibility findings carry WCAG references and stay separate from general usability issues; they never mix into participant incidence for a different source.
Keyboard-only session guide
Setup. Ask the participant to put the mouse and trackpad aside for the session. Confirm they can share their screen and that you can see the focus indicator. Do a thirty-second practice on a neutral page so they settle into tabbing before anything counts.
Running it. Read each task scenario as written, never the steps. Let the participant drive entirely from the keyboard. Do not narrate the interface for them. When they get stuck, use the same hint and assist discipline as any session, and log each one; a keyboard trap that you talked them out of is still an assisted outcome, not a clean success.
What to watch and record. Focus order that does not follow reading order; a focus indicator that disappears or never appears; a control that cannot be reached or activated by keyboard; a keyboard trap the participant cannot escape; a custom widget that ignores arrow keys or Enter and Space; and any point where the participant loses their place after an update. Note the WCAG criterion for each barrier where you can (for example 2.1.1 Keyboard, 2.4.3 Focus Order, 2.4.7 Focus Visible). Mark a clip at the moment the barrier appears.
Scoring. Score task outcomes exactly as elsewhere: unassisted, assisted, or a failure type, never collapsing an assisted completion into a clean success. A task a sighted mouse user finishes easily can be a hard failure by keyboard; that gap is the finding.
Screen-reader session guide
Setup. The participant uses the screen reader and browser they use every day; record which ones and the versions. Ask them to share both screen and audio so their screen reader's speech is on the recording. Confirm the pace: let the screen reader speak fully before you say anything, and never read the screen aloud for them.
Running it. Read the scenario, then stay quiet and listen. The screen reader's speech is your primary evidence, so protect the audio: avoid talking over it. Let the participant navigate by their own strategy, whether by headings, landmarks, forms, or links. Use hints and assists sparingly and log every one; reading out a label the interface failed to expose is an assist.
What to watch and record. Unlabeled controls and images that announce as "button" or nothing; form fields with no associated label or error text; content that changes silently with no announcement; reading order that does not match the visual order; a heading structure that gives no map of the page; a focus that jumps somewhere unannounced; and any place the participant cannot tell what state a control is in. Note the WCAG criterion where you can (for example 1.1.1 Non-text Content, 1.3.1 Info and Relationships, 4.1.2 Name Role Value, 4.1.3 Status Messages). Mark a clip at the announcement, or the silence, that reveals the barrier.
Scoring. Score outcomes honestly and keep the barrier separate from the task result: a participant can complete a task and still hit a serious announcement failure worth reporting on its own. Record both.