Design
Eight sessions recorded, and the same hesitation in six of them
Session notes turned into a task-by-task table of where people stalled and what they said.
What the worker does
Usability testing synthesis with an AI worker produces a table rather than a narrative. The worker reads session notes from Google Drive and returns one row per task attempted: how many participants completed it, where they hesitated, what they said at that moment, and a severity score against the rubric in its skill file. Deciding what to change in response stays with the design team.
- Input
- Session notes and transcripts
- Connections
- Google Drive, Notion, Linear
- Output
- A task-by-task table
The findings that survive are the ones somebody remembered
After eight sessions, the team remembers the participant who got angry and the one who could not find the settings. Those two moments shape the redesign. The quieter pattern, where five people hesitated for three seconds at the same step and then recovered, does not get discussed at all, because recovering looks like success in a note.
A task-level table catches that pattern, because it counts hesitations rather than failures.
The shape of the delivered table
One row per task, one column per thing you will argue about.
| Column | What goes in it |
|---|---|
| Task | The task as the participant was asked to perform it |
| Completed | Count of participants who finished unaided |
| Assisted | Count who finished after a prompt from the moderator |
| Hesitation point | The specific step where the pause happened |
| Verbatim | What the participant said at that moment, in their words |
| Severity | Scored against your rubric, with the rule cited |
How this differs from research synthesis
Both read transcripts. They answer different questions and are briefed differently.
Usability synthesis is task-shaped
The unit is a task attempt and the outcome is a completion rate with an observed failure point. It tells you where the interface is wrong.
Discovery synthesis is theme-shaped
The unit is a topic raised across participants, with quotes and counts. It tells you what problem people have, which is a different question.
Mixing them buries the interface finding
A hesitation at step four is specific and actionable. Folded into a theme about confidence in the product, it stops being either.
Turning rows into work
Each severity-one row can become a task in the design workstream with the verbatim quote in the description, which is the single most effective way to keep the finding alive through three weeks of implementation.
The Focus lane holds the ones being fixed this week across three horizons, so a usability finding does not slide behind whatever arrived on Monday.
Questions people ask
+Does the worker watch session recordings?
No. It reads notes and transcripts you have already stored in Google Drive. Moderating sessions and taking observational notes stay with the researcher, and the quality of those notes sets the ceiling for the synthesis.
+How does it score severity?
Against the rubric written in its SKILL.md file, and it cites which rule it applied for each row. A rubric like blocked without help is severity one, hesitated and recovered is severity three makes the scoring checkable rather than a matter of tone.
+Can it compare two rounds of testing?
Yes, if both rounds are in folders it can read and the tasks are named consistently. The comparison is usually delivered as a second table showing completion counts side by side, which is the fastest way to see whether a fix worked.
+What if the sessions were unmoderated?
Unmoderated sessions produce fewer verbatims and no assisted-completion column, so the table will be sparser. The worker reports the columns it could fill and names the ones it could not rather than leaving the gap ambiguous.
Related
Design work an AI teammate can take, and the part it cannot
Five jobs around the work, none of them the work itself. Taste stays where it belongs.
Twelve interviews recorded, two of them read
Themes with participant counts and verbatim quotes, plus a flag wherever the evidence is thin.
Twenty minutes of every design review goes on finding the file
The agenda written the day before, with last review's unresolved threads at the top.
Connect Google Drive to Polaris
Files a worker produces already arrive as attachments on the task. Files your team already keeps in Drive are the part still waiting.
Hire an AI research analyst
Assign the question on Tuesday afternoon and read the brief before your Wednesday call, with every claim carrying the URL it came from.
Focus lane
A commitment for a horizon, not a filter over everything you have.
Acceptance criteria
Written before the work, checkable after it, and binary either way.