Design

Eight sessions recorded, and the same hesitation in six of them

Session notes turned into a task-by-task table of where people stalled and what they said.

What the worker does

Usability testing synthesis with an AI worker produces a table rather than a narrative. The worker reads session notes from Google Drive and returns one row per task attempted: how many participants completed it, where they hesitated, what they said at that moment, and a severity score against the rubric in its skill file. Deciding what to change in response stays with the design team.

Input
Session notes and transcripts
Connections
Google Drive, Notion, Linear
Output
A task-by-task table

The findings that survive are the ones somebody remembered

After eight sessions, the team remembers the participant who got angry and the one who could not find the settings. Those two moments shape the redesign. The quieter pattern, where five people hesitated for three seconds at the same step and then recovered, does not get discussed at all, because recovering looks like success in a note.

A task-level table catches that pattern, because it counts hesitations rather than failures.

The shape of the delivered table

One row per task, one column per thing you will argue about.

ColumnWhat goes in it
TaskThe task as the participant was asked to perform it
CompletedCount of participants who finished unaided
AssistedCount who finished after a prompt from the moderator
Hesitation pointThe specific step where the pause happened
VerbatimWhat the participant said at that moment, in their words
SeverityScored against your rubric, with the rule cited

How this differs from research synthesis

Both read transcripts. They answer different questions and are briefed differently.

  • Usability synthesis is task-shaped

    The unit is a task attempt and the outcome is a completion rate with an observed failure point. It tells you where the interface is wrong.

  • Discovery synthesis is theme-shaped

    The unit is a topic raised across participants, with quotes and counts. It tells you what problem people have, which is a different question.

  • Mixing them buries the interface finding

    A hesitation at step four is specific and actionable. Folded into a theme about confidence in the product, it stops being either.

Turning rows into work

Each severity-one row can become a task in the design workstream with the verbatim quote in the description, which is the single most effective way to keep the finding alive through three weeks of implementation.

The Focus lane holds the ones being fixed this week across three horizons, so a usability finding does not slide behind whatever arrived on Monday.

Questions people ask

+Does the worker watch session recordings?

No. It reads notes and transcripts you have already stored in Google Drive. Moderating sessions and taking observational notes stay with the researcher, and the quality of those notes sets the ceiling for the synthesis.

+How does it score severity?

Against the rubric written in its SKILL.md file, and it cites which rule it applied for each row. A rubric like blocked without help is severity one, hesitated and recovered is severity three makes the scoring checkable rather than a matter of tone.

+Can it compare two rounds of testing?

Yes, if both rounds are in folders it can read and the tasks are named consistently. The comparison is usually delivered as a second table showing completion counts side by side, which is the fastest way to see whether a fix worked.

+What if the sessions were unmoderated?

Unmoderated sessions produce fewer verbatims and no assisted-completion column, so the table will be sparser. The worker reports the columns it could fill and names the ones it could not rather than leaving the gap ambiguous.

Your next hire takes 60 seconds.

The software is free — unlimited people, tasks, workstreams and docs. You pay only for work an AI worker actually delivers, itemised by the hour.

Get started free

Last checked .