Benchmark
AI scribe benchmark: Hanah and six other AI scribes
Run on 6 October 2026 · 13 recorded consultations · 7 AI scribes
We played the same 13 recorded consultations into Hanah, Heidi, Lyrebird, Freed, PatientNotes, CliniScripts and Preve, then graded every note against what was actually said. Every transcript, note and grade is published on this page, with the reason for each one.
100%
of Hanah's notes free of major hallucination flags
The average across the other six scribes was 48%.
66% fewer
clinically significant transcript errors
0.59 per consultation against an average of 1.75.
11% more
required clinical details captured
93.7% against an average of 84.2%.
Results
Scribes are listed by the share of required clinical details their notes captured. Each scribe's note for the first prompt of each recording is scored.
| AI scribe |
Clinical details captured |
Notes with every clinical detail |
Major hallucination flags per note |
Notes with no major hallucination flag |
Significant transcript errors per consultation |
Word Error Rate, English consultations |
| Hanah13 notes |
93.7% |
62% |
0.00 |
100% |
0.59 |
1.6% |
| Heidi13 notes |
91.4% |
38% |
0.23 |
77% |
1.15 |
3.3% |
| CliniScripts13 notes |
90.8% |
62% |
0.23 |
77% |
0.69 |
2.1% |
| Lyrebird13 notes |
90.2% |
46% |
0.62 |
62% |
1.92 |
3.4% |
| Freed†12 notes |
85.7% |
42% |
0.83 |
33% |
1.25 |
19.4% |
| PatientNotes13 notes |
74.7% |
15% |
2.77 |
8% |
1.92 |
3.9% |
| Preve*10 notes |
72.2% |
30% |
1.10 |
30% |
3.54 |
6.7% |
| Average of the other six |
84.2% |
39% |
0.96 |
48% |
1.75 |
6.4% |
A significant transcript error is one a blind rater judged could mislead a clinician: a wrong drug, dose, number, side or symptom, a lost negation or red flag, or clinical content dropped. Word Error Rate counts every word the same, fillers included.
The jaw surgery review was spoken mostly in Mandarin. Each scribe ran it with its input language set to Mandarin where it offers that setting. Heidi made the fewest significant transcript errors on it (3 in each of three runs), then CliniScripts (4) and Hanah (4 to 5 across three runs). PatientNotes made 9, Freed 7, Lyrebird 14 and Preve 24. Preve has no input language setting, and Freed detects the language itself.
How it was run
The recordings
The 13 recordings are Hanah's accuracy corpus, the same set behind the release results on our transparency page. They cover an ICU handover dictation with and without ward noise, physiotherapy, dietetics, occupational therapy, myotherapy and community nursing consults, a child with asthma whose parent speaks limited English, a quad bike fall recorded outdoors, a 50 minute myotherapy session with about seven minutes of speech in a noisy treatment room, and a jaw surgery review spoken mostly in Mandarin. Each has a verbatim reference transcript and a checklist of the details a note must capture.
Capture
Each scribe ran in its own web app in Chrome, signed in to a standard account. The recording was played through the browser's microphone input, the same file for every scribe. After recording, each recording's prompts were typed into the scribe's own document or chat feature word for word, for example "Generate SOAP notes please" and "Summarise this consult in two sentences for a referral letter". Each scribe was run once per recording on 6 October 2026, with English as the input language where a scribe asks for one. For the Mandarin review, scribes with an input language setting were set to Mandarin: Heidi to English and Chinese (Mandarin), Lyrebird and PatientNotes to Chinese (Mandarin, Simplified). CliniScripts produced no transcript with Chinese and English selected on two attempts, so its standard setting is used. Heidi and Hanah ran the Mandarin review three times each, and their transcript results are averaged.
Grading by a Claude jury
The notes were graded by Claude, Anthropic's model. For each recording, every scribe's output was shuffled behind a letter with product names removed, and graded against the reference transcript with a fixed rubric. Each checklist item is marked yes, partial or no. Every statement the transcript doesn't support is listed as a hallucination flag: major when it could change care or misstate the clinical picture (an invented symptom, finding, drug, diagnosis or plan, or a wrong side or drug name), minor otherwise. The grades were unblinded only after every recording was graded. The full rubric is in our methodology.
Transcript errors
Each transcript was aligned word by word with the verbatim reference. Differences in filler words were set aside, and Claude judged each remaining difference blind, using the rule above for what makes a transcript error significant. Word Error Rate is calculated by code with the same strict normalisation as our own release reports, treating equivalent written forms as the same words: 22 mm and 22 millimetres, US and UK spellings, Chinese and Arabic numerals, and Simplified and Traditional characters. Chinese is scored per character.