Benchmark

AI scribe benchmark: Hanah and six other AI scribes

We played the same 13 recorded consultations into Hanah, Heidi, Lyrebird, Freed, PatientNotes, CliniScripts and Preve, then graded every note against what was actually said. Every transcript, note and grade is published on this page, with the reason for each one.

100% of Hanah's notes free of major hallucination flags The average across the other six scribes was 48%.
66% fewer clinically significant transcript errors 0.59 per consultation against an average of 1.75.
11% more required clinical details captured 93.7% against an average of 84.2%.

Results

Scribes are listed by the share of required clinical details their notes captured. Each scribe's note for the first prompt of each recording is scored.

AI scribe Clinical details captured Notes with every clinical detail Major hallucination flags per note Notes with no major hallucination flag Significant transcript errors per consultation Word Error Rate, English consultations
Hanah13 notes 93.7% 62% 0.00 100% 0.59 1.6%
Heidi13 notes 91.4% 38% 0.23 77% 1.15 3.3%
CliniScripts13 notes 90.8% 62% 0.23 77% 0.69 2.1%
Lyrebird13 notes 90.2% 46% 0.62 62% 1.92 3.4%
Freed†12 notes 85.7% 42% 0.83 33% 1.25 19.4%
PatientNotes13 notes 74.7% 15% 2.77 8% 1.92 3.9%
Preve*10 notes 72.2% 30% 1.10 30% 3.54 6.7%
Average of the other six 84.2% 39% 0.96 48% 1.75 6.4%

* Preve was scored on its standard clinical note, because it follows a custom instruction only after a sample PDF is uploaded to build a template. Preve doesn't generate a note from less than 600 characters of transcript, and its three shortest recordings have no note.

† Freed kept recording after Stop on the shortest recording and produced no transcript or note, on two attempts. Freed is scored on the other 12.

A significant transcript error is one a blind rater judged could mislead a clinician: a wrong drug, dose, number, side or symptom, a lost negation or red flag, or clinical content dropped. Word Error Rate counts every word the same, fillers included.

The jaw surgery review was spoken mostly in Mandarin. Each scribe ran it with its input language set to Mandarin where it offers that setting. Heidi made the fewest significant transcript errors on it (3 in each of three runs), then CliniScripts (4) and Hanah (4 to 5 across three runs). PatientNotes made 9, Freed 7, Lyrebird 14 and Preve 24. Preve has no input language setting, and Freed detects the language itself.

By consultation

Clinical details captured, with the number of major hallucination flags underneath. Select a consultation to read every note and grade.

ConsultationHanahHeidiCliniScriptsLyrebirdFreedPatientNotesPreve
ICU handover dictation 100%no major 100%no major 100%no major 100%no major 100%2 major 75%2 major no note
ICU handover dictation, ward noise 100%no major 100%no major 100%no major 92%no major 75%1 major 75%2 major no note
Community nurse home visit 100%no major 100%no major 79%no major 100%1 major 100%1 major 64%1 major 86%1 major
Dietitian, IBS review 92%no major 92%no major 100%no major 92%no major 100%no major 75%1 major 92%1 major
Knee pain, short consult 83%no major 83%no major 83%no major 83%no major no note 100%no major no note
Knee replacement prehab 100%no major 88%no major 100%no major 100%1 major 100%1 major 100%2 major 100%1 major
Myotherapy, 50 minutes in a noisy room 100%no major 88%no major 56%no major 100%no major 81%1 major 94%2 major 100%no major
OT home safety assessment 100%no major 100%1 major 100%no major 93%no major 93%1 major 93%3 major 93%4 major
Child with asthma, parent with limited English 100%no major 86%no major 100%no major 100%no major 100%2 major 93%4 major 71%1 major
Jaw surgery review in Mandarin and English 81%no major 88%no major 88%1 major 69%3 major 62%no major 47%9 major 25%1 major
Quad bike fall, outdoors 86%no major 86%1 major 100%1 major 93%2 major 93%1 major 64%5 major 64%2 major
Rugby knee prehab 83%no major 100%no major 100%no major 100%1 major 67%no major 50%3 major 100%no major
Rugby hamstring, sideline 100%no major 86%1 major 93%1 major 86%no major 86%no major 86%2 major 79%no major

Every note, transcript and grade

Choose a consultation and a scribe. You can play the recording, read the reference transcript, and see each checklist verdict and hallucination flag with the reason it was given.

Loading…

How it was run

The recordings

The 13 recordings are Hanah's accuracy corpus, the same set behind the release results on our transparency page. They cover an ICU handover dictation with and without ward noise, physiotherapy, dietetics, occupational therapy, myotherapy and community nursing consults, a child with asthma whose parent speaks limited English, a quad bike fall recorded outdoors, a 50 minute myotherapy session with about seven minutes of speech in a noisy treatment room, and a jaw surgery review spoken mostly in Mandarin. Each has a verbatim reference transcript and a checklist of the details a note must capture.

Capture

Each scribe ran in its own web app in Chrome, signed in to a standard account. The recording was played through the browser's microphone input, the same file for every scribe. After recording, each recording's prompts were typed into the scribe's own document or chat feature word for word, for example "Generate SOAP notes please" and "Summarise this consult in two sentences for a referral letter". Each scribe was run once per recording on 6 October 2026, with English as the input language where a scribe asks for one. For the Mandarin review, scribes with an input language setting were set to Mandarin: Heidi to English and Chinese (Mandarin), Lyrebird and PatientNotes to Chinese (Mandarin, Simplified). CliniScripts produced no transcript with Chinese and English selected on two attempts, so its standard setting is used. Heidi and Hanah ran the Mandarin review three times each, and their transcript results are averaged.

Grading by a Claude jury

The notes were graded by Claude, Anthropic's model. For each recording, every scribe's output was shuffled behind a letter with product names removed, and graded against the reference transcript with a fixed rubric. Each checklist item is marked yes, partial or no. Every statement the transcript doesn't support is listed as a hallucination flag: major when it could change care or misstate the clinical picture (an invented symptom, finding, drug, diagnosis or plan, or a wrong side or drug name), minor otherwise. The grades were unblinded only after every recording was graded. The full rubric is in our methodology.

Transcript errors

Each transcript was aligned word by word with the verbatim reference. Differences in filler words were set aside, and Claude judged each remaining difference blind, using the rule above for what makes a transcript error significant. Word Error Rate is calculated by code with the same strict normalisation as our own release reports, treating equivalent written forms as the same words: 22 mm and 22 millimetres, US and UK spellings, Chinese and Arabic numerals, and Simplified and Traditional characters. Chinese is scored per character.