← All posts Article

Hanah vs The Rest: an objective, transparent benchmark

We played the same 13 recorded consultations into Hanah, Heidi, Lyrebird, Freed, PatientNotes, CliniScripts and Preve, then graded every note against what was actually said. The full results, with every transcript, note and grade, are on our benchmark page.

The results

AI scribe Clinical details captured Major hallucination flags per note Significant transcript errors per consultation
Hanah 93.7% 0.00 0.59
Heidi 91.4% −2.3% 0.23 +0.23 1.15 +0.56
CliniScripts 90.8% −2.9% 0.23 +0.23 0.69 +0.10
Lyrebird 90.2% −3.4% 0.62 +0.62 1.92 +1.33
Freed 85.7% −8.0% 0.83 +0.83 1.25 +0.66
PatientNotes 74.7% −19.0% 2.77 +2.77 1.92 +1.33
Preve 72.2% −21.5% 1.10 +1.10 3.54 +2.95

Hanah's notes captured 11% more of the required clinical details than the average AI scribe. None of Hanah's 13 notes had a major hallucination flag, and fewer than half of the other scribes' notes were free of one. Hanah's transcripts had 66% fewer clinically significant errors than the average, and 76% fewer on the English consultations. A clinically significant transcript error is one that could mislead a clinician: a wrong drug, dose, number, side or symptom, a lost negation or red flag, or clinical content dropped.

What a major hallucination flag looks like

A major hallucination flag is something in the note that wasn't said and could change care. These are from the notes in this benchmark, and each one is published with the transcript it came from.

  • In an ICU handover, PatientNotes named dapagliflozin as the cause of the patient's ketoacidosis. The clinician said empagliflozin. In the same dictation recorded with ward noise, PatientNotes recorded a heart rate of 91. The 91 was the patient's oxygen saturation, which was left out of the note.
  • For a child with asthma, PatientNotes wrote betamethasone where the clinician said dexamethasone, and recorded yellow sputum when the parent described yellow vomit.
  • A patient who came off a quad bike became a bicycle fall in Heidi's summary and a motorcycle fall in Lyrebird's note. Lyrebird also put the injuries on the right side, which was never said.

Transcripts can go wrong as well. In a 50 minute myotherapy session with about seven minutes of speech in a noisy treatment room, Freed's transcript filled the quiet stretches with speech nobody said, including "Thank you for watching!" and a passage about the history of a town. Lyrebird's transcript of the same session included Zolmitriptan, a migraine drug nobody mentioned. Hanah's transcript of that recording had no clinically significant errors.

Mandarin and English

One consultation was a jaw surgery review held mostly in Mandarin, with the clinical terms in English. Every scribe that offers an input language setting was set to Mandarin for it, and Heidi and Hanah each ran it three times.

Heidi transcribed it best, with 3 clinically significant errors in each of its three runs. CliniScripts made 4 and Hanah made 4 or 5. Hanah's were English clinical words spoken inside Mandarin, such as "pus" written as "pass" and chlorhexidine written as "hexadene". PatientNotes made 9, Freed 7 and Lyrebird 14. Preve has no input language setting and made 24. Hanah's note from the review still had no major hallucination flags.

How we ran it

The recordings are Hanah's accuracy corpus, the same 13 recordings behind the release results on our transparency page. They include dictations, physiotherapy, dietetics, occupational therapy, myotherapy and community nursing consults, a child whose parent speaks limited English, an outdoor recording and the Mandarin review. Each has a verbatim reference transcript and a checklist of what the note must capture.

Each scribe ran in its own web app with a standard account and default settings. The same audio file went into every scribe, and the same prompts were typed in word for word. Each scribe ran once per recording, on 6 October 2026, apart from the Mandarin repeats above.

The notes were graded by Claude, Anthropic's model. For each recording, every scribe's notes were shuffled behind letters with the product names removed, then graded against the reference transcript using a fixed rubric. The letters were matched back to products only after all 13 recordings were graded. Each transcript was compared with the reference word by word, and Claude judged every difference blind to decide whether it was clinically significant. The rubric and the rules behind every grade are in our methodology.

Check it yourself

Every grade on the benchmark page comes with its reason. You can play each recording, read the reference transcript, and compare it with what each scribe wrote.

See Hanah in your own practice.

Book a demo