---
title: "AI scribe benchmark: Hanah vs Heidi, Lyrebird, Freed and more"
description: "We played the same 13 recorded consultations into seven AI scribes and graded every note and transcript blind. Hanah's notes had no major hallucination flags, captured 93.7% of the required clinical details against an average of 84.2%, and its transcripts had 66% fewer clinically significant errors. Every transcript, note and grade is published."
source: https://hanah.health/blog/hanah-vs-the-rest-benchmark/
---

[← All posts](https://hanah.health/blog/) Article

# Hanah vs The Rest: an objective, transparent benchmark

The Hanah team · 6 October 2026

![](https://hanah.health/assets/blog/hanah-vs-the-rest-benchmark.webp)

We played the same 13 recorded consultations into Hanah, Heidi, Lyrebird, Freed, PatientNotes, CliniScripts and Preve, then graded every note against what was actually said. The full results, with every transcript, note and grade, are on our [benchmark page](https://hanah.health/benchmark/).

100%of Hanah's notes free of major hallucination flagsHanah_100%_Average_48%_66% fewerclinically significant transcript errorsHanah_0.59_Average_1.75_11% morerequired clinical details capturedHanah_93.7%_Average_84.2%_6m 04ssaved per note against writing it yourselfHanah_6m 04s_Average_4m 09s_

## The results

| AI scribe | Clinical details captured | Major hallucination flags per note | Significant transcript errors per consultation |
| --- | --- | --- | --- |
| Hanah | 93.7% | 0.00 | 0.59 |
| Heidi | 91.4% −2.3% | 0.23 +0.23 | 1.15 +0.56 |
| CliniScripts | 90.8% −2.9% | 0.23 +0.23 | 0.69 +0.10 |
| Lyrebird | 90.2% −3.4% | 0.62 +0.62 | 1.92 +1.33 |
| Freed | 85.7% −8.0% | 0.83 +0.83 | 1.25 +0.66 |
| PatientNotes | 74.7% −19.0% | 2.77 +2.77 | 1.92 +1.33 |
| Preve | 72.2% −21.5% | 1.10 +1.10 | 3.54 +2.95 |

Hanah's notes captured 11% more of the required clinical details than the average AI scribe. None of Hanah's 13 notes had a major hallucination flag, and fewer than half of the other scribes' notes were free of one. Hanah's transcripts had 66% fewer clinically significant errors than the average, and 76% fewer on the English consultations. A clinically significant transcript error is one that could mislead a clinician: a wrong drug, dose, number, side or symptom, a lost negation or red flag, or clinical content dropped.

## What a major hallucination flag looks like

A major hallucination flag is something in the note that wasn't said and could change care. These are from the notes in this benchmark, and each one is published with the transcript it came from.

-   In an ICU handover, PatientNotes named dapagliflozin as the cause of the patient's ketoacidosis. The clinician said empagliflozin. In the same dictation recorded with ward noise, PatientNotes recorded a heart rate of 91. The 91 was the patient's oxygen saturation, which was left out of the note.
-   For a child with asthma, PatientNotes wrote betamethasone where the clinician said dexamethasone, and recorded yellow sputum when the parent described yellow vomit.
-   A patient who came off a quad bike became a bicycle fall in Heidi's summary and a motorcycle fall in Lyrebird's note. Lyrebird also put the injuries on the right side, which was never said.

Transcripts can go wrong as well. In a 50 minute myotherapy session with about seven minutes of speech in a noisy treatment room, Freed's transcript filled the quiet stretches with speech nobody said, including "Thank you for watching!" and a passage about the history of a town. Lyrebird's transcript of the same session included Zolmitriptan, a migraine drug nobody mentioned. Hanah's transcript of that recording had no clinically significant errors.

## Mandarin and English

One consultation was a jaw surgery review held mostly in Mandarin, with the clinical terms in English. Every scribe that offers an input language setting was set to Mandarin for it, and Heidi and Hanah each ran it three times.

Heidi transcribed it best, with 3 clinically significant errors in each of its three runs. CliniScripts made 4 and Hanah made 4 or 5. Hanah's were English clinical words spoken inside Mandarin, such as "pus" written as "pass" and chlorhexidine written as "hexadene". PatientNotes made 9, Freed 7 and Lyrebird 14. Preve has no input language setting and made 24. Hanah's note from the review still had no major hallucination flags.

## How we ran it

The recordings are Hanah's accuracy corpus, the same 13 recordings behind the release results on our [transparency page](https://hanah.health/au/trust/transparency/). They include dictations, physiotherapy, dietetics, occupational therapy, myotherapy and community nursing consults, a child whose parent speaks limited English, an outdoor recording and the Mandarin review. Each has a verbatim reference transcript and a checklist of what the note must capture.

Each scribe ran in its own web app with a standard account and default settings. The same audio file went into every scribe, and the same prompts were typed in word for word. Each scribe ran once per recording, on 6 October 2026, apart from the Mandarin repeats above.

The notes were graded by Claude, Anthropic's model. For each recording, every scribe's notes were shuffled behind letters with the product names removed, then graded against the reference transcript using a fixed rubric. The letters were matched back to products only after all 13 recordings were graded. Each transcript was compared with the reference word by word, and Claude judged every difference blind to decide whether it was clinically significant. The rubric and the rules behind every grade are in our [methodology](https://hanah.health/au/trust/wer-methodology/#mth-comparative).

## Check it yourself

Every grade on the [benchmark page](https://hanah.health/benchmark/#evidence) comes with its reason. You can play each recording, read the reference transcript, and compare it with what each scribe wrote.

See Hanah in your own practice.

[Book a demo](https://hanah.health/index.html#demo)
