---
title: "AI scribe benchmark: Hanah and six other AI scribes"
description: "13 recorded consultations played into Hanah, Heidi, Lyrebird, Freed, PatientNotes, CliniScripts and Preve. Every transcript, note and grade is published, with the reason for each grade."
source: https://hanah.health/benchmark/
---

Benchmark

# AI scribe benchmark: Hanah and six other AI scribes

Run on 6 October 2026 · 13 recorded consultations · 7 AI scribes

We played the same 13 recorded consultations into Hanah, Heidi, Lyrebird, Freed, PatientNotes, CliniScripts and Preve, then graded every note against what was actually said. Every transcript, note and grade is published on this page, with the reason for each one.

100% of Hanah's notes free of major hallucination flags The average across the other six scribes was 48%.

66% fewer clinically significant transcript errors 0.59 per consultation against an average of 1.75.

11% more required clinical details captured 93.7% against an average of 84.2%.

## Results

Scribes are listed by the share of required clinical details their notes captured. Each scribe's note for the first prompt of each recording is scored.

| AI scribe | Clinical details captured | Notes with every clinical detail | Major hallucination flags per note | Notes with no major hallucination flag | Significant transcript errors per consultation | Word Error Rate, English consultations |
| --- | --- | --- | --- | --- | --- | --- |
| Hanah13 notes | 93.7% | 62% | 0.00 | 100% | 0.59 | 1.6% |
| Heidi13 notes | 91.4% | 38% | 0.23 | 77% | 1.15 | 3.3% |
| CliniScripts13 notes | 90.8% | 62% | 0.23 | 77% | 0.69 | 2.1% |
| Lyrebird13 notes | 90.2% | 46% | 0.62 | 62% | 1.92 | 3.4% |
| Freed†12 notes | 85.7% | 42% | 0.83 | 33% | 1.25 | 19.4% |
| PatientNotes13 notes | 74.7% | 15% | 2.77 | 8% | 1.92 | 3.9% |
| Preve*10 notes | 72.2% | 30% | 1.10 | 30% | 3.54 | 6.7% |
| Average of the other six | 84.2% | 39% | 0.96 | 48% | 1.75 | 6.4% |

\* Preve was scored on its standard clinical note, because it follows a custom instruction only after a sample PDF is uploaded to build a template. Preve doesn't generate a note from less than 600 characters of transcript, and its three shortest recordings have no note.

† Freed kept recording after Stop on the shortest recording and produced no transcript or note, on two attempts. Freed is scored on the other 12.

A significant transcript error is one a blind rater judged could mislead a clinician: a wrong drug, dose, number, side or symptom, a lost negation or red flag, or clinical content dropped. Word Error Rate counts every word the same, fillers included.

The jaw surgery review was spoken mostly in Mandarin. Each scribe ran it with its input language set to Mandarin where it offers that setting. Heidi made the fewest significant transcript errors on it (3 in each of three runs), then CliniScripts (4) and Hanah (4 to 5 across three runs). PatientNotes made 9, Freed 7, Lyrebird 14 and Preve 24. Preve has no input language setting, and Freed detects the language itself.

## By consultation

Clinical details captured, with the number of major hallucination flags underneath. Select a consultation to read every note and grade.

| Consultation | Hanah | Heidi | CliniScripts | Lyrebird | Freed | PatientNotes | Preve |
| --- | --- | --- | --- | --- | --- | --- | --- |
| ICU handover dictation | 100%no major | 100%no major | 100%no major | 100%no major | 100%2 major | 75%2 major | no note |
| ICU handover dictation, ward noise | 100%no major | 100%no major | 100%no major | 92%no major | 75%1 major | 75%2 major | no note |
| Community nurse home visit | 100%no major | 100%no major | 79%no major | 100%1 major | 100%1 major | 64%1 major | 86%1 major |
| Dietitian, IBS review | 92%no major | 92%no major | 100%no major | 92%no major | 100%no major | 75%1 major | 92%1 major |
| Knee pain, short consult | 83%no major | 83%no major | 83%no major | 83%no major | no note | 100%no major | no note |
| Knee replacement prehab | 100%no major | 88%no major | 100%no major | 100%1 major | 100%1 major | 100%2 major | 100%1 major |
| Myotherapy, 50 minutes in a noisy room | 100%no major | 88%no major | 56%no major | 100%no major | 81%1 major | 94%2 major | 100%no major |
| OT home safety assessment | 100%no major | 100%1 major | 100%no major | 93%no major | 93%1 major | 93%3 major | 93%4 major |
| Child with asthma, parent with limited English | 100%no major | 86%no major | 100%no major | 100%no major | 100%2 major | 93%4 major | 71%1 major |
| Jaw surgery review in Mandarin and English | 81%no major | 88%no major | 88%1 major | 69%3 major | 62%no major | 47%9 major | 25%1 major |
| Quad bike fall, outdoors | 86%no major | 86%1 major | 100%1 major | 93%2 major | 93%1 major | 64%5 major | 64%2 major |
| Rugby knee prehab | 83%no major | 100%no major | 100%no major | 100%1 major | 67%no major | 50%3 major | 100%no major |
| Rugby hamstring, sideline | 100%no major | 86%1 major | 93%1 major | 86%no major | 86%no major | 86%2 major | 79%no major |

## Every note, transcript and grade

Choose a consultation and a scribe. You can play the recording, read the reference transcript, and see each checklist verdict and hallucination flag with the reason it was given.

Consultation ICU handover dictation ICU handover dictation, ward noise Community nurse home visit Dietitian, IBS review Knee pain, short consult Knee replacement prehab Myotherapy, 50 minutes in a noisy room OT home safety assessment Child with asthma, parent with limited English Jaw surgery review in Mandarin and English Quad bike fall, outdoors Rugby knee prehab Rugby hamstring, sideline

Loading…

## How it was run

### The recordings

The 13 recordings are Hanah's accuracy corpus, the same set behind the release results on our [transparency page](https://hanah.health/au/trust/transparency/). They cover an ICU handover dictation with and without ward noise, physiotherapy, dietetics, occupational therapy, myotherapy and community nursing consults, a child with asthma whose parent speaks limited English, a quad bike fall recorded outdoors, a 50 minute myotherapy session with about seven minutes of speech in a noisy treatment room, and a jaw surgery review spoken mostly in Mandarin. Each has a verbatim reference transcript and a checklist of the details a note must capture.

### Capture

Each scribe ran in its own web app in Chrome, signed in to a standard account. The recording was played through the browser's microphone input, the same file for every scribe. After recording, each recording's prompts were typed into the scribe's own document or chat feature word for word, for example "Generate SOAP notes please" and "Summarise this consult in two sentences for a referral letter". Each scribe was run once per recording on 6 October 2026, with English as the input language where a scribe asks for one. For the Mandarin review, scribes with an input language setting were set to Mandarin: Heidi to English and Chinese (Mandarin), Lyrebird and PatientNotes to Chinese (Mandarin, Simplified). CliniScripts produced no transcript with Chinese and English selected on two attempts, so its standard setting is used. Heidi and Hanah ran the Mandarin review three times each, and their transcript results are averaged.

### Grading by a Claude jury

The notes were graded by Claude, Anthropic's model. For each recording, every scribe's output was shuffled behind a letter with product names removed, and graded against the reference transcript with a fixed rubric. Each checklist item is marked yes, partial or no. Every statement the transcript doesn't support is listed as a hallucination flag: major when it could change care or misstate the clinical picture (an invented symptom, finding, drug, diagnosis or plan, or a wrong side or drug name), minor otherwise. The grades were unblinded only after every recording was graded. The full rubric is in our [methodology](https://hanah.health/au/trust/wer-methodology/#mth-comparative).

### Transcript errors

Each transcript was aligned word by word with the verbatim reference. Differences in filler words were set aside, and Claude judged each remaining difference blind, using the rule above for what makes a transcript error significant. Word Error Rate is calculated by code with the same strict normalisation as our own [release reports](https://hanah.health/au/trust/wer-methodology/#mth-layer-a), treating equivalent written forms as the same words: 22 mm and 22 millimetres, US and UK spellings, Chinese and Arabic numerals, and Simplified and Traditional characters. Chinese is scored per character.
