---
title: "AI scribe accuracy benchmark, United States — Hanah"
description: "Hanah's AI scribe accuracy benchmark for the United States: 1.6% Word Error Rate over 11 recordings graded 30 July 2026 on the Australian deployment, which runs the same stack, against 113 pre-registered clinical facts."
source: https://hanah.health/us/trust/transparency/
---

Validation transparency

# AI scribe accuracy benchmark — United States

These results were graded on the Australian deployment, which runs the same speech recognition and AI drafting stack as the United States. The [methodology](https://hanah.health/us/trust/wer-methodology/) behind these results, including how each number was measured and its limitations, is in the methodology document.

## Headline metrics

## Coverage breakdown

## Jury composition

Each detail is graded by every juror configured in the panel. Verdicts disagree on hard cases — that's the signal. Per-juror reasoning is shown verbatim below.

## Sessions

Click any recording to drill in: source audio, ground-truth transcript, what the STT pipeline heard, every generated document, and every per-LLM verdict with its reasoning.

## Methodology & provenance

Methodology

Word Error Rate, panel-of-LLM jury, prompt-scoped expectations, calibration sanity-check, and corpus graduation roadmap (synthetic → role-play → real-world).

[Read the methodology document →](https://hanah.health/us/trust/wer-methodology/index.html)

Source data

The full machine-readable artefact this page renders from. Every number on this page is computable directly from this file.

[Open raw data →](https://hanah.health/us/trust/transparency/data.js)

Reproducibility

Corpus, juror prompts, and judge model versions are all version-controlled. Anyone can re-run the harness against the same commit and reproduce these numbers.

[Trust Center home →](https://hanah.health/us/trust/index.html)
