Hanah ← Trust Center United States · Validation transparency
Validation transparency

AI scribe accuracy benchmark — United States

These results were graded on the Australian deployment, which runs the same speech recognition and AI drafting stack as the United States. The methodology behind these results, including how each number was measured and its limitations, is in the methodology document.

Headline metrics

Coverage breakdown

Jury composition

Each detail is graded by every juror configured in the panel. Verdicts disagree on hard cases — that's the signal. Per-juror reasoning is shown verbatim below.

Sessions

Click any recording to drill in: source audio, ground-truth transcript, what the STT pipeline heard, every generated document, and every per-LLM verdict with its reasoning.

Methodology & provenance

Methodology

Word Error Rate, panel-of-LLM jury, prompt-scoped expectations, calibration sanity-check, and corpus graduation roadmap (synthetic → role-play → real-world).

Read the methodology document →
Source data

The full machine-readable artefact this page renders from. Every number on this page is computable directly from this file.

Open raw data →
Reproducibility

Corpus, juror prompts, and judge model versions are all version-controlled. Anyone can re-run the harness against the same commit and reproduce these numbers.

Trust Center home →