Open the box on accuracy.
Every recording in our corpus runs end-to-end through the production pipeline. We publish what was said, what we transcribed, what was generated, and the verdict each LLM juror reached on whether the document captured what mattered. Numbers and methodology, no PR gloss.
Headline metrics
Two layers of measurement. Layer A is mechanical — Word Error Rate against verbatim ground-truth. Layer B is qualitative — an LLM jury grades each clinically-important detail we declared up front (coverage: yes / partial / no), and separately enumerates every claim in the generated document that the transcript does not support (hallucinations: fabricated content the model invented).
Coverage breakdown
Jury composition
Each detail is graded by every juror configured in the panel. Verdicts disagree on hard cases — that's the signal. Per-juror reasoning is shown verbatim below.
Sessions
Click any recording to drill in: source audio, ground-truth transcript, what the STT pipeline heard, every generated document, and every per-LLM verdict with its reasoning.
Methodology & provenance
Word Error Rate, panel-of-LLM jury, prompt-scoped expectations, calibration sanity-check, and corpus graduation roadmap (synthetic → role-play → real-world).
Read the methodology document →The full machine-readable artefact this page renders from. Every number on this page is computable directly from this file.
Open raw data →Corpus, juror prompts, and judge model versions are all version-controlled. Anyone can re-run the harness against the same commit and reproduce these numbers.
Trust Center home →