VALIDATION TRANSPARENCY
What changed.
Our previously published benchmark against the one we publish today, on the same corpus, in this region. Both reports remain available in full — the numbers below are computed from them, not typed in.
The grading panel changed too.
Our earlier run was graded by a smaller, cheaper pair of jury models. We found they
made avoidable mistakes — miscounting sentences, misreading document structure — so
the current run is graded by stronger models. We have deliberately not
re-graded the older run to match. Its numbers stand as they were measured on the day.
That means part of the coverage movement below reflects better grading rather than a
better product, and we would rather say so than quietly restate history.
What we changed
- Recording pipeline. Fixed voice-activity gating and a decoder lag that silently dropped the end of longer recordings. This accounts for most of the word error rate improvement.
- Ground truth. Corrected reference transcripts where they disagreed with the audio on rendering rather than on words — spoken numbers written out, regional spellings, unit shorthand. Genuine mis-hearings were left in place; those are what the benchmark is for.
- Post-recording correction. A second pass now re-listens to the audio and repairs transcription errors before the note is written. Previously this was failing silently on most recordings.
- Note-generation model. Updated to a newer model, then measured rather than assumed. It reduced invented content; it did not meaningfully move transcription accuracy, because it is not what transcribes.
Why we publish this
Model providers change what sits behind an API without announcing it, and a vendor claiming an upgrade "will be better" is not evidence. We run our own benchmark, per region, through the same browser path a clinician uses, and publish what we saw on that date. If a future run regresses, that will be published too — either we can explain it, or we have found something to fix.