---
title: "What changed — Hanah Validation (AU)"
description: "How Hanah's accuracy benchmark moved between our last published run and the current one, and exactly what changed to cause it."
source: https://hanah.health/au/trust/comparison/
---

VALIDATION TRANSPARENCY

# What changed.

Our previously published benchmark against the one we publish today, on the same corpus, in this region. Both reports remain available in full — the numbers below are computed from them, not typed in.

**The grading panel changed too.** Our earlier run was graded by a smaller, cheaper pair of jury models. We found they made avoidable mistakes — miscounting sentences, misreading document structure — so the current run is graded by stronger models. We have deliberately _not_ re-graded the older run to match. Its numbers stand as they were measured on the day. That means part of the coverage movement below reflects better grading rather than a better product, and we would rather say so than quietly restate history.

## What we changed

-   **Recording pipeline.** Fixed voice-activity gating and a decoder lag that silently dropped the end of longer recordings. This accounts for most of the word error rate improvement.
-   **Ground truth.** Corrected reference transcripts where they disagreed with the audio on rendering rather than on words — spoken numbers written out, regional spellings, unit shorthand. Genuine mis-hearings were left in place; those are what the benchmark is for.
-   **Post-recording correction.** A second pass now re-listens to the audio and repairs transcription errors before the note is written. Previously this was failing silently on most recordings.
-   **Note-generation model.** Updated to a newer model, then measured rather than assumed. It reduced invented content; it did not meaningfully move transcription accuracy, because it is not what transcribes.

## Why we publish this

Model providers change what sits behind an API without announcing it, and a vendor claiming an upgrade "will be better" is not evidence. We run our own benchmark, per region, through the same browser path a clinician uses, and publish what we saw on that date. If a future run regresses, that will be published too — either we can explain it, or we have found something to fix.
