Updating to new model and transcription algorithm.
We've updated the model that drafts your notes, along with the algorithm that records and transcribes a consult. This post explains what changed, what to expect, and what stays the same.
What our benchmarking shows
Measured against our Australian corpus, comparing this run with the one we published on 29 June:
Word error rate
1.6% −4.3pt vs 29 JuneInvented details, confirmed
1 −4 vs 29 JuneClinical facts captured
94.7% +3.2pt vs 29 JuneView data and results in more depth.
In our testing, the new Gemini Flash 3.6 model captured more clinical facts and invented far fewer details. In a clinical context completeness should never come at the cost of accuracy, and with this update we didn't have to trade one for the other.
Templates Remain the Same
Your templates are unaffected. The same templates should produce the same structure, the same sections, and the same fields. Nothing you have set up needs revisiting, and no action is required from you.
What changed
We updated both the note generation model from Gemini Pro 3.1 to Gemini Flash 3.6 and updated our transcription algorithm.
The transcription algorithm. Reworked how speech is detected and how the decoder keeps pace with a long consult, so more of what is said reaches the transcript — particularly at the start of a sentence and through the end of a longer recording. We also improved name detection, and updated our benchmarking corpus to include lesser known names.
The note-generation model. Updated to Gemini Flash 3.6. Combined with the improved transcription accuracy this led to fewer hallucinations and tightened language around note generation.
The numbers, and how to check them
Both runs are published in full: the current report, the side-by-side, and past results. You can read every transcript, every generated note, and every grader's reasoning.
Important note: We upgraded the LLM Jury to use larger models for increased confidence in our benchmarking. We found running smaller models multiple times often produced worse benchmarking reports than large models once.
Why we publish this at all
We believe software for healthcare providers should not be treated like SaaS 'move fast and break things'. We are brutally transparent and share any changes we make or plan to make to ensure you are adequately informed about how we process data.
The results from these updates are overall quite positive, however as we wrote at the start of this month, the number only means something if we also publish when they are bad and that transparency is our commitment to you.
If you feel we have regressed, or believe our benchmarking corpus is not reflective of your practice please don't hesitate to reach out at hello@hanah.health.