---
title: "Deploying AI in your practice: six questions to ask before you sign — Hanah"
description: "Every AI scribe demo looks the same. The differences turn up later — in the invoice, the subprocessor list, and the first consult you record in a room with no signal. Six questions to put to every vendor, including us, with our own answers written down."
source: https://hanah.health/blog/deploying-ai-in-your-practice/
---

[← All posts](https://hanah.health/blog/) Article

# Deploying AI in your practice: six questions to ask before you sign

Ahmed, Founder, Hanah · 3 September 2026

AI medical scribes are conceptually simple and without going into the weeds of accuracy, deployment and training can be built relatively easily, making them all look the same.

The issues show up later, when it struggles to cope with multilingual conversation, speaker seperation, hidden charges for basic features like software integration, data transparency and even robust offline modes - when you lose your signal just as you start a house visit or the wifi goes out in your clinic.

Hanah is built for the real world, and these are the questions we think you should ask every AI scribe on the market, that we answer forthright.

## 1\. What is the real per-clinician price, with everything you need switched on?

The headline number on a pricing page is rarely what a practice pays, because the parts that make the product usable are often sold separately. The one to watch is the connection to your practice management system. A scribe that cannot file into your PMS is a second tab and a copy-paste habit, so if the integration sits behind a higher tier or an add-on line, the advertised price was never the price.

Check the pricing page, make sure you're getting everything you need for true workflow improvements, otherwise you replace documentation with juggling browser tabs.

We don't beleive it's ethical to upcharge for bread and butter features like PMS integerations, our [pricing](https://hanah.health/pricing/) includes a free Basic plan with the patient portal, exercise prescriptions and outcome measures, no card required. Practice at $100 per practitioner per month in AUD, $80 billed annually, with a one-month free trial — that tier includes the scribe, document generation, and the browser extension and API integerations into leading practice management software (contact us if we're missing yorus!).

## 2\. Where does the data actually go — all of it?

"Data sovereignty" is usually sold as a flag on a map, and a flag is not an answer. An ambient AI product is a chain of companies: the app, the database, the object storage, the backups, the speech-to-text engine, the language model, the email sender, the edge in front of it all.

Ask for the trust center, every subprocessor (third party the software relies on), what it does, what categories of data reach it, and the country it processes in. Then ask a second question that separates the careful vendors from the confident ones: _which link in that chain leaves the country, and what is done about it?_ Almost every chain has one. A vendor whose answer is a flat "everything stays onshore" deserves to be challenged.

Each Hanah region — [Australia](https://hanah.health/au/trust/), [the United Kingdom](https://hanah.health/uk/trust/), [New Zealand](https://hanah.health/nz/trust/) — runs on its own servers, with its own database, its own object storage, its own backups and its own speech and language stack. Nothing is pooled across regions. Each one publishes its own subprocessor table in its own [Trust Center](https://hanah.health/trust/), because the answers genuinely differ by region.

And here is our awkward link, published on our own trust page rather than left for you to discover in a security review: in the Australian deployment, streaming speech-to-text is processed in the United States. The audio is transient and is not retained once the transcript is produced, our terms prohibit training on it, and the resulting transcript is stored in Australia — but it is a cross-border disclosure under Australian Privacy Principle 8, and your privacy impact assessment needs to say so. We would rather you read that from us on a Tuesday than from a due-diligence questionnaire the week before you sign.

> Ask every vendor for the subprocessor table. The ones who have it ready are the ones who have thought about it.

## 3\. Can you reproduce their accuracy claim?

Everyone in this market claims to be accurate. Almost nobody publishes a number you could check, and a number you cannot reproduce is marketing wearing a lab coat — [we wrote about that in July](https://hanah.health/blog/ai-health-tech-transparency/), including the part where our own product got three things wrong.

A satisfaction statistic is not an accuracy measurement. "97% of users rate the notes as excellent" tells you how a survey went, not how often the transcript was wrong. When a vendor quotes you a percentage, ask four things: measured against what corpus, scored by what method, graded by whom, and where is the previous run.

That last one matters most. Any vendor can publish a good quarter. Publishing the bad one next to it is the part that makes the number mean something. Our [methodology](https://hanah.health/au/trust/wer-methodology/), our [current results](https://hanah.health/au/trust/transparency/) — every transcript, every generated note, every grader's reasoning — and [every past run](https://hanah.health/au/trust/benchmarks/) are all published, left exactly as they were measured on the day.

## 4\. Which kind of wrong does their model prefer to be?

This is the question almost nobody asks, and in a clinical record it is the one that matters.

There are two ways an AI note can be wrong. It can miss something that was said, or it can state something that was never said. Both are errors; they are not the same error. A missing detail is visible to the clinician reading the draft against their own memory of the consult. An invented detail — a laterality, a dose, a denial of symptoms that never happened — reads as perfectly plausible, which is exactly what makes it dangerous once it is signed and sitting in the record.

So we measure both, on every grading pass. Word error rate against a verbatim ground-truth transcript, and a panel of language models grading the generated note twice over: coverage (did it capture each fact registered in advance?) and hallucination (does it contain a claim the transcript does not support?).

Measuring both forces you to choose between them, and we have. When we [changed models in July](https://hanah.health/blog/benchmark-update-july-2026/), the incoming model captured slightly fewer details than the one it replaced and invented far fewer. We took that trade deliberately: word error rate fell from 5.9% to 1.6%, confirmed invented details went from three to none, and clinical facts captured sits at 94.7%. In a clinical context, completeness should not come at the cost of accuracy. A clinician can add the detail the model left out. They cannot reliably un-read one it made up.

Ask your vendor which way they made that trade, and whether they measured it at all.

## 5\. What happens when the consult isn't in English?

Australian and British practices are multilingual and their software mostly is not. In our own benchmark corpus, two of the recordings are a community nurse visiting a Vietnamese-speaking household and an occupational therapist in a Hindi-speaking one — not because it makes the numbers look good, but because a corpus of confident Anglophone speakers in quiet rooms would tell us nothing about a Tuesday in the northern suburbs.

Three things matter here, and they come apart.

-   **The consult.** Hanah identifies the spoken language automatically rather than being told in advance, so a session that switches between languages mid-sentence is handled rather than fought. We learned that the hard way: an earlier build carried a fixed list of expected languages, which quietly biased every session in every region and, on one accented English utterance, produced an opening line of hallucinated Arabic. We removed the hints. Detection alone is better than a guess applied to everybody.
-   **The record.** The clinical note is written in English, because that is the record's language and the medico-legal artefact your college and your insurer expect.
-   **The patient's copy.** This is the part that changes outcomes. A patient can set a preferred language — 38 of them — and their care plans and documents are translated into it. The original is never altered; a separate translated copy is generated and kept in step automatically whenever the clinician edits the plan, so the version the patient reads at home never quietly drifts from the version the clinician wrote. Clinicians also list the languages they speak on their profile, so a practice can match a patient to a person rather than to a translation.

One deliberate gap, since this is a post about asking hard questions. Auslan and BSL appear in our language list for clinicians but are excluded from the patient translation feature, because a text translation pass cannot serve a signed language and pretending otherwise would be worse than saying so.

## 6\. What happens when the wifi drops mid-consult?

Every scribe demo happens on good wifi. Real practice includes basements, home visits, rural clinics, hospital corridors, and the one treatment room at the back where the signal has never worked. If the answer to this question is a shrug, the product is a demo.

The failure everyone imagines is a spinner. The failure that actually costs you is silent: the recording appeared to work, the clinician moved on to the next patient, and the audio is gone. There is no recovering a consult that has already happened.

Hanah is built the other way around. Every recording is written to local storage on the device as it happens — losslessly compressed, not only when the connection is down but always, so the local copy exists before it is ever needed. When the connection is healthy, segments transcribe live and upload as they close. When it drops, recording simply continues into local storage, and the segments captured during the outage are marked for catch-up.

When the network returns, those segments upload and transcribe, and the transcript fills in. Uploads retry on a widening backoff and _never_ give up, because dropping audio on a guess is the one unrecoverable outcome. If the tab crashes or the browser is closed mid-consult, the next session finds the orphaned recording and rescues it rather than treating it as rubbish left behind. In our own testing, a fifty-minute consult recorded through a full connection dropout had all eighty of its audio chunks uploaded and transcribed thirteen seconds after the clinician pressed stop.

Ask any vendor you are evaluating to record a consult in aeroplane mode and show you the note afterwards. It is a two-minute test and it is very hard to fake.

## The short version

Ask

What a good answer looks like

Price with everything switched on

One number per clinician, PMS connection included, no add-on line for the integration

The subprocessor list

A table: who, what for, which data, which country — including the links that leave the country

The accuracy claim

A published method, a named corpus, and the previous run still visible next to this one

Missed vs invented detail

Both measured separately, and a stated preference between them

Non-English consults

Automatic language detection, an English record, and patient material in the patient's language

Aeroplane-mode test

The consult records, and the note arrives when the signal does

None of these questions are unfair, and none of them are about us. Put them to every vendor on your shortlist, put them to us, and give more weight to the ones who answer in specifics than the ones who answer in adjectives. Then run a real trial, on real consults, before you commit.

If you want our answers argued out in more detail — or you think one of them is wrong — write to me at hello@hanah.health.

See Hanah in your own practice.

[Book a demo](https://hanah.health/index.html#demo)
