Vietnamese speech to text for hospitals and clinics
01/10/2026

The clinic is packed. The doctor has seen twenty-odd patients since morning, and his eyes are still fixed on the screen as he types up the record. The patient across the desk watches the keyboard more than the doctor's face. Anyone who has been to a clinic knows the feeling. Vietnamese speech to text exists partly to fix exactly this: the doctor speaks, the machine writes, and people get to look at each other again.
This is the direction AIVISION is working in with its E1.0 speech to text model. But dictating clinical notes is much harder than transcribing a speech, and I want to be clear about where the difficulty lies.
Why hospital speech is its own problem
A recognition model that works well on news, podcasts or customer service calls will not necessarily work in a hospital corridor. The reasons are very practical.
- Background noise: machines, queue announcements, relatives talking, air conditioning humming.
- Specialised vocabulary: medical terms, drug names, anatomical names, with Vietnamese, English and Latin mixed in a single sentence.
- Regional accents: doctors and patients come from everywhere, and Northern, Central and Southern pronunciations differ, along with local speech that never appears in textbooks.
- Clipped, fast speech: a doctor on call does not dictate like a school exercise. They mutter, as if thinking aloud.
Numbers cause trouble too. A measurement with one digit misheard can mean something entirely different, which is exactly why the output cannot go straight into a record without a human looking at it.
From raw text to a structured note
Transcription is only the first step. A clinical note has a structure: reason for visit, history, examination, plan. What a doctor says does not follow that mould. This is where a language model comes in. Once speech becomes text, an LLM can rearrange it into the sections the facility uses, fix the spelling of terms, and flag passages that were unclear so the doctor can look again.
We find it sensible to separate two layers. The first transcribes what was said, faithfully. The second reorganises it. With the split, when there is an error you know where it lives: the ear heard wrongly or the hand arranged wrongly. Merged into one block, tracing errors is miserable.
Another principle: where the model is unsure, it should show it. A word underlined or marked for checking is more useful than a beautiful transcript that is quietly wrong. False polish is the enemy of correctness.
A human still signs
I want to say this plainly, because it shapes the whole design. This tool supports note-taking. It does not replace the doctor in writing the record. The doctor reads, edits and signs. The model does not diagnose, does not suggest medication, and does not add to the record anything the doctor did not say. If a product version starts filling in blanks on its own, it is heading the wrong way.
This is also where interface design matters more than people expect. If the transcript looks too complete, a tired doctor will be tempted to sign without checking. Some product teams show two columns, the original speech beside the organised note, to make comparison easy. I am not saying it is the only way, but it reflects something true: a good tool is not only accurate, it also makes checking easy.
Deployment: the part outside the model
Most of the time in a project like this goes not into the model but into everything around it.
Microphones and environment
Audio quality matters enormously. A lapel mic or a conference mic placed well can help more than another round of training. Test in a real clinic, not a quiet meeting room.
Patient consent
Recording a consultation is sensitive. You need clear notice, consent, and a rule for how long the original audio is kept, or whether it is kept at all. Many organisations process the recording and delete it, keeping only the text a doctor has approved.
Fitting the existing workflow
If the doctor has to open a separate app and copy and paste, the tool will be abandoned within two weeks. The text has to flow to the place where they already work.
On capability, I will only say what is known
AIVISION has released the E1.0 speech to text model for Vietnamese and trains on a cluster of 24 NVIDIA H200 GPUs and 8 NVIDIA B300 GPUs. We have not published an error rate for any hospital environment, because that figure depends on the clinic, the microphone, the team's accents and the specialty vocabulary, and a generic number would only mislead. What we are willing to do is measure on the organisation's real data, with appropriate consent, and look together at where it goes wrong before talking about wider rollout.
Fine-tuning per specialty matters as well. A cardiology ward, a paediatric clinic and a laboratory use different vocabularies. Teaching the model a department's own words is usually a reasonable step, as long as the data is handled properly and someone evaluates the results.
Paperwork and what is worth keeping
People often say technology saves doctors time. I think a more accurate way to put it is that it can move time from the screen to the patient. If a consultation loses a few minutes of typing and the doctor looks at the patient a little longer, that is a small but real change. If the tool only adds a checking step without removing any, it has not earned its place in the clinic. The simplest test is still to ask the people who use it every day.