Medical terms and drug names in Vietnamese: why LLMs slip

01/10/2026

Medical terms and drug names in Vietnamese: why LLMs slip

Try a small experiment. Give a general-purpose model a hastily written Vietnamese clinical note, the kind with abbreviations and Latin words dropped in, and ask it to rewrite the note clearly. Most of the time the result reads smoothly. But if you work in medicine and read closely, you will sometimes find an abbreviation read with the wrong meaning, or a drug name swapped for one that sounds similar. This is what people working on Vietnamese medical terminology with AI run into, and it is the reason fine-tuning a medical LLM is on the table.

This article explains why the errors happen, how far fine-tuning can go, and where no amount of training will help. The viewpoint is that of AIVISION, which trains and fine-tunes Vietnamese language models.

Why general models slip

There is no single cause. Several layers stack up.

Thin training data exactly where it needs to be thick

A model learns from the text it has read. English medical writing has an enormous body of material. Vietnamese medical writing has far less, and much of it sits in formats that are hard to use: scanned PDFs, internal documents, records that are not public for good reasons. The result is that a model understands English terms better than Vietnamese ones, and when it meets a Vietnamised name, it guesses.

One concept, many names

The same thing may be called by a Sino-Vietnamese term, a Latin name, an English name, a colloquial name, and a department's own abbreviation. A practitioner understands from context. The model often lacks that context. A two-letter abbreviation can mean different things in cardiology and in paediatrics.

Drug names are a special case

Drug names have two tiers: the active ingredient and the brand name, and brand names are chosen by each company, sometimes differing by only a letter or two. They are transcribed into Vietnamese in several ways, and users type them without diacritics, mistype them, or hear them over the phone and write them down. For a statistical model, two names that look alike on the page are easily merged into one. That error is dangerous because the two products may be used for entirely different purposes.

Carefree confidence

A model has no feeling of hesitation. When it meets a name it has never seen, it does not say it does not know. It produces a plausible answer. That is not bad design on purpose. It is a consequence of how the model learns: predict the most likely next word.

What fine-tuning helps with

Fine-tuning means further training on data from a narrow field. Done well, it brings a few meaningful changes.

  • The model gets used to how Vietnamese medical staff actually write, including abbreviations and mixed-language habits.
  • It recognises variants of a name better: with and without diacritics, transliterated, familiar misspellings.
  • It understands the structure of document types: clinical notes, prescriptions, labels, leaflets.
  • It keeps terminology intact when summarising or rewriting, instead of rephrasing it to sound nicer.

But I want to avoid the language of miracles. Fine-tuning does not turn a model into a doctor or a pharmacist. It makes the model better at the language of the field, reading and writing more precisely. Understanding terminology is not clinical judgment, and we do not use fine-tuning to cross that line.

Data decides most of the outcome

In projects like this, time spent on data usually exceeds time spent on training. You need a glossary reviewed by practitioners, with variants and abbreviations per department. You need real example sentences, processed so they reveal no personal information. You need easily confused name pairs labelled specifically so the model can learn to tell them apart.

What I find most effective, unglamorous as it is, is a glossary written by end users. A pharmacist knows which product customers call by which name. An internal medicine doctor knows three abbreviations for the same diagnosis. That knowledge is not in books, and no model will pick it up on its own.

Plan for change as well. New products appear, nomenclature is updated, departments add new names. The glossary should be a living document with an owner who maintains it, not something built once.

How to tell whether it improved

Reading something and feeling that it is smooth is not a measurement. A more trustworthy approach is a small but difficult test set: confusable name pairs, ambiguous abbreviations, sentences stripped of context. Run it before and after fine-tuning and have experts grade by hand. Record not just the rate of correct answers but the kinds of mistakes, because omitting something is very different from inventing something.

I will not give a figure for the improvement here, since it depends on the data and the specific problem, and a number borrowed from elsewhere only misleads. If someone offers you a precise percentage without saying which dataset it was measured on, ask.

Putting it in its place

AIVISION trains on a cluster of 24 NVIDIA H200 GPUs and 8 NVIDIA B300 GPUs, has released the L1.0 LLM for Vietnamese along with the E1.0 speech to text model, and takes on fine-tuning for fields such as healthcare and pharmacy. Our approach is to build a solid Vietnamese base and then tune it for each domain together with the people who will use it, with data handled properly.

Even so, there is one reminder I keep for everyone using tools like this: when a term or drug name appears in a model's output and it matters for a decision, check it against the source document. A better model makes fewer errors, it does not make zero. Doctors and pharmacists remain the final decision makers, and the tool should be designed to make their checking fast, not to make checking unnecessary.

Related insights

See all insights