Why a Vietnamese LLM Needs Its Own Training

01/10/2026

Why a Vietnamese LLM Needs Its Own Training

A friend who runs customer support sent me a screenshot one afternoon. The customer had typed, with no diacritics at all, a quick question about a warranty. The chatbot replied with a polished, fluent paragraph that missed the point entirely, because it read one unaccented word as something else. Nobody on the team laughed. It is the kind of small failure you only see in production, and it is the main reason we at AIVISION take the idea of a dedicated Vietnamese LLM seriously.

To be fair: the large multilingual models are decent at Vietnamese today, and plenty of tasks do not need anything special. But 'decent' and 'reliable enough to put in front of customers' are far apart, and the gap shows up in four specific places.

The tokenizer: a cost hiding in plain sight

A model does not read characters, it reads tokens. The tokenizer is built from the original training corpus, and when that corpus is mostly English, English fragments get merged efficiently while Vietnamese gets chopped into small pieces. A Vietnamese sentence with full diacritics can cost noticeably more tokens than its English equivalent. We will not quote a number because it varies by tokenizer, but you can measure it yourself in five minutes: take the same paragraph in both languages and count.

The consequences stack up. API bills grow because pricing is per token. The context window effectively shrinks, so fewer pages fit. Generation slows down because each syllable may take several steps. Dedicated training lets you choose or extend the vocabulary so that common syllables are single tokens. The trade-off is real: you have to retrain the embeddings, and if you do it carelessly the model can forget some of what it already knew.

Data: more is not the same as cleaner

There is a lot of Vietnamese text online, but much of it is republished news, comment threads, ads, spam and raw machine translation. A model trained on that inherits the stiff rhythm of machine-translated prose. You have probably read a Vietnamese paragraph and known instantly it was translated from English: every word correct, the sentence something nobody would actually say. Models produce paragraphs like that because they have read a great many of them.

The bigger shortage is in specialised domains: administrative text, legal documents, clinical records, drug information. These are either not public or locked in poor-quality scanned PDFs. Getting a model to speak a field's language means somebody has to collect, clean and normalise that material. It is tedious work, and it matters more than the choice of architecture. We cover it separately in a post on training data.

Tone marks are not decoration

Vietnamese has six tones, and the marks carry meaning. Six different words can be built on the same syllable 'ma' just by changing the mark. Human readers can often recover unaccented text from context; models cannot always. In practice, users type without accents all the time, especially on phones. Telex typos give you strings no dictionary contains. Voice transcripts drop or misplace marks too.

A model trained only on clean, edited text gets confused by all this. A model that has seen enough real, messy examples learns to guess correctly. This matters even more for voice products: Vietnamese speech to text produces the text that the LLM then reads, so an accent error at the recognition layer flows straight into the understanding layer. That is why we treat speech recognition and language modelling as two halves of one problem rather than separate products.

How people actually talk

This is the hardest part to measure. Vietnamese pronouns shift with age, rank and relationship. Particles such as 'a', 'nhe' and 'nha' change the tone of a whole sentence. A bank employee writes something warm and casual, the model answers in textbook formality, and while nothing is technically wrong, it feels stiff. Vocabulary differs between northern, central and southern speakers. Younger users mix English into sentences. Abbreviations are everywhere.

A multilingual model usually copes with most of this, but sometimes it 'corrects' what did not need correcting, or answers in a register that is too formal. For support bots and internal assistants, tone is part of the product.

When you do not need dedicated training

I do not want this to read as an advertisement. If you only need email summaries, draft documents or ticket classification with limited data, a multilingual model with a well-written prompt is often enough. If the problem is missing or fresh knowledge, retrieval (RAG) fits better. Training or fine-tuning earns its cost in a few situations: token cost for Vietnamese becomes a burden at scale, you need consistent tone and terminology, sensitive data must stay on private infrastructure, or the field demands precise vocabulary, as in healthcare and pharmacy. We compare these options in a separate post on fine-tuning LLMs.

Where AIVISION stands

AIVISION is an AI company in Vietnam that trains LLMs on a cluster of 24 NVIDIA H200 and 8 NVIDIA B300 GPUs. On the hardware side, the H200 offers 141GB of HBM3e memory at roughly 4.8TB/s of bandwidth, and the B300 offers roughly 288GB of HBM3e. Larger memory means longer contexts and bigger models with less sharding. We have released the L1.0 LLM for Vietnamese and the E1.0 speech-to-text model, and we focus on making Vietnamese AI understand the language as people speak it, not only as it appears in books.

We also do training and fine-tuning for specific domains such as healthcare and pharmaceuticals. In those areas the model is only an information and administrative aid. It does not diagnose, prescribe or recommend dosages, and the final decision always belongs to qualified medical professionals.

Spanish-speaking and Filipino markets face very similar issues: skewed tokenizers, thin data, dense dialects. What we learn from Vietnamese can travel to markets like Mexico and the Philippines, as long as teams measure before they believe.

Related insights

See all insights