Vietnamese LLM Training Data: Cleaning, Anonymising, Copyright

01/10/2026

Vietnamese LLM Training Data: Cleaning, Anonymising, Copyright

There is a saying in this field I am tired of hearing and still believe: a model is only as good as its data. Tired because it sounds like a slogan. True because every time a project fails and we trace it back, the cause is almost always some data file nobody read carefully. This post is about Vietnamese LLM training data: the dull work that has to be done and the traps we at AIVISION have seen repeatedly.

Cleaning: the work nobody wants to show off

Raw Vietnamese text is full of debris. Leftover HTML, odd characters from encoding errors, diacritics broken in conversion from legacy formats, paragraphs repeated dozens of times, headers and footers from scanned documents wedged into the middle of sentences. Unicode alone is a headache. The same letter with a tone mark can be stored precomposed or as a base letter plus a combining mark. They look identical to the eye and are different strings to a machine. If you do not normalise to one form, the tokenizer treats them as different words and the model learns a thinner signal. Unicode normalisation is nearly always the first step.

Next comes quality filtering. Drop text that is too short, full of symbols, or has an unusual share of non-Vietnamese characters. Remove duplicates, both exact and near-duplicate. Raw machine translation should be dropped or down-weighted, because the model picks up its stilted voice. Do not filter too hard, though. I have seen a team filter so aggressively that all unaccented and abbreviated text disappeared, and the trained model could not understand real user messages. Some of the 'dirt' users actually type is exactly what you need to keep, selectively.

Anonymisation: easy to say, hard to do

Business data, especially customer chats, medical records and internal documents, is full of personal information: names, phone numbers, addresses, national ID numbers, emails, licence plates, patient codes. Before training, these have to be masked or replaced.

The practical approach is regex rules for structured patterns such as phone numbers, emails and ID numbers; a named-entity model for people and addresses; then consistent placeholders, so the same person stays the same token within a conversation and the text still flows. After that, a human must check a random sample. Vietnamese names are harder than you would think: several common given names are also ordinary nouns or adjectives. Abbreviated or accent-free addresses are harder still to catch.

Two admissions. Anonymisation is rarely perfect, so you add other safeguards: restricted access, training on private infrastructure, audit logs. And some information can re-identify a person when combined, even if each piece is masked. For medical data the bar must be much higher, and someone responsible for legal or compliance matters should be involved from the start, not after training is done.

Copyright and permission to use

'Can we train on this data?' has no universal answer. It depends on origin, licence, the source site's terms, client agreements and local law. We are not legal advisers and will not draw general conclusions, but some working rules help.

  • Record the source and licence of every dataset. Without that, two years later nobody can answer where something came from.
  • Client data is used only within the agreed scope, and the agreement should state in writing whether it is used for training and who owns the fine-tuned model.
  • Copyrighted content such as books, news and paid documents needs separate review. Do not assume it is fair game because it is online.
  • When in doubt, ask a lawyer or drop the source. Model quality rarely hinges on one questionable source.

The traps we see most

Leakage between training and test sets is the classic. Evaluation scores soar because the model has seen the exam. Remove near-duplicates before splitting.

Bias in voice and region. If most chat data comes from one northern branch, the model speaks that way and answers southern customers awkwardly. If all medical documents come from one hospital, the model reflects that hospital's habits and terminology.

Inconsistent or wrong labels. When several people write the sample answers, each with their own style or even conflicting views on a procedure, the model learns the mess. Clear writing guidelines and a cross-review round cost far less than retraining.

Outdated data. Prices, rules and procedures change, yet the model keeps answering with the old version, confidently. For information that changes, keep it outside the model and retrieve it at run time instead of baking it into the weights.

Synthetic data from other models. It is convenient and cheap, but unchecked, errors and machine style get amplified generation after generation. Use it selectively, with human review, and never treat it as a source of truth.

A habit worth keeping

My suggestion is simple: before training, have the team sit down and read 100 random samples, actually read, not skim. Once we discovered a whole segment of samples opening with the same long-winded greeting purely this way. Automated filters may miss it; human eyes do not.

In healthcare and pharmacy, data needs to be reviewed by medical experts even more. The model is only a tool for information and administrative support. It does not diagnose, prescribe or advise on dosage. A sample that accidentally teaches the model to sound like it issues medical instructions should be removed immediately.

AIVISION trains on a cluster of 24 NVIDIA H200 and 8 NVIDIA B300 GPUs, but much of that compute is pointless if the input is dirty. For us, investing in data always pays better than investing in more hardware, which sounds backwards coming from a company with plenty of GPUs.

Related insights

See all insights