Fine-Tuning a Vietnamese LLM: Data, SFT and Evaluation

01/10/2026

Fine-Tuning a Vietnamese LLM: Data, SFT and Evaluation

The hardest part of a fine-tuning project is rarely the moment you press run. It is the two weeks before, when the team argues about what a 'correct' answer even looks like. This post walks through the Vietnamese LLM fine-tuning workflow we at AIVISION usually follow, in the real order of the work, with the places people tend to trip.

Step zero: describe the problem with examples

Before touching data, write down 30 to 50 examples you want the model to handle well, with the answers you would accept. It sounds crude, but it is the most valuable asset in the whole project. It forces the client to be specific: 'a summarising assistant' means how long, what to keep, what to drop, what tone. Both sides often discover they understood the task differently right here, and finding that out now is far cheaper than finding it after training.

Collecting data

Data usually comes from three places. The organisation's own material: chat history, emails, minutes, procedure documents. Public data relevant to the field. And newly written examples by experts, usually the smallest and most precious part. Our advice is to start small but clean. A few thousand high-quality samples read by people in the profession often beat hundreds of thousands of unchecked ones.

One small trick: let a domain expert look at a few dozen random samples before you scale up. They spot errors engineers miss, such as a term used slightly wrong or a procedure that changed last year. Cleaning, anonymisation and copyright matter enough that we cover them in a separate post on training data.

Formatting and splitting

SFT (supervised fine-tuning) data takes the form of instruction and response pairs, or multi-turn conversations. You need one consistent template: is there a system prompt, how is the user addressed, what format is used. If one sample says a warm 'Dạ anh' and another says a stiff 'Dear customer', the model learns a muddled voice.

Then split into three sets: training, validation for monitoring during the run, and a final test set kept sealed until the end. The classic mistake is leakage, where the same question with a few words changed lands in both training and test. Scores look beautiful and mean nothing. Remove duplicates and near-duplicates before splitting.

SFT: the choices that actually matter

Training sounds like high science, but the practical decisions are fairly concrete. Which base model: what size, how much Vietnamese it already handles, what licence. Full fine-tuning or LoRA. Learning rate and number of epochs. With small data, too many epochs make the model memorise and lose generality, so it answers stiffly and forgets common language. The familiar sign: training loss keeps falling while validation loss starts to rise.

Another phenomenon is catastrophic forgetting: the model gets good at the new task and worse at old ones. Mitigations include mixing in general, diverse data, using a lower learning rate, or using LoRA. Every option has a price and there is no universal recipe, so run several small configurations and compare.

On infrastructure, we train on a cluster of 24 NVIDIA H200 and 8 NVIDIA B300 GPUs. The H200 offers 141GB of HBM3e at roughly 4.8TB/s, and the B300 roughly 288GB of HBM3e. For fine-tuning, large memory allows bigger batches and longer contexts with less sharding, so experiments move faster. Do not read that as a quality guarantee: data still decides the outcome.

Evaluating before deployment

This is the step most often shortened, and the one most worth protecting. We usually evaluate in three layers.

  • Automatic checks on the sealed test set: correct format, correct terminology, correct content against reference answers where a machine can verify it.
  • Expert grading by hand on a representative sample, with a clear rubric, especially for terminology, ambiguous cases and cases where the model should refuse.
  • Adversarial testing: deliberately off-topic questions, attempts to force the model to invent facts, text without diacritics, abbreviations and typos.

You also need a baseline: the original model with a good prompt. If the fine-tuned model is merely equal, you spent effort for nothing, and you should say so plainly. We explain why benchmark scores alone fall short in a separate post on evaluating Vietnamese LLMs.

Deployment and the loop afterwards

Going live is not the end. Start with a narrow scope or a shadow mode where it runs alongside people who still make the decision, gather real feedback and log the cases where the model is wrong. Those cases are raw material for the next fine-tuning round. Track cost, latency and how user questions drift over time, because real data always moves away from the training data.

In fields like healthcare and pharmacy we add another control layer. The model supports information and administrative work only. It does not diagnose, prescribe or advise on dosage, and every relevant output goes to a medical professional for review and the final decision.

AIVISION has released the L1.0 LLM and the E1.0 speech-to-text model for Vietnamese, and builds fine-tunes for specific domains. We do not offer generic quality numbers, because a number only means something attached to one task and one evaluation set. What we can say is that this workflow, together with the habit of measuring before believing, spares many projects their most unpleasant surprises.

Related insights

See all insights