Evaluating Vietnamese LLMs Properly: Why Benchmarks Fall Short
01/10/2026

Imagine hiring someone because they scored high on the entrance exam, and on day one they cannot work out how to answer a customer. The exam score was not wrong; it simply measured something other than the job. LLM benchmark scores are often like that. For evaluating Vietnamese LLMs the gap is wider still, because most public benchmarks were designed for English and translated, or built on multiple-choice exam formats far from how people really use models.
What benchmarks measure, and what they miss
Public benchmarks have value. They give a common baseline for comparison and help catch models with serious problems in knowledge or basic reasoning. We still look at them. But four limits are worth remembering.
Content: multiple-choice knowledge questions measure the ability to pick an answer, not to write a customer email in the right tone. They do not measure how a model handles accent-free messages, abbreviations or vague requests.
Translation: a benchmark translated from English carries the phrasing and culture of the source. A question about another country's laws or procedures says nothing about understanding Vietnamese context.
Contamination: if benchmark questions leaked into training data, a high score only reflects memory. With long-standing public benchmarks this risk is not small, and from the outside it is nearly impossible to verify.
Targeting: once a number becomes what everyone chases, people optimise for the number. Models are no exception. That is why we do not present benchmark scores as the main evidence for any AIVISION product, and why we are wary of anyone who does.
Your own evaluation set: small but on point
The most effective approach we know is to build an evaluation set for your own problem. It need not be big. A few hundred representative cases drawn from real situations: questions customers actually asked (anonymised), real documents to summarise, hard cases that tripped up the old system. Each case comes with clear grading criteria: what information must appear, what must not, what tone, when to refuse.
Include dedicated groups for Vietnamese-specific challenges: no diacritics, Telex typos, abbreviations, English mixed in, pronouns that follow social roles, dialects. And groups that probe limits: out-of-scope questions, attempts to force invention, questions with false premises. A good model does not only answer correctly when it knows; it also says 'I am not sure' or 'the document does not mention this' when it does not. That ties directly to reducing hallucination, which we cover in its own post.
Who grades, and how
There are three ways to grade, each with strengths and weaknesses.
- Automatic rule-based checks: fast and repeatable, good when answers are clear, such as valid JSON, correct figures or the presence of a certain code. Not enough for free-form text quality.
- Human grading: the most trustworthy for nuance, naturalness and professional correctness, but slow, expensive, and graders disagree with each other. You need a clear rubric and a measure of agreement between graders.
- Another model as judge: cheap, fast, scalable. But judge models have their own biases, such as favouring long answers or prose that resembles their own, and may grade specialised Vietnamese wrongly. Use one only after comparing it against hand grading on a small sample, so you know how far it drifts.
A practical combination: automatic checks for what machines can verify, a judge model to screen large volumes, human graders for a representative sample and every high-stakes case. In fields such as healthcare or pharmacy, medical experts should grade directly, because a small terminology error can have large consequences even when the paragraph reads fluently.
Measurement mistakes we see often
Unfair comparisons: one model gets a carefully tuned prompt, the other a default one. The result reflects the prompt writer, not the model.
Running once and believing it. LLM output has randomness. On a small evaluation set, a difference of a few points may be noise. Repeat runs, look at the spread, and do not celebrate an improvement that sits inside the noise.
Looking only at the average. A model can be excellent on 95 percent of cases and seriously bad on the other 5. If those 5 percent are about drug doses or legal rules, the average hides the most important thing. Break results down by group and read the failures one by one.
Forgetting non-quality measures: latency, cost per call, stability with long contexts, behaviour when users try to break it. A model a few points better but twice as slow can be the worse choice.
Not measuring after launch. Real users ask differently from your evaluation set, and drift further over time. Sample real conversations periodically, grade them again and add them to the set.
How AIVISION sees this
We train LLMs on a cluster of 24 NVIDIA H200 and 8 NVIDIA B300 GPUs, and have released the L1.0 LLM for Vietnamese along with the E1.0 speech-to-text model. We do not offer rankings or generic accuracy percentages here, because a number detached from a specific task misleads more than it helps. What is worth comparing is the result on your own questions, with your own data.
If you are choosing between models or weighing LLM fine-tuning, my advice is to spend the first week building the evaluation set, not training. It makes every later decision less a matter of feeling. More about how we work is at AIVISION.