Vietnamese LLM for Legal and Administrative Documents

01/10/2026

Vietnamese LLM for Legal and Administrative Documents

There is an odd paradox in Vietnamese legal texts: the more precision they need, the harder they are to read. A guidance letter cites a circular, the circular cites a decree, and the decree has been amended twice. To answer a seemingly simple question you have to follow the whole chain. That is the kind of work a Vietnamese LLM can do fairly well when used properly, and does very badly when treated as a virtual lawyer. This article is about that boundary.

Three jobs worth giving to a model

Semantic search

Traditional keyword search makes you guess the exact words the document uses. You type "quit without notice" but the law says "unilateral termination of the labor contract in violation of regulations". Semantic retrieval combined with an LLM that reads the retrieved passages bridges the two phrasings. It is a familiar architecture: retrieve relevant passages from a document store, hand them to the model, and ask for an answer with citations to specific articles and clauses.

Drafting

Reply letters, proposals, notices, template contracts. Most of these follow a form with repeated phrasing. An LLM produces a first draft quickly in the right format and administrative tone, and the officer edits it. In a government office the time saved is not in typing but in not staring at a blank page.

Comparison and cross-checking

This one is rarely mentioned but very valuable. Compare two versions of a document and list what changed. Compare a contract against the company's standard template and flag deviating clauses. Check whether a draft contradicts itself, say Article 5 sets ten days while Article 12 sets fifteen. Doing this by hand is tiring and easy to get wrong, while a machine reads without getting bored.

Where an LLM goes wrong

The most dangerous thing in this field is that an LLM can invent. Not in a silly way, but in a way that looks real: a plausible document number, a clause that seems right but does not exist, or a provision that has expired cited as if it were in force. I have seen a model give a coherent, complete answer citing "Article 47", and when the real document was opened, Article 47 was about something entirely different.

So here are the minimum principles I would insist on:

  • Answer only from the loaded corpus: the system must refuse when it finds no basis, rather than fall back on the model's general knowledge.
  • Verifiable citations: every claim points to the exact document, article and clause, and the user can click through to the original text.
  • Validity management: the corpus must know which documents are in force and which have been amended or replaced. The model does not know this on its own.
  • A responsible person: the output is reference material for a qualified person, not a legal conclusion.

I stress the last point. Such a tool is an aid, not legal advice. If someone uses a machine answer to make a legally consequential decision without a lawyer or specialist reviewing it, the fault lies in the process, not only in the model.

What makes legal Vietnamese hard

Vietnamese legal language has traits that trouble general-purpose models. Very long sentences with nested clauses. Dense Sino-Vietnamese vocabulary. Cross-references like "as provided at point b, clause 2, Article 15 of this Law" demand that the model track hierarchical structure exactly. And many everyday words carry a distinct technical meaning in legal context, such as the several Vietnamese words for different kinds of "time limit".

That is why finetuning an LLM on domain-specific Vietnamese text makes sense. AIVISION trains and finetunes LLMs for Vietnamese and for specific fields, has released the L1.0 model, and legal and administrative work is one application direction we consider reasonable. I am not quoting an accuracy figure, because for legal text an average says nothing: a single error in a key clause is enough to do harm.

Where to start

Choose a narrow scope, such as an organization's internal regulations or one area of law like labor. Load a clean corpus that distinguishes validity status. Ask a few specialists to put the hard questions they actually meet to the system, then compare each answer with the source text. After a few rounds you will know exactly where the system can be trusted.

To learn more about Vietnamese language models, see the AIVISION site.

Related insights

See all insights