Reducing LLM Hallucination in Specialised Domains
01/10/2026

During an internal test, we once asked a model about a very specific regulation, the kind with an article number and a year of issue. It answered fluently, cited the article number, even gave the year. Everything sounded plausible. Except that the article did not exist. This is not rare, and to a non-expert it is nearly impossible to spot. It is LLM hallucination: the model produces fluent, confident content that is wrong or unsupported.
Why models make things up
It helps to understand the mechanism so expectations are right. An LLM is trained to guess the most plausible next word, not to look up truth. On things it knows well, a plausible guess usually coincides with the correct answer. On rare, recent or deep topics it still guesses, and the guess sounds just as real. At no point does the model 'know' it is inventing. Nor does it have a built-in way to stop and say 'I am not sure', unless that behaviour has been trained in.
In specialised domains the situation is worse for three reasons. Deep material is thin in general training corpora. One wrong word can change the meaning: a drug name close to another, a unit of measure, a dropped 'not'. And the reader is often inclined to trust because the paragraph looks professional. Vietnamese adds a layer: specialist terms often have several names, drugs have both generic and brand names, and source documents are frequently poor scans.
No silver bullet, only layers of defence
Honestly, nobody has eliminated hallucination entirely. What can be done is to lower its frequency, make errors easier to catch, and design the system so one invented sentence does not travel straight to a consequence. These are the layers we use.
Give the model documents to lean on
RAG (retrieve, then generate) has the clearest effect. Instead of letting the model recall, we find relevant passages and instruct it to answer only from them. This cuts invention noticeably for private or fast-changing knowledge. But RAG is not magic: if retrieval fetches the wrong passage, the model will confidently answer from it. The quality of chunking, indexing and ranking matters more than people expect. Measure retrieval separately from generation.
Require citations
Ask every claim to come with a quoted passage and its location in the document, so users can click through and verify. On the system side, you can automatically check that the quoted passage really exists and supports the claim. An answer that cannot cite a source should be flagged or blocked.
Teach the model to say 'I do not know'
This is where LLM fine-tuning earns its keep. Fine-tuning does not teach many new facts, but it teaches behaviour very well: politely decline when the documents lack the information, hand off to a person when the question is out of scope, ask back when the premise looks wrong. For that, the dataset needs plenty of refusal examples, otherwise the model learns to always answer. Note the trade-off: over-training refusal makes the model dodge questions it could have answered. Measure both directions.
Narrow the scope and the output
A model with a narrow job invents less than one expected to do everything. Limit topics, use structured output (fill in fields instead of free writing), and a low sampling temperature for tasks that need precision. For numeric data such as doses, prices and dates, pull values directly from a structured database and insert them, rather than letting the model 'write out' numbers.
Check after generation
A separate verification step can compare the answer with the sources, catch drug names or terms not in the domain dictionary, or flag sentences containing figures that need a human look. It is imperfect, but catching part of the errors is far cheaper than leaving users to catch them.
Keep a person in the loop
For consequential decisions, a qualified person must review before action. Design the interface to make that easy: show sources next to the answer, highlight where the model is unsure, allow quick edits and record the edits as data for the next improvement round.
Healthcare and pharmacy: clear boundaries
AIVISION trains and fine-tunes LLMs for specific domains such as healthcare and pharmaceuticals. There we set firm boundaries: the model is a tool for information and administrative support, such as summarising documents, looking up sourced information and drafting text. It does not diagnose, prescribe or advise on dosage. Medical professionals make the final decision. Designing this way is not only a legal or ethical matter; it is also the most realistic way to accept that hallucination cannot yet be fully removed.
You can only reduce what you measure
You cannot reduce what you do not measure. Build a dedicated hallucination set: questions answerable from the documents, questions with no answer, questions with false premises, questions about entities that do not exist. Measure the invention rate before and after every change, and read the failures by hand. A high general score says little about this, as we discuss in our post on evaluating Vietnamese LLMs.
We train on a cluster of 24 NVIDIA H200 and 8 NVIDIA B300 GPUs and have released the L1.0 LLM for Vietnamese, but raw compute does not produce reliability by itself. Reliability comes from clean data, layered system design and the habit of doubting your own output. For more on how AIVISION approaches these problems, see our homepage.