RAG vs Finetuning an LLM for Internal Company Chatbots
01/10/2026

A client once told me, "Just stuff all our company documents into the model and you're done." I understand why they thought so. But how you put those documents in, through RAG or by finetuning an LLM, leads to two very different systems in cost, reliability and upkeep. This post describes how we think about that choice when building internal chatbots for Vietnamese companies.
Two approaches, two different jobs
RAG (retrieval-augmented generation) works like this: when a user asks something, the system finds a few relevant passages in your document store, hands them to the model along with the question, and the model answers from what it just read. The model learns nothing new. It is given reference material, like an employee opening a handbook before replying.
Finetuning is different. You continue training the model on your own data, and the change really lives in the weights. What the model picks up is usually style, format, terminology and ways of reasoning within a field. It is a poor way to cram a thousand pages of HR policy into the model, because knowledge absorbed that way is easily remembered wrongly or only partially.
An image I use often: RAG is giving an employee access to a library. Finetuning is sending them through a vocational course. The library helps them look up the right fact. The course helps them speak and work the way the profession does.
When RAG should come first
If the chatbot mostly answers questions from existing documents, such as procedures, policies and product guides, start with RAG. The reasons are practical:
- When a document changes, you update the store instead of retraining.
- You can cite sources, so users can check where an answer came from.
- Access control is easier: accounting only sees accounting documents.
- The starting cost is much lower.
RAG is not a cure-all, though. The hardest part is usually not the model but the data. I have watched teams spend weeks just dealing with blurry scanned PDFs, broken tables, and documents that exist in three conflicting versions. If retrieval pulls the wrong passage, the model will answer wrongly with confidence, and users will struggle to notice. For Vietnamese, chunking and semantic search are further affected by diacritics, compound words, and the unaccented writing common in internal messages.
When finetuning is truly worth it
Finetuning starts paying off when the problem is in how the model answers, not only what it knows. Some typical situations:
- The model repeatedly misreads domain terms, for example drug names, active ingredients and abbreviations in clinical records.
- Output must follow a fixed format, such as forms, coding schemes or templated reports.
- Tone must stay consistent with a brand or internal standard.
- You want a smaller, cheaper model that is good enough for one narrow task.
The price of finetuning is that you need clean data, labeled or high-quality examples, and a proper way to evaluate. Bad data gives a bad model, and finetuning can even erode some general ability. That is not rare. It is why we always want to see real data and a test question set before promising anything.
Combining both in practice
Most good internal chatbots I know use both, in an order that makes sense. Build RAG first, run it with real employees for a few weeks, and log the wrong answers. Then sort the errors:
- Errors from retrieving the wrong document or from missing documents: fix the RAG layer and the data.
- Errors from the model misreading terms, using the wrong format, or sounding off: these are candidates for finetuning.
This avoids the most expensive trap: finetuning early to patch a problem that actually came from the document store. I once saw a team finetune twice in a row before discovering the old policy was still sitting in the store.
| Situation | Lean toward |
|---|---|
| Looking up policies and procedures that change often | RAG |
| Need citations for verification | RAG |
| Domain terms misunderstood again and again | Finetune |
| Output must match a fixed template | Finetune |
| Need both fresh knowledge and domain fluency | Combine both |
An example from hospitals and pharmacy
Picture a chatbot helping pharmacists look up internal drug-usage guidance. The guidance changes with each update, so RAG handles lookup and citation. But generic names, brand names, department abbreviations and the way pharmacists phrase questions differ a lot from ordinary text. Finetuning helps the model interpret the question correctly from the start. The two complement each other rather than compete.
AIVISION is an AI company in Vietnam that does LLM training and finetuning for Vietnamese and for specific domains such as healthcare and pharmaceuticals, on a cluster of 24x NVIDIA H200 and 8x NVIDIA B300. We have released the L1.0 LLM and the E1.0 speech-to-text model for Vietnamese. When advising, we usually suggest customers begin with the cheap, reversible parts. More about our direction is on AIVISION.
Common mistakes
- Evaluating by gut feeling. Collect dozens to a few hundred real employee questions with reference answers, and rerun them after every change.
- Forgetting permissions. An internal chatbot that lets anyone ask for the payroll table is an incident waiting to happen.
- Treating the chatbot as a one-off project. Documents change, users change how they ask, and the system needs someone to look after it.
- Ignoring the answer "I don't know." A chatbot willing to say it found nothing is always more trustworthy than one that always has an answer.
If I had to give one short piece of advice: build RAG first, measure real errors, then decide which parts to finetune. The order is unglamorous, but it saves money and avoids weeks of unnecessary work.