What Is LLM Fine-Tuning, and When Beats Prompting or RAG

01/10/2026

What Is LLM Fine-Tuning, and When Beats Prompting or RAG

The question I hear most from clients is not 'which model is best' but 'do we need to fine-tune?'. And the honest answer, more than half the time, is not yet. That sounds odd coming from a team that does LLM fine-tuning for a living, but we would rather say it than take a project that leaves both sides disappointed.

What fine-tuning really is

After its initial training, an LLM is a huge block of weights that has absorbed language and general knowledge. Fine-tuning continues training that block on a much smaller, deliberate dataset to change how it behaves. It is not 'teaching facts' in the naive sense. What fine-tuning does best is change behaviour: tone, output format, handling of terminology, how it refuses, how consistent it stays across thousands of queries. It does worse at absorbing new facts that keep changing.

There are several flavours. You can update all weights, or use lightweight methods like LoRA that touch a small fraction of parameters. The lightweight route is cheaper, faster, easier to swap, and a sensible starting point for most business problems. Full fine-tuning makes sense when you need deeper change, such as adding a language or a whole domain.

Three tools, three different problems

The easiest way to remember is to ask whether your model does not know, does not follow, or does not speak the right way.

  • Prompting fixes the case where the model has not understood what you want. It is the cheapest and fixable in minutes. The limit is that long prompts cost tokens on every call and can still be ignored in long contexts.
  • RAG fixes missing information, especially private or fast-changing material such as price lists, internal procedures and product documents. The data lives outside the model, so you update it by swapping documents, not retraining, and you can cite sources.
  • Fine-tuning fixes the case where the model knows enough but behaves wrongly for your use, or where you need something smaller, faster and cheaper at scale.

These are not mutually exclusive. A good system often uses all three: a fine-tuned model for tone and format, RAG for knowledge, prompts for per-situation instructions.

Signs you should fine-tune

In our experience fine-tuning makes sense when a few things line up. You have tried careful prompting and few-shot examples, yet outputs still drift in the same cases. You need a strict output format, such as JSON against a fixed schema or minutes in an organisation's template, and you do not want to pay for a thousand-token prompt on every call. You want a smaller model to cut cost and latency while keeping quality on a narrow task. Your data must stay on private infrastructure. Or your field has terminology that general models often misuse, such as healthcare, pharmacy or law.

For Vietnamese there is one more reason: if the base model handles the language poorly at the tokenizer or tone level, fine-tuning can pull its behaviour closer to how real users talk. We explain that in our post on why a Vietnamese LLM needs dedicated training.

Signs you should not

If you have no evaluation set, do not fine-tune. Without a yardstick you will not know whether the tuned model is better or worse, only whether it feels better. If your information changes often, do not fine-tune, because each change means retraining. If you have fewer than a few hundred examples of poor quality, the problem is the data, not the method. And if one carefully rewritten prompt solves 90 percent of the issue, stop there; running a dedicated model is not free.

A common trap is using fine-tuning to 'teach' facts and then being surprised when the model still invents things. Fine-tuning can make a model more confident about a domain it does not really master, which can make hallucinations sound more convincing. For information that must be right word for word, RAG with source checking is safer.

Where the real cost sits

People assume the cost is GPU time. GPUs are only part of it. Most of the effort goes into preparing data, writing labelling guidelines, checking quality, building an evaluation set and maintaining the model after launch. A few thousand clean examples reviewed by domain experts are usually worth more than a few hundred thousand scraped automatically. We describe the process in our post on the Vietnamese fine-tuning workflow.

On hardware, AIVISION trains on a cluster of 24 NVIDIA H200 and 8 NVIDIA B300 GPUs. The H200 has 141GB of HBM3e and the B300 roughly 288GB, which leaves room to iterate on fine-tuning experiments rather than waiting in a queue. But hardware only shortens run time. It does not replace understanding the problem.

A short way to decide

Before starting, answer three questions with real examples rather than impressions. Is the model missing information, or missing behaviour? Do you have an evaluation set of at least a few dozen representative cases? Are the cost and latency of your current setup actually a problem? If information is missing, go the RAG route. If behaviour is missing and you have a yardstick, then fine-tuning becomes a serious candidate.

For healthcare and pharmacy we add a constraint: the model supports information and administrative work only, such as drafting summaries or looking up documents. It does not diagnose, prescribe or advise on dosage. Medical professionals make the final call, and the system should make their checking easier, not harder.

AIVISION has released the L1.0 LLM for Vietnamese and builds fine-tunes for specific domains. If you want a sense of how we think about AI more broadly, the AIVISION homepage has an overview.

Related insights

See all insights