Fine-tuning LLMs for pharma: drug lookup, labels, pharmacy support

01/10/2026

Fine-tuning LLMs for pharma: drug lookup, labels, pharmacy support

Picture a Saturday afternoon at a crowded pharmacy counter. The pharmacist is advising one customer while another cuts in: what is in this product, how is this box different from that one, the leaflet is printed so small nobody can read it. Most of these questions are lookups, and lookups are something machines are good at. That is why fine-tuning LLMs for pharmaceutical information is attracting so much attention, and why it deserves a careful discussion of what it should and should not do.

At AIVISION we train and fine-tune Vietnamese language models for specific domains, including pharmacy and drug information. This article is the view of someone building the product, not an advertisement.

Why a general model is not enough for drug information

Ask a general-purpose model about a medicine and you will get an answer that sounds polished. The trouble sits in the word sounds. Drug information is the kind of data where one small detail changes the meaning: brand name versus active ingredient, dosage form, pack size, the headings in a package leaflet. General models are trained on huge amounts of internet text, in which Vietnamese drug information is a small and uneven slice.

The familiar result is a model that mixes up two products with similar names, or invents a section that does not exist in the source document. It is not malicious. It was optimised to write fluent sentences, not to stop when it is unsure.

What fine-tuning is for, and which data it needs

Fine-tuning means continuing to train a base model on data from a narrow field, so it gets used to that field's vocabulary, document structure and typical questions. In pharmacy, three kinds of work come up most often.

Looking up ingredients and composition

A user asks what a product contains, or which products in the pharmacy's catalogue contain a given active ingredient. This is a mapping problem between names. The model has to understand brand names and generic names, and also the accented and unaccented spellings that customers type.

Reading and summarising labels and leaflets

Turning a long leaflet into a readable summary that keeps the original order of sections and the original content. The model re-presents what the document says. It does not add recommendations of its own.

Supporting pharmacy staff

Searching the catalogue quickly, recalling internal procedures, drafting reply text for the pharmacist to edit. The person behind the counter gets a free hand, and the decision stays theirs.

Training data should come from authoritative sources with clear usage rights: drug documents approved by the regulator, the organisation's own catalogue, internal procedures. The cleaner the data, the less cleanup later. What we avoid is mixing in material of unclear origin, because one wrong sentence in the training set can become a wrong sentence spoken with confidence.

The boundary to hold firmly

Some things the model does not do. That is a design decision, not a defect to be fixed later.

  • It does not advise doses for an individual.
  • It does not suggest a medicine to treat a particular symptom or condition.
  • It does not replace the pharmacist in judging interactions, contraindications or a user's circumstances.
  • When someone describes symptoms, the system points them to a qualified person.

Drug information is different from drug advice. The first is document lookup: how this product is packaged, which heading the leaflet uses. The second requires knowing who the user is, what they already take, what conditions they have, and that is work for someone trained for it. We treat this boundary as a condition for the product to exist at all, not as a small disclaimer at the bottom of the page.

Retrieval or baking knowledge into the weights

One technical decision is often misunderstood: should all drug knowledge be pushed into the model through fine-tuning? In many teams' experience, no. Catalogues change constantly. New products arrive, old ones are withdrawn, leaflets are updated. If knowledge lives in the weights, every change means retraining, and you never quite know whether the model has forgotten the old version.

A more sensible split: fine-tune so the model is fluent in the language of the field, understanding names, abbreviations and the style of pharmaceutical documents, and keep the specific content in a document store that is updated separately. The model reads from it and answers with citations. When a document changes, you update the store. When nothing relevant is found, the system says nothing was found instead of improvising.

Being honest about capability

AIVISION trains on a cluster of 24 NVIDIA H200 GPUs and 8 NVIDIA B300 GPUs, has released the L1.0 LLM for Vietnamese, and takes on fine-tuning for fields such as healthcare and pharmacy. The infrastructure lets us try many training configurations. But quality on a specific pharmacy's problem is only known after measuring, on that pharmacy's own catalogue and the questions its customers actually ask. So the process we favour starts with a test set written by the organisation's own pharmacists, not with an impressive number borrowed from somewhere else.

You might ask why the pharmacists should write the test set. Because they are the ones who know. The hard questions are usually ones only they would think of: a customer mispronouncing a drug name in a regional way, using a habitual abbreviation, asking half a question and trailing off.

What is worth trying first

If you run a pharmacy or a small chain and are weighing this up, I would start on the staff side, not the customer side. Give a few pharmacists an internal tool for searching the catalogue and leaflets, let them use it for a month, and log where it is wrong. Only then think about putting it in front of buyers. It is slower, but it gives you real evidence of whether the tool is right or wrong, and pharmacists are not pushed to trust something they have not verified.

The technology is useful when it makes lookup faster and lets qualified people spend their energy on the part only they can do. When it starts pretending to do that part, it has gone too far.

Related insights

See all insights