Medical data security with LLMs: on-premise, anonymisation, access

01/10/2026

Medical data security with LLMs: on-premise, anonymisation, access

A hospital CTO once told me, half joking, that his staff needed no training to learn how to paste a patient record into a public chatbot and ask for a summary. The joke is not funny because it is true. When we talk about medical data security with LLMs, the biggest risk is often not a hacker but the convenience at a user's fingertips.

As a builder of Vietnamese language models, AIVISION treats security as part of the design from day one, not a coat of paint at the end. This article covers three layers we think about: where the model runs, how data is handled before it reaches the model, and who is allowed to see what.

Layer one: where the model runs

The first question, before any discussion of quality, is where the data goes. There are three common deployment styles, and each has different trade-offs.

  • A public cloud service called through an API: fast and cheap to start, but data leaves your infrastructure. For patient records, many organisations cannot accept that.
  • A private cloud or isolated data region: data stays in an environment whose terms you control, a balance between flexibility and safety.
  • On-premise: the model runs inside the hospital's own data centre. Data never leaves the walls, but you carry the hardware, operations and updates.

I do not believe on-premise is always the right answer. For a small clinic, running its own GPU server can cost more effort than it is worth, and a badly secured home-built server is worse than a well-managed cloud. But for a large hospital, or data under strict legal requirements, running the model inside your own infrastructure is reasonable and many choose it. Our design direction is to support in-house deployment as well, where the model goes to the data instead of the other way round.

Layer two: anonymise before processing

Many tasks do not need to know who the patient is. Summarising a note, normalising terminology, finding the matching procedure: none of this depends on a name or an ID number. So why feed them in?

Anonymisation is a filter before data touches the model. It detects names, phone numbers, addresses, patient IDs and birth dates, and replaces them with placeholders. After the model returns a result, if needed, the real values are restored from a mapping table stored separately and protected by permissions.

The difficulty lives in the details. Vietnamese personal names often collide with ordinary words, so a name in one sentence is a common noun in another. Addresses are abbreviated, incomplete, buried in descriptive text. A rare detail, such as an unusual occupation plus a small place of residence, can identify someone even with no name present. So anonymisation is never perfect, and we do not promise it is. It reduces risk, it does not remove it. It needs testing on the organisation's own kind of data, with a periodic manual review of samples.

An intermediate option is worth considering: send the model only the minimum data the task requires. Summarising one course of treatment does not need the patient's administrative history attached. Data minimisation sounds dry but is among the cheapest and most effective measures.

Layer three: access control

Even when the model runs in-house and the data is anonymised, a question remains: who can ask what. A record search system attached to an LLM can quietly become a back door if it can read everything and answers anyone.

  • Role-based permissions: nurses, doctors and billing staff see different parts.
  • The model retrieves only from documents that the person asking is allowed to view, not from the whole store.
  • Audit logs: who asked what, when, and which documents the answer came from.
  • Limits on bulk export, and alerts for unusual queries.

The second item is where many projects stumble. If you index every record into one store and let a model search it, but forget to attach access rights to each document, you have built a tool that lets a staff member ask about a patient outside their care scope. Permissions must be enforced at the retrieval layer, before content reaches the model. You cannot rely on the model to behave.

The things that get forgotten

There are a few leaks people rarely think about. System logs can contain the full text of questions and accidentally become a second sensitive data store. Backups need encryption. Data used for further training or fine-tuning needs its own rules: when it may be used, whether it has been anonymised, who approves. And the provider, including us, should state in writing whether data is retained, for how long and for what purpose.

I also want to mention people. The best process loses to an employee who pastes data into a public tool because the official one is too slow or awkward. Good security is security that colleagues actually want to use. If the internal tool is more convenient than the shortcut, people will take the main road.

What we will not claim

AIVISION trains on a cluster of 24 NVIDIA H200 GPUs and 8 NVIDIA B300 GPUs, has released the L1.0 LLM and the E1.0 speech to text model for Vietnamese, and does fine-tuning for fields such as healthcare. We do not make absolute promises of safety, because no system achieves that and anyone who promises it deserves suspicion. The sensible path is to assess risk per context, agree with the organisation's legal and information security teams, test first on synthetic data, and only then touch real records. And whatever the technology, anything a model produces in a medical setting still needs human review before it is used.

Related insights

See all insights