On-Premise LLM vs Cloud: Choosing for Sensitive Data

01/10/2026

On-Premise LLM vs Cloud: Choosing for Sensitive Data

One meeting stuck with me. A hospital asked a very short question: does our patient data leave this building? The room went quiet for a few seconds. That question, not whether the model is big or small, decided the whole architecture. This post is about choosing between on-premise LLM deployment and the cloud when your data is sensitive, written from the perspective of a team building a Vietnamese LLM in Vietnam.

I do not have one answer for everyone. Some organizations should run fully in-house, and for others the cloud is the saner choice. What I can offer is a set of questions so you can decide for yourself instead of following a trend.

What sensitive data actually means

People hear the word sensitive and think of state secrets. The day-to-day reality is more modest: medical records, prescriptions, customer contracts, payroll, call-center recordings that contain ID numbers, financial reports not yet published. A single recorded customer call can contain a name, an address and sometimes an account number.

Before picking infrastructure, sort your data into a few groups: public, internal, confidential, and bound by law or contract. Each group may take a different path. I once watched a company force everything into one solution and then spend three months unwinding it, because 80% of their data was never sensitive in the first place.

On-premise: what you gain and what it costs

Running the model inside your own infrastructure, or in an isolated private zone, gives the clearest control. Data never leaves your network, access logs are yours, and when an auditor asks where the data went, the answer is short and easy to prove.

The cost is real too. You need GPUs, which are expensive and often have lead times. You need people to run them: driver updates, thermal monitoring, handling a dead node at 2 a.m. Models also get updated by hand, so without a strong team an on-premise system tends to age badly within a year. I have seen several projects start with enthusiasm and fade because nobody owned operations.

  • Good fit: hospitals, banks, agencies with strict data residency rules, companies with large and steady query volume.
  • Poor fit: small teams, experiments, irregular usage.

Cloud: more reasonable than you might think

The cloud gives you elasticity. Want to test an idea for two weeks? Rent GPUs for two weeks instead of buying a system. Large providers also invest more in security than a mid-sized company can manage alone. To be blunt, a server in an unattended equipment room is not automatically safer than a properly configured cloud region.

The issue lies elsewhere. Your data leaves your direct reach, so you rely on contracts, configuration, and the provider not using your data to train anything else. For data with domestic storage requirements, it also depends on where that cloud region physically sits. Have your legal team read that part carefully rather than letting engineering guess.

A hybrid setup is often the practical answer

What works best in my experience is splitting by sensitivity. Regulated data goes through a model hosted internally. Low-risk work, like drafting marketing copy or summarizing public documents, runs in the cloud. In between sits an anonymization step: names, phone numbers and patient IDs are replaced with placeholders before anything is sent out.

This needs an orchestration layer, and that layer can fail. If the anonymizer misses one phone number, you just pushed real data outside. Test that step seriously instead of treating it as magic.

CriterionOn-premiseCloud
Data controlHighDepends on contract and setup
Upfront costLargeSmall
Cost at heavy, steady useOften better long termCan climb quickly
Operational burdenHeavyLight
Speed of experimentationSlowFast

Why Vietnamese complicates the picture

Open models are usually strongest in English. Vietnamese brings its own trouble: tone marks, regional vocabulary, the abbreviations office workers invent, and specialized terms in medicine and pharmacy. To make a model understand your field, you usually have to finetune an LLM on internal data. That is exactly where on-premise gets important, because finetuning data is the most sensitive kind. It shows precisely how your business works.

AIVISION is an AI company in Vietnam. We train LLMs on a cluster of 24x NVIDIA H200 and 8x NVIDIA B300, have released the E1.0 speech-to-text model and the L1.0 LLM for Vietnamese, and do training and finetuning for domains such as healthcare and pharmaceuticals. We know customers often need data to stay with them, so we aim to discuss each case on its own terms rather than push a single deployment model. You can read more about how AIVISION works on our website.

Five questions to ask before deciding

  • Which data is bound by law or contract on where it is stored, and what share of your total data is that?
  • How many queries per day, and is the volume steady?
  • Who is on call when the system fails after hours?
  • If data leaked, how bad would it be, legally and reputationally?
  • Do you need to finetune on internal data, and is that data allowed to leave your network?

If most of your data falls in the restricted group, usage is steady, and you have or can hire an operations team, on-premise deserves serious consideration. If you are only experimenting and most data is public, starting on the cloud and migrating later is nothing to be embarrassed about.

One last note for engineering teams: do not pick infrastructure before you have a concrete problem and a few hundred real samples to test with. Many expensive decisions in this field get made before anyone has measured how fast or slow a model runs on their own data.

Related insights

See all insights