Quantization, Inference and GPU Cost When Deploying LLMs
01/10/2026

The model runs beautifully on an engineer's machine. Then it goes to production, crawls, and the end-of-month GPU bill makes the finance team wince. I have seen this script repeat across many teams building a Vietnamese LLM. The root cause is usually three concepts few people explain well: quantization, inference, and how GPUs are billed. I will try to describe them for someone who is not deep in the technical weeds.
What inference is, and why it becomes the long-term cost
Training is when the model learns. Inference is when it works: receiving a question and generating an answer. Training is expensive but happens once, or a few times. Inference happens every time a user shows up, every day, every hour. For a product with real users, total inference spend often passes training spend after a long enough period.
One commonly misunderstood point: the model generates text one token at a time. Each new token requires reading almost all of the model's weights from GPU memory. So response speed is limited by memory bandwidth more than by raw compute. Once you grasp this, it becomes clear why shrinking the model helps it run faster.
Quantization: compressing a model like compressing a photo
A model's weights are billions of numbers. By default each one is stored at high precision, taking 16 bits. Quantization stores them with fewer bits, typically 8 or 4. It is like compressing a photograph: the file gets much smaller, it looks nearly the same to the naked eye, but zoom in and a few details look smudged.
The benefits are very practical. The model takes less memory, so it fits on cheaper or fewer GPUs. Weights load faster, so generation speed usually improves. The price is that quality may drop, and by how much depends on the model, the compression method and the task. For simple question answering, the difference is sometimes negligible. For multi-step reasoning, arithmetic, or specialized text with rare terms, I find quality is more easily affected.
There is no magic number that applies everywhere. The only reliable way is to run the compressed version against your own real questions and compare it with the original. For Vietnamese, check cases with tricky diacritics, proper names and figures separately, because errors tend to show up there.
How GPU cost is built up
Setting detail aside, the cost of running an LLM service has a few parts:
- Hardware: buying or renting GPUs. Memory size decides which models fit.
- Run time: rented GPUs are billed by the hour, so leaving them on all day costs money even if only a couple of people ask something at midnight.
- Utilization: a GPU running at 10% means the other 90% is burned money.
- People: those who monitor, tune and handle incidents, often forgotten in budgets.
The question people skip is how many users hit the service at once during peak hours. An internal system for two hundred employees has a very different peak from a consumer-facing app. Design for the peak and you waste money in quiet hours. Design for the average and you choke when it gets busy.
Levers for lowering cost
In my experience the order of consideration runs from cheap and low-risk to more laborious:
- Pick a model that is just big enough. A small model finetuned for the exact task can beat a large general one, and cost far less.
- Quantization. Cuts memory and speeds things up, but you must measure quality.
- Batching. Process several requests together so the GPU is not idle. The trade-off is that each user may wait slightly longer.
- Caching. If many people ask the same thing, or share the same opening instructions, do not compute it from scratch every time.
- Length limits. Long prompts and answers cost money per token. Many systems stuff a full page of instructions into every call and nobody remembers why.
| Lever | Benefit | Trade-off |
|---|---|---|
| Smaller model | Cheaper, faster | Needs good finetuning, weaker general ability |
| Quantization | Less memory, faster | Quality may drop, needs testing |
| Batching | Efficient GPU use | Per-request latency may rise |
| Caching | Less repeated compute | More complexity when content changes |
A story from speech to text
Cost shows up even more clearly in voice workloads. A call center with thousands of calls a day needs them transcribed. If you need the text while the call is happening, you must stream and accept GPUs that stay on. If you only need transcripts afterward for quality review, you can batch them overnight, much cheaper. With the same Vietnamese speech to text model, two business requirements produce two different bills. Ask how fast the business truly needs it before buying hardware.
What to measure before deciding
- Time to first token: how long a user waits before seeing the first word.
- Generation speed while running.
- How many concurrent requests the system handles before slowing down.
- Cost per thousand requests, based on your real traffic.
- Quality on your own test set after every compression or configuration change.
AIVISION is an AI company in Vietnam that trains LLMs on a cluster of 24x NVIDIA H200 and 8x NVIDIA B300, and has released the E1.0 speech-to-text model and the L1.0 LLM for Vietnamese. We care about deployment because a great model that cannot run within a customer's budget has solved nothing. You can learn more at AIVISION.
One habit I want every team to have: record cost per request from the very first prototype. The number looks harmless at first, but it tells you whether the product can survive when users grow tenfold. Finding out early beats finding out when the invoice arrives.