What a 24-GPU H200 Cluster Does When Training an LLM
01/10/2026

Imagine copying a huge library by hand in a single night. One person will never finish, so you call in helpers and give each a few shelves. Then new problems appear: who keeps the master copy, who checks mistakes, how do people share notes without losing half the evening. Training a large language model on AIVISION's 24-GPU H200 cluster is much like that, except the copyists are chips and the pages are billions of numbers.
What an H200 is, without the jargon
The H200 is an NVIDIA GPU built for AI computing at scale. Two public figures are worth remembering: each card has 141GB of HBM3e memory, and memory bandwidth of about 4.8TB/s. Think of memory as the desk and bandwidth as the width of the aisle that brings documents to the desk. A bigger desk lets you spread out more at once, and a wider aisle means less standing around.
For LLMs, waiting is the main enemy. Much of a chip's time during training is spent not computing but waiting for data to arrive from memory. That is why bandwidth matters almost as much as raw compute.
Add it up and 24 GPUs at 141GB each give roughly 3.3TB of GPU memory across the cluster. That is simple multiplication from public specs, not our own measurement, and it does not tell you how large a model fits, because memory also holds much more than the model itself.
Why one GPU is not enough
During training, memory does not just hold the model weights. It also holds intermediate results from every pass, gradients used to update the model, and optimizer state. Together these need several times more memory than the model alone. No matter how strong a single GPU is, it hits the ceiling quickly.
The answer is to split the work. There are several ways, and engineers usually combine them:
- Split the data: each GPU holds a copy of the model and processes a different batch, then the group merges results.
- Split the model: spread layers or large operations across GPUs when one cannot hold the whole thing.
- Split the training state: scatter gradients and optimizer state across machines so each carries less.
All of these share one cost: the GPUs must talk to each other constantly. The network between them matters as much as the chips. A cluster with great GPUs and a slow network is like a brilliant team of copyists who can only communicate by post.
What a training run really looks like
Honestly, most of an engineer's time is not spent admiring a falling loss curve. A long run tells a messier story. Some days a node hangs on a small fault, the whole job stops, and it has to reload from the latest checkpoint. Some days the loss suddenly spikes and half a morning goes into working out whether a bad batch of data or a learning rate set too high is to blame.
That is why checkpointing, saving state periodically, is a survival habit. Save often and you pay in disk time; save rarely and you lose more when something breaks. It is a classic trade-off with no universal answer. Every team has to find the balance for its own system.
Another topic rarely mentioned is running cost. A 24-GPU cluster draws serious power and gives off serious heat, which brings cooling, backup power and monitoring into play. We will not quote figures here, but anyone who thinks owning GPUs means buying them and switching them on should talk to someone who runs a machine room.
What the H200s do at AIVISION
The H200 cluster is the heavy lifting for training and finetuning Vietnamese LLMs. Continued training on Vietnamese data, domain finetuning for areas like healthcare and pharma, and rerunning experiments after changing how data is processed all need many hours of continuous GPU time.
We are not stating which model trained for how long or on how many tokens, since that information is unpublished. What we can say in general: for LLM finetuning, the advantage of a private cluster is agency. If the data team finishes cleaning a batch of medical-administrative documents in the afternoon, a trial run can start that evening and the results are waiting in the morning. That short feedback loop is worth more than people expect.
An important note on medicine: even when we finetune for healthcare and pharma, the aim is information and administrative support. The model does not diagnose, prescribe or advise on dosage, and medical expertise always makes the final call.
Trade-offs worth knowing
A big cluster is not always the answer. For small experiments, coordinating 24 GPUs costs more effort than running on a few. Sometimes an idea should be tested at small scale first, and only the survivors move to the big cluster.
Also, AIVISION's H200 and B300 clusters serve different parts of the work. The B300, with roughly 288GB of HBM3e per GPU, has its own uses, covered in our post on Blackwell Ultra. Dividing work correctly between the two is part of the craft.
If you are weighing whether to build your own infrastructure or rent it for a Vietnam AI project, our most practical advice is to measure real demand first. How many GPU hours per month, steady or bursty, and whether you have people to operate it. Those three questions decide most of it, and none concerns which chip is currently advertised the most.
You can read more about how we work at AIVISION.