Who Is AIVISION? A Vietnamese AI Team Training Its Own LLMs

01/10/2026

Who Is AIVISION? A Vietnamese AI Team Training Its Own LLMs

One question keeps coming up whenever we meet people outside the field: what does AIVISION actually do? Are we a chatbot shop, an API reseller, another name riding the AI wave? The short answer is that AIVISION is an AI company in Vietnam that trains large language models (LLMs) on its own GPU infrastructure. This post fills in the details, including the parts that are still unfinished.

What AIVISION does, plainly

The core of the work is training and finetuning LLMs for Vietnamese. We do not take a foreign model off the shelf and wrap it in a nicer interface. We sit inside the real training loop: prepare data, train, evaluate, fix, train again. It sounds simple, but anyone who has done it knows how much time and money that loop eats.

AIVISION has released two models so far. E1.0 is a speech-to-text model for Vietnamese, turning spoken audio into written text. L1.0 is an LLM for Vietnamese. We deliberately kept the names modest, because 1.0 is exactly what they are: first versions, with plenty of work ahead.

Beyond those base models, we train and finetune LLMs for specific domains, including healthcare and pharmaceuticals. Those two deserve their own section, because their constraints look nothing like a customer-service bot.

Why own the hardware

People often ask why we do not just rent cloud GPUs by the hour. Renting is still a sensible choice for many teams and we do not argue otherwise. But when training is a daily activity rather than a one-off experiment, owning a cluster changes how you work. You dare to test a hunch, let a job run overnight, and retrain because the data was just cleaned.

Our cluster consists of 24 NVIDIA H200 GPUs and 8 NVIDIA B300 GPUs. By NVIDIA's public specifications, each H200 carries 141GB of HBM3e memory with bandwidth around 4.8TB/s, and the Blackwell Ultra generation B300 offers roughly 288GB of HBM3e. The two groups complement each other rather than compete, and we cover both in separate posts on AIVISION.

One honest note: good GPUs do not automatically produce a good model. Most of the quality lives in the data, in the evaluation, and in the patience of the people involved. Hardware only makes the trial-and-error loop faster. Anyone who says hardware is everything is probably selling hardware.

Vietnamese is its own problem

Foreign LLMs usually handle Vietnamese at a usable level, but not always naturally. You may have noticed grammatical sentences with an odd tone, formal vocabulary where plain words would do, or confusion over how to address the reader. Vietnamese has tones, a complicated system of pronouns, and many regional variations. A good Vietnamese LLM has to cope with all of that.

So we treat Vietnamese LLMs and Vietnamese speech to text as two sides of one problem: machines need to hear people speaking Vietnamese, and write and answer in Vietnamese that fits the context. We are not quoting an accuracy figure here, because a number without a dataset and a measurement method means nothing.

Healthcare and pharma: slower on purpose

When the topic is medicine, everything slows down. We finetune LLMs for healthcare and pharmaceutical use, but the role is deliberately narrow: supporting information and administrative work, such as summarizing documents, organizing information, and drafting text for a qualified person to review.

  • The model does not diagnose.
  • The model does not prescribe or advise on drug dosage.
  • The final decision always belongs to doctors, pharmacists and other medical professionals.

This is not boilerplate. A language model can write a fluent answer that is wrong, and in medicine a fluent mistake is more dangerous than a clumsy one. We design products so a human is always the last checkpoint.

Three sites, three audiences

AIVISION runs three websites: aivision.vn in Vietnamese, aivgroups.com in English for international readers, and aivision.mx for Mexico. Separate sites are not decoration. Each market has its own language, habits and needs. For markets such as the Philippines or Mexico, where local languages and accents are as varied as in Vietnam, the lessons from Vietnamese are valuable experience, although we promise nothing about a roadmap.

What we have not done

A few things are worth saying out loud. We have not published comparison benchmarks, we are not stating parameter counts, and we will not claim our model beats anyone else's until there is a public measurement rigorous enough to support it. If an AI company's page says it is number one everywhere, be skeptical.

What we can show is which road we are on: training our own models, owning our infrastructure, releasing real models, and focusing on the Vietnamese problems that general-purpose systems tend to skip. If you are an engineer, researcher or product person following Vietnam AI, the next posts go deeper: the H200 cluster, the B300 nodes, the E1.0 and L1.0 models, and why Vietnamese speech recognition is hard.

We write this blog to record the work, including what runs well and what breaks. Each post will contain something concrete you can check or push back on, and we welcome both.

Related insights

See all insights