FAQ: Vietnamese Speech to Text and LLMs from AIVISION

01/10/2026

FAQ: Vietnamese Speech to Text and LLMs from AIVISION

We get a lot of repeated questions from teams considering Vietnamese speech to text and a Vietnamese LLM for real work. Rather than answering piecemeal over email, we collect them here. Some we can answer directly, and for others we have to say plainly that it needs a case-by-case conversation. We would rather be honest than quote a pretty number that has not been measured on your problem.

About AIVISION and the models

Who is AIVISION?

AIVISION is an AI company in Vietnam. We train LLMs on a cluster of 24x NVIDIA H200 and 8x NVIDIA B300, and we do training and finetuning of LLMs for Vietnamese and for specific domains such as healthcare and pharmaceuticals. Our websites are aivision.vn for Vietnamese, aivgroups.com for English, and aivision.mx for Mexico.

What are E1.0 and L1.0?

E1.0 is a speech-to-text model, turning speech into text, and L1.0 is a large language model (LLM), both for Vietnamese. We have released both. Put simply: E1.0 listens, L1.0 understands and generates text. They can work together, for example listening to a call and then summarizing it.

How many parameters do the models have?

We do not publish parameter counts in this post. The number also rarely tells you whether a model fits your problem, since a small, well-finetuned model can suit you better than a large general one. If you need technical details to assess infrastructure, get in touch and we will discuss based on your specific needs.

What is the accuracy percentage?

We deliberately do not answer this with a single number. Speech recognition accuracy depends heavily on audio quality, regional accents, noise, domain vocabulary and how it is measured. A figure from a studio says nothing about your call center. A more trustworthy approach is to test on your own real recordings and measure together. Contact us to set that up.

About Vietnamese speech to text

What makes Vietnamese speech to text hard?

Vietnamese has six tones, and two words that differ only in tone mark can mean entirely different things. Add northern, central and southern accents, regional words, fast speakers who swallow syllables, multi-speaker conversations, and noise. In fields like medicine, drug names and rare terms make it harder still. So we focus on data and evaluation that stay close to real usage conditions.

What audio conditions give better results?

A few general lessons: place the microphone close to the speaker, limit echo, avoid heavy audio compression, and separate channels per speaker if possible. This does not depend on any one model. Sometimes moving the microphone improves results more than changing the model.

Can it handle domain-specific terms?

This is one of the reasons we do finetuning for individual domains such as healthcare and pharmaceuticals. Domain terms are where general models often slip. How much improvement you get must be measured on your data, and we do not quote numbers before that measurement exists.

About Vietnamese LLMs and finetuning

How is finetuning an LLM different from using an off-the-shelf model?

An off-the-shelf model works immediately but may misread terminology, answer in the wrong format, or sound off. Finetuning continues training on your data so the model becomes familiar with your field and way of working. In exchange you need good data and serious evaluation. Not every problem needs finetuning. Sometimes RAG or a well-written prompt is enough.

Should I use RAG or finetuning?

If the chatbot mainly looks up documents that change often, start with RAG. If the problem is the model misreading terms or producing the wrong format, finetuning is worth considering. Many good systems use both. We usually suggest measuring real errors first and then deciding.

How much does it cost?

We do not give a general price because cost depends on data volume, how much cleaning is needed, model size, infrastructure and operational requirements. Contact us to discuss your needs, and we will be clear about what is billed where.

About deployment and data

Can it run on-premise?

For many customers, especially in healthcare and finance, whether data leaves the organization is the first question. We are open to discussing deployment options, including placement inside the customer's own environment, depending on requirements and the operating capacity of both sides. The specific option needs to be discussed case by case.

How is my data handled?

The principle we follow is that customer data belongs to the customer. Details on storage, deletion timelines, anonymization and permitted use should be written into the agreement, and we advise every customer to read this part carefully with any provider, not only us.

Does the system support languages other than Vietnamese?

E1.0 and L1.0 were released for Vietnamese. We do not claim support for other languages at this time. If you need another language, let us know and we will discuss it as a separate problem.

How do I start a trial?

The most practical way is to start with a narrow task and a small amount of real data: a few dozen recordings, or a few dozen typical questions. Agree on success criteria together before running anything. That gives you concrete evidence instead of promises. More information is available at AIVISION.

Why do the answers here contain so few numbers?

Because we want to be honest. Accuracy, latency and cost all depend on your problem, data and infrastructure. A generic number sounds attractive but misleads easily, and we would rather you measure directly and judge for yourself. If you are comparing several providers, use the same test dataset for all of them, including us.

Related insights

See all insights