Multilingual AI: Lessons from a Vietnamese LLM for Mexico

01/10/2026

Multilingual AI: Lessons from a Vietnamese LLM for Mexico

Building an LLM for Vietnamese taught me many things English-language documentation rarely mentions. The language has six tones, writing full of diacritics, users who type without accents when in a hurry, and regions that use different words for the same object. When I look toward markets like Mexico or the Philippines, I see these problems do not disappear. They change shape. This post discusses multilingual AI as a problem and a general approach, not as an announcement about any particular model.

To be clear from the start: AIVISION's current models, E1.0 and L1.0, were released for Vietnamese. What follows is thinking about method, drawn from our Vietnamese work, and not a claim about supporting any other language.

What Vietnamese taught us

Three big lessons. First, high-quality data in a lower-resource language is scarcer than you expect, especially specialized data like medical or pharmaceutical text. Second, text normalization takes more effort than anyone plans for: character encoding, misplaced diacritics, abbreviations, social media slang. Third, evaluation. Benchmark scores translated from English do not tell you whether a model understands how real people talk.

Voice adds another layer. A Vietnamese speech to text system must cope with northern, central and southern accents, call-center noise, and speakers who interrupt each other. Every new language brings a similar set of problems, differing in the details.

Spanish in Mexico: same language, different market

Many assume Spanish has plenty of resources so the problem is easy. The true part is that general data is abundant. The underestimated part is that Mexican Spanish has vocabulary, forms of address, idioms and customer-service tone that differ from Spain or Argentina. A support chatbot using a stiff European register sounds off immediately. Add the mixing of English in business settings, especially near the border.

I think the Vietnamese lessons transfer: do not treat Spanish as one block. Decide the target variety early, collect data for exactly that variety, and have native speakers from the right region review it. A reader from Spain may not notice what feels awkward to someone in Mexico.

The Philippines: many languages and code-switching

The Philippines is complicated in a different way. There is Filipino and many regional languages, while English is widely used at work. Users often switch between English and a local language within a single sentence, which is called code-switching. For speech recognition this is nasty: the model has to decide which language each word belongs to mid-sentence.

Vietnamese has its own version, when tech workers mix English terms into everyday speech, like asking whether the deploy is done. Experience says you cannot handle this by treating every sentence as monolingual. Training data must reflect how people actually speak, even when it is nonstandard by textbook rules.

Technical problems that repeat in every language

  • Tokenizers. Tokenizers designed around English often shatter Vietnamese, and many other languages, into too many fragments. The result is more tokens, slower and costlier inference, and sometimes weaker understanding.
  • Domain data. Medical, legal and financial vocabulary in each language has to be sourced and verified separately. Machine translation from English easily produces terms that sound right but nobody uses.
  • Dialects and accents. Pooling everything together tends to average the model out and make it mediocre everywhere.
  • Safety and culture. What is sensitive in one country may be ordinary in another. Safety filters translated word for word often misfire.

Evaluating without being fooled by pretty numbers

The mistake I see most is measuring a multilingual model with machine-translated English test sets and concluding that it understands the target language. Translated test sets carry English sentence structure and cultural context. More trustworthy is a native test set: real questions from real users, written by native speakers, in the right domain, including typos and everyday phrasing.

The same goes for speech recognition: you need real recordings from the actual usage environment, the right accent groups, the right noise level. A clean studio gives lovely results that do not survive contact with reality.

How we think about expansion

AIVISION is an AI company in Vietnam that trains LLMs on a cluster of 24x NVIDIA H200 and 8x NVIDIA B300, has released the E1.0 speech-to-text model and the L1.0 LLM for Vietnamese, and does training and finetuning for domains such as healthcare and pharmaceuticals. We run aivision.vn for Vietnam, aivgroups.com for international readers, and aivision.mx for Mexico, so localization is a real question in our work.

The approach we lean toward is to go market by market with discipline instead of announcing support for a hundred languages. Concretely: define the domain and language variety to serve, work with native speakers to build data and test sets, measure on real tasks, then expand. It is slower, and in return you know what the model can and cannot do.

What I would ask if you build AI for a new market

  • Which language variety, in which region, for which user group?
  • Where does the domain data come from legally, and who verifies it?
  • Which native speakers take part in evaluation, and do they represent real users?
  • Do users code-switch or abbreviate, and does your data reflect that?
  • What do that country's data protection rules require about storage location?

Answering these five before writing the first line of code often saves months. Language is not just letters. It is habit and context, and a model is only good when the data places it in the right context.

Related insights

See all insights