Why Vietnamese Speech Recognition Is Hard: Accents, Noise, Jargon

01/10/2026

Why Vietnamese Speech Recognition Is Hard: Accents, Noise, Jargon

A person from Nghe An says one thing. A person from Saigon says another. Both are speaking Vietnamese, and to a listener neither sounds strange, but for a Vietnamese speech recognition system they are two very different problems. We at AIVISION build the E1.0 speech-to-text model and have seen quite a variety of errors. This post walks through the three biggest groups of difficulty, with an insider's opinion attached.

Regional accents: one language, many sound systems

Vietnamese has six tones, and how they are pronounced varies by region. In parts of central Vietnam, certain tones sound quite different from standard northern or southern speech. Initial consonants differ too: some areas distinguish tr from ch and s from x, others do not. Some regions pronounce d, gi and r almost the same, while others keep three distinct sounds.

Human listeners compensate with context. We hear a sound and guess the word from the sentence. A machine trained mostly on one accent will have higher error rates on the others. That is not an algorithm flaw but a data diversity flaw.

The real difficulty is collecting data for less common accents. You need real speakers, real recordings, and people who listen back and write every word down. It is expensive, slow, and biased toward regions that are easy to reach. We do not claim E1.0 has handled every accent, since no public measurement lets us say so. It is an area of continuous improvement.

One more thing: people do not have a single voice. The same person talks differently in casual conversation and formal speech, tired and alert. Many Vietnamese living far from home mix their native accent with the one where they now live, producing a hybrid that no regional chart captures.

Noise: the spoiler outside the lab

Most speech recognition models look fine in clean conditions. Trouble starts in the real world. A motorbike passes, a fan hums, air conditioning, dishes clattering in a cafe, someone else cutting in. For Vietnamese, noise is especially harmful because many important distinctions lie in small acoustic details like the tail of a tone or a final consonant. Cover those details and two words with different meanings become one.

A situation product teams run into often: a phone placed on a meeting table three or four meters from the speaker. Sound bounces off walls and people sitting close drown out people sitting far. The result is poor, not because the model is bad but because the signal was damaged before it reached the model. Sometimes the most effective advice is very ordinary: put the microphone closer to the speaker.

Sometimes noise suppression before recognition makes results worse. An aggressive denoiser can erase the small details of speech along with the noise. It is a subtle trade-off that few outside the field think about.

Jargon and proper names

This group of errors irritates users the most, because it hits the very words that matter. A meeting can be transcribed almost perfectly except for the client's name, the project name and a few internal abbreviations.

The reason is easy to understand: these words are rare in training data. A model is used to everyday speech and tends to replace an unfamiliar word with a familiar one that sounds close. A drug ingredient, a legal procedure, a technical part: all easily turn into similar-sounding common words.

Take pharmacy as an example. Drug names are long, of foreign origin, and pronounced in inconsistent Vietnamized ways: one person reads it English-style, another French-style, a third in the habit of the pharmacy counter. A recognizer must face all those readings. This is exactly why we stress that in healthcare and pharma, speech to text should only support administrative note-taking, and a qualified person must check all content. The system does not diagnose, prescribe or advise on dosage.

Common approaches include:

  • Adding domain vocabulary to training data or during finetuning.
  • Placing a language model behind the recognizer to fix errors from context, for example a Vietnamese LLM that rereads the transcript and corrects nonsense.
  • Letting users supply lists of proper names and terms to prioritize.

None of these is magic and each has its own price. Using an LLM for correction can make text prettier while drifting from what the speaker actually said, a risk to take seriously when the transcript serves as a record.

Difficulties rarely mentioned

Beyond the three big groups, a few more things trouble the team. Vietnamese speakers often mix in English mid-sentence: deadline, meeting, feedback, even tech product names. The boundary between Vietnamese and English inside one sentence is something single-language recognizers handle clumsily. Then there is text normalization: spoken as two thousand two hundred, written as 2,200 or 2200? How are dates written? We have no single answer, because each user has its own conventions.

The same problem exists in markets AIVISION follows, such as Mexico and the Philippines, where people routinely mix languages in the same sentence. The shared lesson: no model solves everything, and knowing exactly where a model fails matters as much as making it better.

Related insights

See all insights