Build a Vietnamese Voice App: STT, LLM and TTS Behind One API Key
05/10/2026

The request sounds simple: a customer opens the app, says something like “can I come by at eight tomorrow for an oil change?”, and the app understands, holds a slot and answers out loud. No forms, no dropdowns. Suppose you are the only developer at a small chain of motorbike repair shops and that request just landed on your desk. It sounds like a big project, but a Vietnamese voice app like this is really just three API calls chained together: listen, think, speak.
What makes this easier on s2speech is that all three steps live under one account, one API key and one balance: Speech-to-Text, the aivision-L1.0 language model and Text-to-Speech. One place to manage access, one place to watch costs. What follows is the path we would suggest to the person who will actually write the code, opinions included, not just a feature list.
Before you write a single line of code
Go record some audio. Not your own voice reading aloud at your desk, but real customers: the fast talker, the older man who hesitates, the office worker who drops “book” and “confirm” into the middle of a Vietnamese sentence, people from the central provinces and the Mekong Delta. Collect a few dozen sentences like that into a small test set. Every decision that follows, from how you listen to how you prompt, should be checked against it.
Then create an account and get your API key. According to s2speech.com, every account currently comes with $10 of free usage per day and no card required, so the testing phase can start right away without topping up. When you do need credit, top-ups are by bank transfer, starting from 50,000 VND.

Listening: streaming or file?
s2speech Speech-to-Text has two doors. Walking through the wrong one is behind a lot of apps that “work, but feel slow”.
- Streaming over WebSocket for live conversation. Audio goes up in small chunks and text comes back while the user is still talking. Use it when the app needs to respond instantly, like a booking assistant.
- REST for files or URLs when users send voice messages, or when you process recordings after a call. Simpler, and easier to retry when the network drops.
One detail people often overlook: results include word-level timestamps. They are more useful than you might think, from highlighting the word being spoken in subtitles to jumping straight to the moment a customer said “reschedule” when you need to listen again. The STT is also built to handle Vietnamese mixed with English, numbers, money, dates and a range of regional accents. But don't take anyone's word for it, ours included: run your own test set and look at the results yourself.
Thinking: give the model a frame
aivision-L1.0 has an OpenAI-compatible API. If your backend already uses a familiar client library, most of the work is changing the base URL, the API key and the model name. The model handles summaries, question answering, meeting minutes, translation and conversation analysis, supports streaming, and works in both Vietnamese and English.
For a booking app, don't let the model chat freely. A few rules keep things from drifting:
- Keep the system prompt short and specific: which hours the shop is open, which services it offers, and when to ask the customer again.
- Ask the model for two things: structured data (service, date, time) for the booking logic, and one short sentence to read back to the customer. Have the backend validate that data before it touches the real calendar.
- Turn on streaming so the reply flows into the speaking step as soon as the first words are ready.
- Keep answers short. Listening is not reading: three long sentences on a screen are fine, but through a phone speaker the customer has forgotten the opening one by the time the third ends.
Speaking: a voice that starts on the first words
s2speech Text-to-Speech offers natural Vietnamese and English voices that stream from the very first words. Combined with streaming from the model, the customer hears the answer while its tail end is still being generated. If the app ever connects to a phone system, 8 kHz G.711 μ-law/A-law output is already there for Asterisk and FreeSWITCH. As for cloning a voice from 20 seconds to 2 minutes of recording, do it only with the voice owner's consent. That is a condition, not an option.
Three classic mistakes in a first demo
- Putting the API key inside the mobile app. Don't. The key lives on your backend, and the app only talks to your server. A leaked key turns your balance into someone else's.
- Not reading back important details. “So that's eight tomorrow morning for an oil change, correct?” One confirmation line is far cheaper than a customer showing up at the wrong time.
- Only testing in an air-conditioned office. Go stand by the road, turn on a fan, play some music. According to AIVISION's published internal figures, the average word error rate across four Vietnamese test sets is 11.84%, from a low of 4.58% on FLEURS-vi read speech up to 20.29% on real business meetings. The messier the environment, the more your app needs to be designed to recover from mistakes.
The list of services and ready-made apps is on the s2speech page at aivgroups.com; to create an account, grab a key and start testing, go to s2speech.com. And if all you need is to record, transcribe and produce minutes without writing code yet, AI Voice Note is a ready-made app that does exactly that, with question answering over what was recorded.
As for the repair shop? Version one does not need to be clever. It just needs to hear the appointment time correctly, ask again when it is unsure, and never promise an oil change after the shop has closed.