AIVISION releases Speech-to-Speech AI: free, realtime, on-premise

02/10/2026

AIVISION releases Speech-to-Speech AI: free, realtime, on-premise

You say a sentence, and a moment later the AI answers back in its own voice. No typing, no reading text off a screen. That is what AIVISION has just released: a set of Speech-to-Speech models for Vietnamese, free to try at s2speech.com, running in real time and deployable on-premise inside a company's own infrastructure.

AIVISION solution demo
AIVISION Speech-to-Speech introduction video

What Speech-to-Speech means, in plain terms

A voice conversation system has to do three things in a row: listen (turn speech into text), understand (a language model works out the reply) and speak (read the answer aloud). When those three steps come from three different vendors, the delays add up and the experience feels disjointed. AIVISION's Speech-to-Speech brings all three together: Speech-to-Text, the aivision-L1.0 language model and Text-to-Speech, on one platform behind one API.

The result is a smoother conversation. Audio is streamed, so the listener starts hearing the answer before the whole reply has been spoken.

Four headline points of the release

  • Free to use: every account on s2speech.com gets $5 of free usage per day, no card required.
  • Realtime: it listens, understands and answers while the conversation is happening.
  • On-premise: companies can run the whole system on their own servers, so voice data never has to leave their infrastructure.
  • Modest hardware: it needs a single GPU with about 8GB of VRAM, in the RTX 2080 Ti class, and serves 8 concurrent conversations (8 CCU).

Why running on an 8GB GPU matters

Many companies want voice AI but hit two walls: they will not send customer recordings to an outside service, and they do not have the budget for an expensive GPU cluster. A model that needs only about 8GB of VRAM changes that. One server with a mainstream graphics card is enough to start, and capacity scales by adding machines at roughly 8 sessions per GPU.

On-premise also keeps conversation data under the company's control, which sectors like banking, healthcare and the public sector often require.

What it is for

Call centres and customer care

Answer repetitive questions in a natural voice, after hours or when lines are overloaded. The platform supports 8 kHz G.711 audio that plays directly on common telephony systems.

Voice assistants inside apps

Add spoken question-and-answer to an app, kiosk or device without stitching together several separate services.

Interpreting and multilingual conversation

s2speech.com already offers ready-made apps such as voice interpreting, live subtitles and multilingual meeting rooms. According to AIVISION, the translation and meeting apps support 24 languages, with Vietnamese as the core.

Accuracy published as numbers

Listening is the foundation of Speech-to-Speech. AIVISION reports an average word error rate of 11.84% for its speech recognition model across four Vietnamese test sets, including medical conversations and real business meetings. These are AIVISION's internal figures, measured on data not used for training, and published on s2speech.com.

Every real environment is different, with its own noise, regional accents and jargon, so the best test is still a trial on your own data before a wide rollout.

How to get started

Sign up and try it now at s2speech.com. Developers can integrate over REST and WebSocket, and the language model uses an OpenAI-compatible API.

For on-premise deployment or business enquiries, contact:

Explore AIVISION's other AI solutions at AIVISION.

Related insights

See all insights