On-Premise Call Center Voice AI on One 8 GB GPU: Needs and Pitfalls
05/10/2026

“Customer recordings do not leave this building.” Imagine that line landing in a Monday meeting, right after the customer care team has presented its voicebot plan. The person saying it heads compliance, and she is not joking. The problem changes shape on the spot: it is no longer about which cloud service to pick, but how to run on-premise call center voice AI inside the company's own server room.
The good news is that this is lighter than many people expect. The s2speech Speech-to-Speech release is free to use, runs in real time, can be deployed on-premise and fits on a single GPU with roughly 8 GB of VRAM, RTX 2080 Ti class, handling 8 concurrent conversations. The less cheerful news: “it installs” and “it runs smoothly every single day” are two different things. Here is what to prepare, and where teams tend to trip.
What needs to be on the table
- A server with a GPU of about 8 GB VRAM. No expensive cluster required; a card in the RTX 2080 Ti class is enough for the Speech-to-Speech release. That said, give that GPU this one job and don't make it share memory with other workloads.
- A link to your phone system. The s2speech Text-to-Speech service already offers 8 kHz G.711 μ-law and A-law output for Asterisk and FreeSWITCH phone systems. For the on-premise release, ask about input and output audio formats up front so you don't end up transcoding in the middle.
- Scripts and business data. A voicebot only answers well when it is connected to wherever the answers live: appointment calendars, order status, return policies. This is usually the most time-consuming part, and no GPU can do it for you.
- Someone who owns the system. Who watches the logs, who listens to failed calls, who fixes the script. Without that role, projects tend to die quietly after a few weeks.

8 concurrent sessions, read correctly
8 CCU means that at any single moment, that GPU handles up to 8 conversations. Not 8 calls a day, and not 8 virtual agents sitting around waiting. If a ninth call arrives while all 8 sessions are busy, it has to go somewhere: into a queue, to a human agent, or to another server.
So before buying anything else, pull your PBX logs and look at what your peak hours really look like. Some lines are quiet all day and only spike first thing in the morning and right after lunch. Others hum along steadily from morning to night. The number of simultaneous calls at the peak is what decides how many GPUs you need, not the total number of calls in a month.
A practical tip: don't design the voicebot to swallow everything. Give it the structured call types, and when sessions are full or a caller asks for a person, route straight to the agent team. Callers don't care who picks up. They care whether anyone picks up at all.
Where things go wrong in production
Phone audio is not demo audio
Demos usually happen with a good microphone in a quiet room. Real calls travel over narrowband phone lines, with motorbikes, children and speakerphones in the background. Test with your own PBX recordings. Even AIVISION's published internal figures show a gap between read speech and real-world audio: a word error rate of 4.58% on FLEURS-vi, but 20.29% on real business meetings, measured on data not used for training.
Numbers, money and dates
s2speech Speech-to-Text is built to handle numbers, amounts, dates and Vietnamese mixed with English. Even so, the script should make the voicebot read back order codes, amounts and appointment dates for the caller to confirm. A quick “let me read that back to you” costs a few seconds and can head off an entire complaint.
Operations after launch day
On-premise means you take care of things a cloud provider would normally handle: GPU temperature and memory monitoring, configuration backups, version updates, and a plan for when the server restarts in the middle of a shift. Always keep a way out: if the AI box stops responding, the PBX should automatically route calls back to people.
Data and access
Keeping data in-house is the main reason to go on-premise, so don't let it leak out the back door. Who can listen to recordings, how long transcripts are kept, whether logs mask customers' account numbers. Those questions deserve written answers before go-live.
A lean pilot plan
If it still sounds like a fit, a lean plan could look like this. Prototype the script through the platform APIs using phrases your own team records or synthetic data, with no real customer recordings yet. According to s2speech.com, every account currently gets $10 of free usage per day with no card required, so this stage can start before anyone has to ask for budget. Once the script is solid, move to the on-premise Speech-to-Speech release, run it alongside human agents on a low-risk call line, listen to failed calls every day, and only scale when the share of calls handed to people has dropped and settled.
Details on each service are on the s2speech product page, and you can create a trial account at s2speech.com. The AIVISION team has deployed voice AI in Vietnam, the USA, Mexico, the Philippines and Thailand, so questions like “our PBX has an unusual setup” have somewhere to go.
And the head of compliance? She doesn't need to know what VRAM is. She only needs to look at the diagram and see that no data arrow crosses the company wall. With on-premise, that is a diagram you can actually draw. Just ask AIVISION to confirm one thing in writing: the on-premise build runs entirely inside your network with no outbound calls.