A second voice
Modern speech to text handles noise well. Another person talking is a different problem. The engine transcribes both voices and hands your agent the mixture.
VIVA real-time models give your AI agent a clean voice to listen to and the sense of timing to hold a real conversation.
Raw microphone
As you see, those words Kumquat. keep on popping up in the Juxtapose. background. And we don't see them over here. Zucchini.
3 interfering words
After Krisp voice isolation
As you see, those words keep on popping up in the background. And we don't see them over here.
0 interfering words
Why agents break
Three things go wrong, and all three start in the audio, not in your LLM.
Modern speech to text handles noise well. Another person talking is a different problem. The engine transcribes both voices and hands your agent the mixture.
Agents wait for silence to decide the turn is over. People pause mid thought. The result is dead air on one side and talk over on the other.
To an agent, a listener saying mhm sounds like an interruption. So it stops talking when nobody asked it to.
Integration
VIVA runs first, in front of your speech to text. Everything after it inherits a clean signal.
Runs where you run
What's in VIVA SDK
Each model does one job well. Run one, or run them together.
Separates the person speaking from everything around them: competing voices, background chatter, room echo. Your STT gets one clean voice, shaped so the engine reads it the way it was trained to.
Predicts when a speaker is done, straight from raw audio. No transcription needed. Kills the awkward pause and the talk-over. Multilingual out of the box.
Classifies speech that lands mid-response: real interruption or a "mhm," takeover or backchannel. Tells your agent when to stop, and when to keep going.
Speech or silence, accurately, and robust against background noise and secondary voices, so you get fewer false triggers.
Accent, gender and synthetic-speech detection in real time. The read on a caller that humans take for granted, as a signal you can route on.
New · Voice Isolation 2.5
We measured this the way a skeptic would. 1,685 real-world recordings, three conditions, ten STT engines from seven vendors, nothing tuned to any single engine.
Two sizes
Most teams cap isolation to the calls they think need it. Lite is light enough to leave on.
New: 70% less compute
Matches the default on typical calls, with a gap only in extreme conditions. Built for high volume server fleets, edge, mobile and embedded.
Choose it when
CPU is the constraint
The default
Strongest in the hardest conditions: extreme noise, dense background voices, badly degraded streams.
Choose it when
Worst-case quality matters most
Integration
One call in front of your speech to text. Nothing downstream changes.
Tell us your use case and we'll set you up with the SDK and a trial to start.
Customers