Your agent is only as good as what it hears.

VIVA real-time models give your AI agent a clean voice to listen to and the sense of timing to hold a real conversation.

Same STT engine on both lanes

Raw microphone

As you see, those words Kumquat. keep on popping up in the Juxtapose. background. And we don't see them over here. Zucchini.

3 interfering words

After Krisp voice isolation

As you see, those words keep on popping up in the background. And we don't see them over here.

0 interfering words

46%
fewer word errors
Real-time
15 ms latency, runs on CPU
1B min
processed monthly
Organizations worldwide trust us
Discord
Twilio
RingCentral
oVice
Daily
Dyte
Aircall
PhoneBurner
Vonage
Symphony
CarrierX
Zoho Voice
Altea Healthcare
Phound
Gather
Interact
Roam
LiveKit
Vapi
Telnyx
Tavus
Discord
Twilio
RingCentral
oVice
Daily
Dyte
Aircall
PhoneBurner
Vonage
Symphony
CarrierX
Zoho Voice
Altea Healthcare
Phound
Gather
Interact
Roam
LiveKit
Vapi
Telnyx
Tavus

Why agents break

Demos work. Real calls don't.

Three things go wrong, and all three start in the audio, not in your LLM.

A second voice

Modern speech to text handles noise well. Another person talking is a different problem. The engine transcribes both voices and hands your agent the mixture.

The wrong moment

Agents wait for silence to decide the turn is over. People pause mid thought. The result is dead air on one side and talk over on the other.

Every "mhm" is a stop

To an agent, a listener saying mhm sounds like an interruption. So it stops talking when nobody asked it to.

Integration

It sits underneath the conversation

VIVA runs first, in front of your speech to text. Everything after it inherits a clean signal.

Input Real audio
Krisp VIVA 15 ms
Transcribe STT
Reason LLM
Speak TTS

Runs where you run

What's in VIVA SDK

Five models. One integration.

Each model does one job well. Run one, or run them together.

Voice Isolation

Separates the person speaking from everything around them: competing voices, background chatter, room echo. Your STT gets one clean voice, shaped so the engine reads it the way it was trained to.

  • 46% lower WER
  • Runs on CPU
  • lite: 3.5x less compute
  • 15 ms
  • 16 kHz

New · Voice Isolation 2.5

Voice isolation your STT
can actually read.

We measured this the way a skeptic would. 1,685 real-world recordings, three conditions, ten STT engines from seven vendors, nothing tuned to any single engine.

46%
fewer word errors
15.3% → 8.2%
70%
on background voices
35.9% → 10.9%
2.15%
on harm on clean audio
vs 2.10% no processing
77.4%
on the best single engine
22.5% → 5.1%

Two sizes

Isolation on every stream,
without the CPU bill

Most teams cap isolation to the calls they think need it. Lite is light enough to leave on.

New: 70% less compute

Voice Isolation 2.5 lite

Matches the default on typical calls, with a gap only in extreme conditions. Built for high volume server fleets, edge, mobile and embedded.

Choose it when

CPU is the constraint

The default

Voice Isolation 2.5

Strongest in the hardest conditions: extreme noise, dense background voices, badly degraded streams.

Choose it when

Worst-case quality matters most

Both models
Latency 15 ms Compute Runs on CPU Bandwidth 16 kHz
How to read the lite numbers
Lite is measured on the same 1,685 recordings as the default model. On typical calls the two are within noise of each other; the gap opens only in the hardest conditions, where the full model keeps more of the target voice intact.

Integration

Drops into what you run

One call in front of your speech to text. Nothing downstream changes.

  1. 01 Connect once session = krisp.viva.connect(model="…v2.5")
  2. 02 Clean each frame, before your STT clean = session.process(frame)   # 15 ms, on CPU
  3. 03 Push it on, to any engine stt.push(clean)
Full specs and everything it runs with
Server-side, on-prem or in your own cloud. Python, Node.js, Go and iOS SDKs, with LiveKit, Pipecat and raw WebRTC pipelines supported out of the box. PCM S16LE, 16 kHz mono in and out.

Already on VIVA? It's a model swap.

Integration is unchanged. Point your pipeline at the new model.

Server SDK docs

Run it on your own calls.

Tell us your use case and we'll set you up with the SDK and a trial to start.

Customers

From voice AI teams

Guarav Agarwal
Guarav Agarwal VP of product at Twilio

“We are thrilled to partner with the Krisp team, who are leading the way in AI-based noise cancellation technology. Together we can provide exceptional audio quality to the billions of conversations on the Twilio Video platform, allowing customers to build engaging and personalized virtual experiences across telehealth, education, remote work, and more.”

Stanislav Vishnevskiy
Stanislav Vishnevskiy CTO & Co-Founder at Discord

“Our users expect world-class quality from all of their communication channels within Discord, and especially with their voice communications. Integrating Krisp’s AI-based Noise Cancellation technology into our desktop and mobile applications was seamless.”

Kumar Saurav
Kumar Saurav Co-founder & CTO at Vodex

“To ensure exceptional customer experiences as we scale, we’re excited to be integrating Krisp’s best-in-class noise cancellation technology. This partnership empowers Vodex to reach new heights in delivering clear and uninterrupted voice interactions.”

Dr. Navdeep Dhaliwal
Dr. Navdeep Dhaliwal Director of Clinical Informatics at Aarista

“By integrating Krisp’s advanced AI Noise Cancellation technology, we are not only enhancing physician dictation accuracy but also revolutionizing telemedicine communication. This partnership underscores our commitment to delivering unmatched clarity and efficiency in healthcare interactions, setting Aarista apart as a leader in transforming healthcare communication and patient care.”

Zach Koch
Zach Koch Founder and CEO at Ultravox

“Integrating Krisp eliminates one of the biggest challenges—unwanted interruptions—ensuring seamless, interruption-free conversations.”

Yukari Minai
Yukari Minai PR at oVice

“oVice customers can now enjoy more comfortable conversations in the business metaverse with access to the world's best AI noise-cancellation technology.”

Howard Lerman
Howard Lerman Founder and CEO

“Roam Inventors are working around the clock to deliver the #1 Cloud HQ meeting experience. Our new partnership with Krisp means our members will get the best available noise cancellation on the planet.”

David Casem
David Casem CEO of Telnyx

“At scale, the biggest challenge in voice AI isn't the model. It's the quality of the signal going into it. Krisp addresses that at the source, which improves everything downstream from transcription to response.”

Have questions?
We've got answers.

What's in VIVA 2.5?
Voice isolation, turn-taking and signal detectors, in one SDK. 2.5 is a big step on the isolation side, which is where the word error rate numbers come from.
Doesn't voice isolation make transcripts worse?
It can, and we measured exactly where. On the examples that were most problematic for Voice Isolation 2.1, it stripped too much and pushed word errors above leaving the audio alone. We published that. 2.5 closes most of that gap, and we published the part that remains, because it is not fully closed. On clean audio it now beats untouched, on average.
Which STT engines does it work with?
Any of them. It cleans the audio before your STT sees it. Benchmarked across ten streaming and non-streaming systems from seven vendors, including Deepgram, NVIDIA, Soniox, ElevenLabs and AssemblyAI.
Do the models need transcription or per-language setup?
No. They operate directly on the audio signal, are language agnostic, and need no per-language tuning.
How do we know the numbers are real?
Run it on your own audio. We will set up a trial on your calls. That is the only number that matters to you anyway.
We can't spare the footprint.
There's a lite model at 70% less compute, and on typical calls it matches the default. A gap may appear in extreme conditions, so if worst-case quality matters more to you than compute, take the default. If CPU is the constraint, take lite.

Run it on your own calls.

Tell us your use case and we'll set you up with the SDK and a trial on your audio.

background for toggle