Over three months our QA team recorded in open-plan offices, on four working call-center floors, and from the front seats of moving cars, so that a benchmark for background speech would be made of real rooms instead of synthetic mixing. Here is how it actually went.
The short answer
The secondary voice benchmark is 265 real conversations, recorded in offices, call centers and cars on the devices people actually use, then transcribed and labeled segment by segment by hand. We built it because no public benchmark covered background speech. Where a second voice is present, Krisp Voice Isolation cuts word error rate by roughly three quarters; on already-clean phone audio it costs a little more than it returns.
265
real recordings
3
scenarios
4
call-center floors
3
months of recording
What it took to record 265 real conversations
The plan fit in one sentence.
Record people talking while someone else talks nearby, in the places where that actually happens, on the devices people actually use. Then label every second of it by hand.
It took three months.
Twenty-four audio devices in the office scenario alone. Eighteen headset models across four call-center floors. Eighteen phone models, and whatever Bluetooth microphone happened to be in each car. By the end nearly everyone at Krisp had recorded something, a lot of them from the driver’s seat.
Some days the problems started before anyone pressed record. Four working call centers had to agree to let a recording team onto the floor, and every session had to fit around their shifts and their customers, not ours.
Then there were the adapters. More than once the team made it to a call center with a bag of professional headsets and a laptop ready to record, and realized the adapters that connect one to the other weren’t in the bag. Eighteen headset models means eighteen different connectors, and it only takes one missing adapter to lose the session.
Even with everything plugged in, plenty of takes never made it into the dataset. The second voice had to come in on cue, talk over the primary, keep going after the primary stopped, then fall silent, and a speaker who started early or quit too soon erased the window we cared about most. Volume was just as fussy: too quiet and it didn’t count as a second voice, too loud and it buried the primary. And some takes were lost only afterwards, when a muted mic or the wrong input turned a session that felt fine in the room into silence on the drive.
The phone calls had their own kind of chaos. A car is not a studio. The radio comes on, the kids in the back have opinions, a passenger joins the conversation. Anna asked nearly everyone at Krisp to record scripts from their cars, then started worrying they’d focus on reading and forget they were driving. So every request came with the same line: eyes on the road, the script can wait. By the team’s count she sent it 41 times. Nobody crashed, and the recordings are in the dataset.
We tell these stories because they are the argument. Every recording here was made by a real person, in a real room, on a real device. That’s the only way to find out what speech-to-text does when the room isn’t quiet, and it’s the part of this dataset that can’t be synthesized.
Why we didn’t just simulate it
The usual way to build an evaluation set is to take clean speech, add a noise recording, convolve it with an impulse response, and call it a room. That works well enough for training. It does not work for evaluation.
The thing you most need to measure is the degradation you did not think to simulate.
Background noise and overlapping voices have long been among the primary obstacles in human communication. Anybody who has tried to hold a conversation in an open-space office or on a busy street knows the feeling: you raise your voice, you repeat yourself, you give up and text instead.
The problem has quietly grown wider. It is no longer just human to human. Millions of conversations now happen between a person and a bot, and when a secondary voice leaks into that channel the speech-to-text model transcribes it alongside the intended speaker. The transcript reaches the LLM, the LLM reads words that were never meant for it, and its behaviour shifts. A customer says “I’d like to cancel my subscription” while a colleague in the next seat says “Can you grab lunch?” and the bot hears both.
Modern speech-to-text has made real strides on noise robustness. Traffic, air conditioning, keyboard clatter: handled. Background speech is a different matter. These models are designed to capture speech robustly, not to judge the importance of one voice over another. A background voice shares the target speaker’s acoustic structure in a way that ordinary noise does not, so the model transcribes every voice indiscriminately rather than dismissing the secondary one.
We went looking for a benchmark that mirrored this. There are good public benchmarks for ambient noise, reverberation and accented speech. For the background-voice problem, the one that happens in every open-plan office, there was a gap.
Degraded audio is a second gap. Bluetooth car microphones, low-bandwidth telephony codecs, a speaker halfway across the room: the signal arrives already compromised, so an isolation model has a much tighter margin. It has to clean the audio without further damaging something fragile. We could not find benchmarking data for that either.
So we recorded our own.
Three scenarios, three realities
We settled on three categories, each reflecting a distinct environment where voice isolation matters.
Work
Working from an open-space office, from home, or from a meeting room. The most universal scenario, where secondary speakers are colleagues, family members, or whoever else is sharing the space. Recorded across small rooms, medium rooms, large conference rooms and open-plan floors, on everything from professional Jabra headsets to consumer earbuds.
29
speakers
96
recordings
24
devices
4
room types
17
scripts
Call center
Collected by visiting four real call-center locations. Many agents talk in parallel, but one of them sits closest to our primary speaker’s microphone. Background voices here are not incidental. They are constant and structured, each carrying its own conversation with its own customer. Scripts spanned banking, insurance, healthcare, telecom and more.
7
speakers
106
recordings
18
headsets
4
locations
28
scripts
Phone calls
A different kind of problem. Almost no deliberate secondary speaker here; the challenge is the audio itself. Most recordings came from cars, calling over a Bluetooth car microphone, with the phone on speaker far from the mouth, or through earbuds while driving. We also captured calls from noisy buildings, shops, streets and metro stations. Street noise, music, cabin rumble, crowd chatter and telephony codecs all degrade the signal before isolation ever sees it.
26
speakers
63
recordings
18
phone models
4
location types
20
scripts
In the Work and Call Center scenarios, different background noise types appear naturally alongside the secondary speakers: office chatter, keyboards, music, babble. That natural variability is the part synthetic mixing cannot replicate.
Labeling every segment
Each recording ships with a hand-written transcript and a JSON metadata file carrying the transcript text, speaker gender, recording device and recording environment. Transcriptions capture speech exactly as it occurs, including hesitations, false starts and self-corrections, so the ground truth reflects what was actually said rather than a cleaned-up version.
Beyond transcription, every audio segment was labeled by hand with one of four tags:
primaryMainly the primary voice. Some ambient or indistinct babble may still be present.
mixBoth primary and secondary speakers audible.
secondaryOnly the secondary speaker audible.
noiseNo speech in the segment, only noise.
These segment-level annotations are what make fine-grained analysis possible. Without them you can only score a recording as a whole, which hides the case that matters most: a model that correctly outputs silence and a model that quietly deletes the primary speaker can post the same overall number.
What the recordings showed
That part is a separate write-up, with the full results across eleven speech-to-text configurations and four Voice Isolation models. The short version: where a second voice is present, word error rate falls by roughly three quarters. On already-clean phone audio, isolation costs a little more than it returns, and we published that too.
Read next
STT handles noise now. It still can't handle a background voice.
The results: 265 recordings through 11 speech-to-text configurations, with and without Krisp Voice Isolation. Includes the condition where we make things worse.
This dataset exists because the QA team went and made it. They organised sessions inside working call centers, brought in readers of different nationalities and accents so the recordings would sound like real customers and real agents rather than one office in one city, carried the gear, chased the adapters, and labeled every segment afterwards.
The dataset, the transcripts, the segment labels and every model’s output on every file are published on Hugging Face. If you use it, we would like to hear what you find.