Every call that goes through Krisp Voice Isolation costs a small slice of CPU time. When you run thousands of calls at once, those slices are your server bill. So we asked a simple question: how many live calls can one server clean, and what is the cheapest way to raise that number?
The answer surprised us. The biggest win was not a new model. It was one setting. On the right CPUs, it lets the same model handle about 2× more calls. On a single 16-vCPU AWS Graviton3 machine, our full model went from 76 live calls to 137.
The short answer
Turning on FP16 with one SDK flag, enableLowPrecisionExecution, lets the same Voice Isolation model hold about 2× the live calls per server on AWS Graviton3 and Intel Sapphire Rapids, with no loss in quality. On Intel Ice Lake, which has no FP16 hardware, it gives nothing.
≈2×
live calls per server on FP16 CPUs
137
live calls on one 16-vCPU Graviton3, from 76
1.48 ms
to clean a 20 ms frame, from 3.00
1 flag
no model or pipeline changes
This post explains what we changed, how we measured it, and where it does not help. If you run Krisp Voice Isolation at scale, the section on what this means for you tells you exactly what to do.
The terms this post relies on
FP16A number stored in 16 bits instead of the usual 32. Half the size, still more precision than voice isolation needs.
Real-time factor (RTF)Processing time divided by audio time. RTF 0.5 means a 20 ms frame takes 10 ms to clean. Lower is better.
vCPUThe unit cloud providers rent. On Graviton3 it is a whole physical core; on Intel it is one of two hyperthreads sharing a core.
Max concurrent streamsThe most live calls one server can clean at once without breaking the CPU, RTF or dropped-frame limits.
The problem: every call needs a slice of the CPU
Voice Isolation works on tiny pieces of audio, 20 milliseconds at a time. Each piece has to be cleaned before the next one arrives. For one call, this is easy. For 100 calls on one server, the CPU has to clean 100 pieces every 20 milliseconds.
At some point the CPU cannot keep up. Audio arrives faster than it is cleaned. Frames get dropped, and callers hear glitches. The number of calls a server can hold just before that point is its real capacity. Everything in this post is about pushing that number up.
Here is the budget for one core. In a single 20 ms window, 6 frames fit at 3.00 ms each with FP16 off. With FP16 on, at 1.48 ms each, 13 fit. That is single-stream p95 latency for Krisp VI Balanced on a 16-vCPU Graviton3; real capacity also depends on the CPU and RTF limits below.
Half the time per frame: 6 frames fit in a 20 ms window with FP16 off, 13 with FP16 on. Single-stream p95 latency, Krisp VI Balanced, 16-vCPU AWS Graviton3
What we changed: half-size numbers
A neural network is mostly multiplication. Our models store their weights and the audio features they work on as float32 numbers: 32 bits each, about 7 decimal digits of precision, and a huge range. That is far more precision than voice isolation needs.
FP16 is the same kind of number packed into 16 bits. This is the whole change. It is not a different algorithm and not a smaller model. It is the same network, with the same layers, storing and multiplying half as many bits.
Two good things follow from that:
Less memory traffic. Weights and intermediate results are half the size, so each frame moves half as much data between RAM and the CPU. These models are often limited by how fast data moves, not by how fast the CPU can multiply. Cutting the data in half is a big part of the speedup.
More math per instruction. Modern CPUs process many numbers with one instruction. A 512-bit register on an Intel chip holds 16 float32 numbers or 32 FP16 numbers. Same instruction, same register, twice the work. Arm NEON registers are smaller (128 bits, so 4 or 8 numbers), but the ratio is the same: 2×.
Here is the second effect in one line of CPU code. The register is the same 512 bits either way; halving each number doubles the lanes.
Same 512-bit register, same instruction: 16 float32 numbers or 32 FP16 numbers.
One 512-bit register, one instruction
Number type
Instruction
Width of each number
Additions per instruction
float32
vaddps zmm0, zmm1, zmm2
32 bits
16
FP16
vaddph zmm0, zmm1, zmm2
16 bits
32
There is one catch. The CPU must have hardware that can do FP16 math natively. Intel Sapphire Rapids (AVX-512 FP16) and AWS Graviton3 (Arm NEON FP16) have it. Older chips like Intel Ice Lake do not. On those, FP16 gives you nothing.
In the Krisp SDK this is a single flag: enableLowPrecisionExecution. When it is on, the SDK checks whether the CPU supports FP16 arithmetic. If yes, the model weights are loaded as float32, converted to FP16, and inference runs in FP16. If not, everything stays on the float32 path. You do not have to change the model, the audio pipeline, or anything else.
And most importantly, FP16 doesn’t degrade the quality.
How we measured it
We did not measure speed on one call. We measured how many calls a server can hold at once, because that is what decides your bill.
Run on one instanceVoice Isolation at 100% suppression, in 20 ms frames, on one cloud instance.
Add live streamsKeep adding concurrent streams, one load level at a time.
Stop at the first broken limitStop when CPU, RTF or dropped frames crosses its limit, and report the last count that held all three.
The three limits a stream count must hold
Limit
Value
Why it matters
CPU utilization (p95)
≤ 85%
Leaves headroom for traffic spikes and the rest of your stack
Real-time factor (p95)
≤ 0.85
Audio must be cleaned faster than it arrives, with margin
Dropped frames
≤ 1%
Callers should not hear glitches
For a single stream of our full model with FP16 on a 16-vCPU Graviton3, RTF is 0.074: a 20 ms frame is cleaned in about 1.5 ms. Every stream you add eats into that headroom until p95 RTF crosses 0.85.
We tested three models. Krisp VI Balanced and Krisp VI Default are the full Voice Isolation models; Krisp VI Lite is the lightweight one. Each ran with FP16 off and on, on three CPU families:
Graviton3 was run twice at every size, on two separate instances, for a reason we explain below. All runs: Linux 6.17 (glibc 2.39), Krisp Python SDK 1.12.0 on Python 3.12.3, August 2026. The benchmark has a safety cap of 508 streams; a “+” in the results means the real ceiling is higher. Single-stream baselines were not recorded on Ice Lake.
Which SDK? Krisp ships separate SDK packages for C/C++, Python, Node.js, Go and Rust, each with its own version number. This benchmark ran on the Python SDK, version 1.12.0. If you want to reproduce it, use the latest release of the SDK for your language (at the time of writing: Python 1.13.3, Node.js 1.8.2, C/C++ 9.22.0) and check its docs for enableLowPrecisionExecution.
Results: how many calls one server can clean
Here is the full model, Krisp VI Balanced, on every instance we tested. Graviton3 numbers are from the conservative instance.
FP16 roughly doubles live calls on Graviton3 and Sapphire Rapids. Ice Lake has no FP16 hardware, so nothing changes. Krisp VI Balanced. Graviton3: Instance A (conservative)
Krisp VI Balanced: max concurrent streams per instance
Instance
FP16 off
FP16 on
Change
Graviton3 · 4 vCPU
23
43
+87%
Graviton3 · 16 vCPU
76
137
+80%
Graviton3 · 32 vCPU
86
171
+99%
Sapphire Rapids · 4 vCPU
11
21
+91%
Sapphire Rapids · 16 vCPU
44
77
+75%
Sapphire Rapids · 32 vCPU
79
144
+82%
Ice Lake · 4 vCPU
7
7
0%
Ice Lake · 16 vCPU
38
38
0%
FP16 roughly doubles capacity on Graviton3 and Sapphire Rapids. Ice Lake has no FP16 hardware, so nothing changes.
The other two models follow the same pattern. Here is every model on the conservative Graviton3 instance, Sapphire Rapids and Ice Lake, FP16 off → on:
Max concurrent streams, FP16 off → on, all three models
Instance
Krisp VI Balanced
Krisp VI Default
Krisp VI Lite
Graviton3 · 4 vCPU
23 → 43 (+87%)
23 → 43 (+87%)
54 → 80 (+48%)
Graviton3 · 16 vCPU
76 → 137 (+80%)
76 → 138 (+82%)
209 → 358 (+71%)
Graviton3 · 32 vCPU
86 → 171 (+99%)
87 → 171 (+97%)
287 → 508+ (+77%+)
Sapphire Rapids · 4 vCPU
11 → 21 (+91%)
11 → 21 (+91%)
35 → 57 (+63%)
Sapphire Rapids · 16 vCPU
44 → 77 (+75%)
45 → 77 (+71%)
127 → 205 (+61%)
Sapphire Rapids · 32 vCPU
79 → 144 (+82%)
79 → 146 (+85%)
250 → 430 (+72%)
Ice Lake · 4 vCPU
7 → 7 (0%)
7 → 7 (0%)
22 → 22 (0%)
Ice Lake · 16 vCPU
38 → 38 (0%)
39 → 39 (0%)
118 → 118 (0%)
508+ means the run hit the benchmark's 508-stream safety cap; the real ceiling is higher.
Across every run, the pattern held:
Graviton3: +61% to +99% for the full models. The lite model gained 13% to 77%, for a reason explained below.
Sapphire Rapids: +61% to +91% across all three models. The most consistent gains we measured.
Ice Lake: no gain. With no FP16 hardware the SDK stays on float32, so FP16 on and off run the same code.
Single-stream latency tells the same story. On a 16-vCPU Graviton3, one frame of Krisp VI Balanced takes 3.00 ms to clean with FP16 off and 1.48 ms with FP16 on. Half the time per frame is what turns into twice the streams per server.
Two side findings are worth knowing:
Krisp VI Balanced costs the same to run as Krisp VI Default. On every instance, the two full models landed within two streams of each other. Upgrading to v2.7 does not cost you capacity.
The lite model carries about 3× the streams. Krisp VI Lite held 2.6–3.3× the streams of the full models at 16–32 vCPU, wherever CPU was the limit. With FP16 on a 32-vCPU Graviton3, it hit our 508-stream cap before it hit any budget.
Where it does not help
We want to be clear about the limits, so you know where to expect the gain.
CPUs without FP16 hardware
On Ice Lake there is no native FP16 math, so there is no speedup to get. You do not have to guard against this yourself. The SDK checks the CPU at startup and runs the FP16 path only where the hardware supports it. Everywhere else it stays on float32. You can turn the flag on across your whole fleet.
To know in advance which instances will see the gain, check the CPU flag the kernel reports:
shell
grep -m1 -owE 'asimdhp|avx512_fp16' /proc/cpuinfo
# asimdhp -> Arm FP16 (Graviton3)
# avx512_fp16 -> Intel AVX-512 FP16 (Sapphire Rapids)
# no output -> no gain; the SDK stays on float32 (Ice Lake)
On Ice Lake today? The gain is one move away. On AWS, Ice Lake runs the c6i, m6i and r6i families. Move to Graviton3 (c7g) or Sapphire Rapids (c7i) and turn the flag on. In the fleet example below, the same 1,000 streams take 432 vCPUs on Ice Lake and 96 on Graviton3.
The lite model’s gains look smaller than they are. In several Graviton3 runs with FP16 on, CPU peaked at only 47–75%. Those runs stopped for another reason, the RTF budget or the 508-stream cap, before the CPU was full. Read the lite gains as floors, not ceilings.
Three things we learned along the way
FP16 was the headline, but the benchmark taught us three more things about running voice models on cloud CPUs.
1. Graviton3 wins per vCPU, and the reason is simple
At 4 and 16 vCPU, Graviton3 held 1.7–2.1× the streams per vCPU of Sapphire Rapids for the full models, with FP16 on or off. Most of that gap is what a vCPU is. On Graviton3, a vCPU is a whole physical core. On Intel, it is one of two hyperthreads sharing a core.
Krisp VI Balanced at 16 vCPU
CPU
One vCPU is
Streams per vCPU
Streams per physical core
Graviton3 (FP16 on, 137 streams)
1 physical core
8.56
8.56
Sapphire Rapids (FP16 on, 77 streams)
½ core
4.81
9.63
Ice Lake (FP16 off, 38 streams)
½ core
2.38
4.75
Per physical core, Sapphire Rapids with FP16 is on par with Graviton3. But you rent vCPUs, not cores, so Graviton3 packs more streams into an instance of the same size. Compare prices per vCPU before assuming the gap carries straight through to your bill.
2. Bigger instances are not proportionally better
If capacity scaled perfectly, streams per vCPU would stay flat as instances grow. Sapphire Rapids comes close: going from 16 to 32 vCPU added 76–110% more streams. Graviton3 does not: the same step added only 13–25% for the full models. On the 32-vCPU Graviton3, CPU peaked as low as 66% when the RTF budget broke; the cores were not the limit. One suspect is cache: the 16- and 32-vCPU Graviton3 both report the same 32 MiB of shared L3, so each core gets half as much at 32 vCPU. We have not isolated the cause yet.
Flat would mean perfect scaling. Graviton3 falls off at 32 vCPU; Sapphire Rapids holds close to flat. Krisp VI Balanced, FP16 on. Graviton3: Instance A
The practical rule is simple: on Graviton3, run more 4–16 vCPU instances instead of fewer 32 vCPU ones.
3. Two instances of the same type can be 43% apart
We ran every Graviton3 size on two separate instances. Same instance type, same software, different physical host. Instance B beat Instance A by 21–43% on the full models.
Krisp VI Balanced on Graviton3: max concurrent streams
Size and setting
Instance A
Instance B
B vs A
4 vCPU · FP16 off
23
32
+39%
4 vCPU · FP16 on
43
52
+21%
16 vCPU · FP16 off
76
96
+26%
16 vCPU · FP16 on
137
182
+33%
32 vCPU · FP16 off
86
112
+30%
32 vCPU · FP16 on
171
221
+29%
The hardware underneath, NUMA placement, and CPU frequency all move the number. Size your fleet on the slower instance. Treat the faster one as upside, not plan.
What this means for you
If you run Krisp Voice Isolation on your own servers
Turn FP16 on by default. The SDK checks the CPU itself. On Graviton3 and Sapphire Rapids it is the cheapest 2× in this report.
On Ice Lake? Move to c7g or c7i. Ice Lake (c6i, m6i, r6i) has no FP16 math, so the flag does nothing there. Graviton3 (c7g) and Sapphire Rapids (c7i) get the 2×.
Prefer Graviton3 at 4–16 vCPU. It leads per vCPU, and scaling out beats scaling up.
Plan around the conservative numbers. Some instances will run 20–40% better. Do not count on it.
Upgrade to Krisp VI Balanced freely. It costs the same to run as Krisp VI Default.
Reach for Krisp VI Lite when density matters most. About 3× the streams, where its quality fits your use case.
A worked example. Say you need to hold 1,000 concurrent Krisp VI Balanced streams. Instances needed = peak streams ÷ max streams per instance, rounded up. Using the conservative numbers, with FP16 on wherever the CPU supports it:
The same 1,000 streams take 96 vCPUs on small Graviton3 instances and 432 to 572 on Ice Lake. Krisp VI Balanced, conservative numbers, FP16 on where supported
Total vCPUs to hold 1,000 concurrent Krisp VI Balanced streams
Instance
Streams each
Instances needed
Total vCPUs
Graviton3 · 4 vCPULeanest fleet
43
24
96
Graviton3 · 16 vCPU
137
8
128
Graviton3 · 32 vCPU
171
6
192
Sapphire Rapids · 4 vCPU
21
48
192
Sapphire Rapids · 16 vCPU
77
13
208
Sapphire Rapids · 32 vCPU
144
7
224
Ice Lake · 16 vCPU (FP16 off)
38
27
432
Ice Lake · 4 vCPU (FP16 off)
7
143
572
The leanest fleet is 24 small Graviton3 instances. The best Intel option needs 2× the vCPUs; Ice Lake needs 4.5×.
If the lite model fits your quality needs, the same 1,000 streams take 3 Graviton3 instances at 16 vCPU, 48 vCPUs in total. The 85% CPU budget already reserves headroom; your application’s own overhead is not included.
What comes next
Three things are still open. First, we want to isolate why the 32-vCPU Graviton3 stops scaling; cache is the leading suspect, but it is not confirmed. Second, our 508-stream cap hid the true ceiling of the lite model with FP16, so we will raise the cap and rerun. Third, we will run the same benchmark on the current 8th-generation AWS instances, such as Graviton4 (c8g) and Xeon 6 (c8i).
The FP16 path already runs on mainstream Arm and x64 Linux, which covers every instance in this report.
Try it on your own servers
You do not have to take our numbers. Set enableLowPrecisionExecution in the SDK and measure your own streams-per-instance before and after. The CPU flag check above tells you in advance whether an instance will see the gain. If you want us to run the benchmark on your target instance type, or share the full per-run results, reach out.
No. It is the same network with the same layers, running in 16-bit math. FP16 doesn’t degrade the quality.
CPUs with native FP16 math: AWS Graviton3 (c7g) and Intel Sapphire Rapids (c7i). Turn the flag on everywhere; the SDK checks the CPU and stays on standard precision where FP16 isn’t supported. On Ice Lake (c6i, m6i, r6i) there is no gain, so moving to c7g or c7i is how you get the 2×.
It depends on the CPU and the model. On a 16-vCPU AWS Graviton3 with FP16 on, the full model holds 137 to 138 live calls and Krisp VI Lite holds 358. Turn on FP16 first; it roughly doubles live calls per server for every model. Then consider the lite model for about 3× the streams.
Appendix: full results
Every run, as measured. Baseline RTF and latency are single-stream p95 values. “CPU at max” is p95 CPU utilization at the reported stream count.
Intel Xeon Platinum 8375C @ 2.90 GHz · FP16 path: none. With no FP16 hardware the SDK runs float32 either way, so FP16-on rows match FP16-off. Single-stream baselines were not recorded.