Inception Labs launched Mercury Voice on 29 September, a voice model built for AI phone agents that answers in a median of 320 milliseconds, fast enough to feel like a person picking up mid-sentence. It plugs straight into LiveKit, Pipecat, Vapi, and Retell, the same voice-agent infrastructure most builders already run. If you've looked at a phone agent for your business and walked away because the pause gave the robot away, that gap just got a lot smaller.
What Is Mercury Voice, Exactly?#
Mercury Voice is a diffusion language model, not the token-by-token kind most chatbots run on. Inception says that lets it refine several words at once instead of generating one token after another, which is where the speed comes from.
According to Inception's own announcement, Inception claims the model is 5.9 times faster than OpenAI's GPT-6 Luna on the same measure, with enough context (128K input, up to 50K output tokens) to hold a long call plus tool calls in memory. Early users named in the launch include Audivi AI, running drive-through ordering with order changes, and OpenCall, which reports seeing around 170 milliseconds in its own production traffic.
The numbers in one place:
- Median time to first spoken response: 320ms (750ms at the 95th percentile)
- Context: 128K input tokens, up to 50K output tokens
- Launch pricing: $0.20/M input, $0.75/M output tokens (regular rate $0.40/$1.50)
- Estimated cost: roughly $0.009 per minute of conversation
- Works with: LiveKit, Pipecat, Vapi, Retell
Does This Change Anything If You're Not Running a Phone Agent?#
If you're handling inbound calls with a receptionist or a basic IVR tree today, nothing here is urgent. There's no action item for you this week, and you can stop reading if phone automation isn't on your roadmap.
If you tried a voice AI vendor like Vapi or Retell before and walked away because the half-second of dead air made callers repeat themselves or hang up, this is the number worth revisiting. A 320 millisecond median sits inside the range most people read as a normal human pause. That's a real shift in what a phone agent can get away with, not a marketing one.
It matters most where a caller is deciding something in real time: a booking confirmation, a car rental query, a payment-plan call. It matters a lot less for outbound dialing or anything fully scripted, where a one-second lag was never what cost you the sale.
The Honest Trade-Offs#
The upside is real. A model fast enough to get interrupted, backtrack, and still hold a tool call mid-sentence closes the biggest gap between an AI phone agent and a human one. Compatibility with LiveKit, Pipecat, Vapi, and Retell means a team can test it without ripping out an existing voice stack. And at under a cent a minute, it's priced to actually run calls, not just demo well.
The caveats are just as real. Inception's benchmark comparisons against GPT-6 Luna, Gemma 4, GLM-5.3-Flash, Qwen3.5, and Gemini 3.5 Flash-Lite come from Inception itself, not an independent lab, and none of the coverage we checked published the raw scores. The 320ms figure also measures the model's own response, not the full call, which is likely why OpenCall's reported 170ms differs from Inception's median. And launch pricing at half off won't be the price you pay in six months.
How We're Already Treating This#
At WebEpex we build most of our client communication automation WhatsApp-first, on AiSensy and n8n, specifically because async text tolerates a slow backend and a live call doesn't forgive a two-second pause the same way. That bias toward text has only gotten stronger since WhatsApp's own per-message pricing took effect today: a slow voice call now competes against a channel that also has a real cost attached to every reply.
One of our GCC-facing communications builds this year paired a traditional IVR with a WhatsApp handoff for exactly that reason: we didn't trust a voice agent to hold a caller's attention long enough to be worth the risk. Meta's own bundled AI pricing split is a reminder that packaged tools come with trade-offs of their own, and we'd rather test a model on our own terms than adopt it because a vendor bundled it in.
We pulled Mercury Voice's numbers into that conversation the day the announcement landed. Our read: the latency finally clears the bar for a narrow pilot, on a low-stakes flow like after-hours booking confirmations or appointment reminders, before we'd put it in front of a caller who's upset or negotiating something. We're scoping that pilot against a scripted, non-urgent flow first, not a full phone-agent swap, and we'll only widen it if the real-world number holds closer to Inception's median than to a worst case.
Same reasoning as cheaper chatbot routing layers and Sakana's model-routing approach earlier this quarter: a faster model earns a swap only after it's tested against your worst case, not the vendor's demo script. We also build chatbots on Voiceflow for clients across the GCC, Europe, and Indian SMB segments, and this is the first voice backend worth pricing out for those flows.
What I'd Tell a Client This Week#
Don't rip out your current phone setup over this. Do ask whether your use case is low-stakes enough to pilot, reminders, confirmations, simple FAQs, before trusting it with an upset customer or a live sale. Ask whoever built your voice stack if they can point Vapi, Retell, LiveKit, or Pipecat at Mercury Voice as a drop-in swap: that's a cheap test, not a rebuild. And hold off on believing any vendor's latency claim, including Inception's, until you've heard it on your own line, with your own accent and background noise.
If you want a straight read on whether your call flow is even ready for something like this, send me what you're running and I'll tell you honestly. Takes two minutes, no pitch attached. [cal.com/webepex/growth-review]