Industry Pulse

Mercury Voice: What a 320ms AI Phone Agent Changes

Inception Labs' new diffusion voice model closes the latency gap for AI phone agents

A rotary-style telephone handset resting in warm, dim light, evoking the gap between old phone systems and instant AI response.
Fallback editorial image (picsum.photos) used after Unsplash hotlink verification could not be completed in this unattended run.

The short answer

Inception Labs launched Mercury Voice on 29 September 2026, a diffusion-model voice backend for AI phone agents that answers in a 320ms median, compatible with LiveKit, Pipecat, Vapi, and Retell. At roughly $0.009 a minute, it is the first voice model fast enough to make a real pilot worth running, though the benchmarks are self-reported and real-world latency still depends on your own telephony setup.

Inception Labs launched Mercury Voice on 29 September, a voice model built for AI phone agents that answers in a median of 320 milliseconds, fast enough to feel like a person picking up mid-sentence. It plugs straight into LiveKit, Pipecat, Vapi, and Retell, the same voice-agent infrastructure most builders already run. If you've looked at a phone agent for your business and walked away because the pause gave the robot away, that gap just got a lot smaller.

What Is Mercury Voice, Exactly?#

Mercury Voice is a diffusion language model, not the token-by-token kind most chatbots run on. Inception says that lets it refine several words at once instead of generating one token after another, which is where the speed comes from.

According to Inception's own announcement, Inception claims the model is 5.9 times faster than OpenAI's GPT-6 Luna on the same measure, with enough context (128K input, up to 50K output tokens) to hold a long call plus tool calls in memory. Early users named in the launch include Audivi AI, running drive-through ordering with order changes, and OpenCall, which reports seeing around 170 milliseconds in its own production traffic.

The numbers in one place:

  • Median time to first spoken response: 320ms (750ms at the 95th percentile)
  • Context: 128K input tokens, up to 50K output tokens
  • Launch pricing: $0.20/M input, $0.75/M output tokens (regular rate $0.40/$1.50)
  • Estimated cost: roughly $0.009 per minute of conversation
  • Works with: LiveKit, Pipecat, Vapi, Retell

Does This Change Anything If You're Not Running a Phone Agent?#

If you're handling inbound calls with a receptionist or a basic IVR tree today, nothing here is urgent. There's no action item for you this week, and you can stop reading if phone automation isn't on your roadmap.

If you tried a voice AI vendor like Vapi or Retell before and walked away because the half-second of dead air made callers repeat themselves or hang up, this is the number worth revisiting. A 320 millisecond median sits inside the range most people read as a normal human pause. That's a real shift in what a phone agent can get away with, not a marketing one.

It matters most where a caller is deciding something in real time: a booking confirmation, a car rental query, a payment-plan call. It matters a lot less for outbound dialing or anything fully scripted, where a one-second lag was never what cost you the sale.

The Honest Trade-Offs#

The upside is real. A model fast enough to get interrupted, backtrack, and still hold a tool call mid-sentence closes the biggest gap between an AI phone agent and a human one. Compatibility with LiveKit, Pipecat, Vapi, and Retell means a team can test it without ripping out an existing voice stack. And at under a cent a minute, it's priced to actually run calls, not just demo well.

The caveats are just as real. Inception's benchmark comparisons against GPT-6 Luna, Gemma 4, GLM-5.3-Flash, Qwen3.5, and Gemini 3.5 Flash-Lite come from Inception itself, not an independent lab, and none of the coverage we checked published the raw scores. The 320ms figure also measures the model's own response, not the full call, which is likely why OpenCall's reported 170ms differs from Inception's median. And launch pricing at half off won't be the price you pay in six months.

How We're Already Treating This#

At WebEpex we build most of our client communication automation WhatsApp-first, on AiSensy and n8n, specifically because async text tolerates a slow backend and a live call doesn't forgive a two-second pause the same way. That bias toward text has only gotten stronger since WhatsApp's own per-message pricing took effect today: a slow voice call now competes against a channel that also has a real cost attached to every reply.

One of our GCC-facing communications builds this year paired a traditional IVR with a WhatsApp handoff for exactly that reason: we didn't trust a voice agent to hold a caller's attention long enough to be worth the risk. Meta's own bundled AI pricing split is a reminder that packaged tools come with trade-offs of their own, and we'd rather test a model on our own terms than adopt it because a vendor bundled it in.

We pulled Mercury Voice's numbers into that conversation the day the announcement landed. Our read: the latency finally clears the bar for a narrow pilot, on a low-stakes flow like after-hours booking confirmations or appointment reminders, before we'd put it in front of a caller who's upset or negotiating something. We're scoping that pilot against a scripted, non-urgent flow first, not a full phone-agent swap, and we'll only widen it if the real-world number holds closer to Inception's median than to a worst case.

Same reasoning as cheaper chatbot routing layers and Sakana's model-routing approach earlier this quarter: a faster model earns a swap only after it's tested against your worst case, not the vendor's demo script. We also build chatbots on Voiceflow for clients across the GCC, Europe, and Indian SMB segments, and this is the first voice backend worth pricing out for those flows.

What I'd Tell a Client This Week#

Don't rip out your current phone setup over this. Do ask whether your use case is low-stakes enough to pilot, reminders, confirmations, simple FAQs, before trusting it with an upset customer or a live sale. Ask whoever built your voice stack if they can point Vapi, Retell, LiveKit, or Pipecat at Mercury Voice as a drop-in swap: that's a cheap test, not a rebuild. And hold off on believing any vendor's latency claim, including Inception's, until you've heard it on your own line, with your own accent and background noise.

If you want a straight read on whether your call flow is even ready for something like this, send me what you're running and I'll tell you honestly. Takes two minutes, no pitch attached. [cal.com/webepex/growth-review]

Sources

  1. Introducing Mercury Voice
  2. Inception launches Mercury Voice for faster AI phone agents

Frequently asked questions

Straight answers to what people ask about Mercury Voice.

Does Mercury Voice replace my current WhatsApp or chatbot automation?
No. Mercury Voice is specifically a voice model built for phone calls, not WhatsApp or text chat. If you are not running a phone agent, nothing about your existing WhatsApp or chatbot setup changes because of this launch.
Is Mercury Voice ready to handle an upset customer call today?
We would not put it there yet. Its benchmarks are self-reported by Inception, and real-world latency depends on your telephony provider and network, so a low-stakes flow like booking confirmations or reminders is the safer place to test it first.
Do I need to switch voice vendors to try Mercury Voice?
No. It is compatible with LiveKit, Pipecat, Vapi, and Retell, the infrastructure most phone-agent builds already run on, so it can usually be tested as a model swap rather than a full rebuild.
How much does Mercury Voice cost to run?
Inception's launch pricing is $0.20 per million input tokens and $0.75 per million output tokens, working out to roughly $0.009 per minute of conversation. That is a 50 percent launch discount off the regular $0.40/$1.50 rate, so expect it to rise later.
Prakhar Vohra
Written by

Prakhar Vohra

Founder & Growth Lead

Founder & CEO - WebEpex & DevAegis, Co-Founder - Tattva Aura Events, I work 1:1 with founders & to build profitable & scalable revenue models

Want this run for you?

We build the system, run the ads and hold the number. Book a call and we will map it in 30 minutes.

Book a call