Two things landed in the same week that, taken separately, look like routine model releases. Microsoft shipped a streaming speech-to-text model and two text-to-speech models on October 1. Inception made its diffusion voice LLM generally available on September 29. Put them side by side and you can see something more interesting: the voice agent stack is being sold to you in pieces now, and the vendors assembling those pieces are not the same vendors.
That matters for any team considering a voice-first workflow - a live meeting assistant, a voice-driven ops bot in Slack, a phone agent for customer intake. The architecture question used to be "which vendor do we pick?" It is now "which components, from which vendors, wired in which order?"
What Microsoft actually shipped on October 1
On October 1, Microsoft launched MAI-Transcribe-2-Streaming, its first streaming speech-to-text model, alongside MAI-Voice-2.1 and MAI-Voice-2.1-Flash for text-to-speech. The point is not just three new models - it is that Microsoft's announcement frames the three releases as components for building conversational agents. That framing is deliberate. This is infrastructure, not a product.
The transcription model is the most immediately useful of the three. MAI-Transcribe-2-Streaming returns provisional text while someone is speaking, updates those results as additional context arrives, and Microsoft says its first partial results arrive just over 100 milliseconds after receiving audio.
Microsoft says it ranks first for accuracy on both final and partial transcripts on Artificial Analysis's AA-WER Streaming evaluation, at 2.5% word error rate with the final transcript ready 0.13 seconds after end of speech.
The number worth computing: streaming costs $0.54 per hour of audio as an introductory rate through the end of 2026. For a team running four hours of meeting transcription a day, that is $2.16/day or roughly $540/year - before you add the LLM or TTS layers.
On the output side, MAI-Voice-2.1-Flash generates 45 seconds of audio at 150 milliseconds end-to-end latency, with inference Microsoft measures as 55% faster than the standard model.
Vercel lists MAI-Voice-2.1 at $22 per million characters and Flash at $15 per million characters.
One caveat worth knowing before you build on this: Microsoft Learn lists the features as public previews without an SLA and not recommended for production. That is not a reason to ignore them - it is a reason to build the abstraction layer now so you can swap them in when they graduate.
Where Inception's Mercury Voice fits in the loop
Microsoft covers the ears and mouth. The middle - the LLM that reasons about what was said and decides what to say - is the bottleneck most voice agents quietly paper over by using a small, fast, cheap model that cannot actually do much.
Mercury Voice is a diffusion language model from Inception, tuned for the LLM slot of a voice agent between speech-to-text and text-to-speech.
Conventional language models generate one token at a time, feeding each token back into the model before predicting the next. Diffusion language models refine multiple token positions in parallel, using GPU capacity more efficiently and increasing generation speed.
The practical result Inception is claiming: Inception made Mercury Voice generally available to enterprise customers on September 29, pitching its diffusion-based language model as a way to give voice agents time to reason and call tools without leaving callers in silence. The company reports a 320-millisecond median time to first answer token on production customer-service prompts, below the roughly 500-millisecond conversational target it uses for comparison.
Mercury Voice supports a 128K context, three reasoning-effort settings, up to 50K output tokens, and an OpenAI-compatible API.
Enterprise access is via Inception's sales team; the model drops into LiveKit, Pipecat, Vapi, and Retell stacks.
The price works out to something specific: the model lists at $0.40 per million input tokens and $1.50 per million output, 50% off at launch, which works out to "about $0.009 per minute of conversation" by Inception's own math. At 1,000 agent-minutes per month, that is $9 for the LLM slice - a meaningful number when you are shopping the full stack.
What this means for teams choosing a voice agent pipeline
The voice-agent market looked like a product choice six months ago: pick ElevenLabs, pick Vapi, pick Retell, pick whatever your CCaaS vendor bundles. Mercury Voice sits in the LLM slot of a voice pipeline: it reasons, calls tools, returns structured outputs and follows long system prompts, served through an OpenAI-compatible chat-completions endpoint so it plugs into LiveKit, Pipecat, Vapi, Retell, or a custom voice stack. Microsoft's stack covers the other two slots. You can now mix and match.
That is genuinely new, and it creates a real evaluation problem. Here is how the three components compare:
| Slot | Model | Vendor | Key number | Price |
|---|---|---|---|---|
| STT | MAI-Transcribe-2-Streaming | Microsoft | 2.5% WER, 0.13s to final | $0.54/audio hr |
| LLM | Mercury Voice | Inception | 320ms median TTFAT | ~$0.009/conv. min |
| TTS | MAI-Voice-2.1-Flash | Microsoft | 150ms end-to-end | $15/1M chars |
The non-obvious implication: if you run a voice agent for 1,000 minutes a month using this stack, the LLM slot costs roughly $9, the TTS slot costs roughly $1-3 depending on output verbosity, and the STT slot at $0.54/hr works out to about $9 for 1,000 minutes of audio. Total: around $20/month per thousand conversation-minutes, before orchestration, hosting, or the voice agent framework itself. That is cheap enough to prototype seriously.
What it does not tell you: Mercury Voice's headline latency was measured on the model alone, on Inception's own test set, and the benchmark win it claims arrives without a single number attached in the launch text. Run your own prompts through it before committing.
Voice AI agent pipeline: common questions
What is the difference between STT, LLM, and TTS in a voice agent?
A voice agent pipeline has three stages: speech-to-text (STT) converts the caller's audio to text, an LLM reads that text and decides what to say, and text-to-speech (TTS) converts the response back to audio. Latency compounds across all three stages. Each can now be sourced from a different vendor and swapped independently.
How fast does a voice AI agent need to respond?
A response that starts within 500 milliseconds feels conversational. Above 800ms, most people perceive a noticeable pause. Inception uses 500ms as its benchmark target for Mercury Voice's LLM slot. STT and TTS each add additional latency on top, so the LLM alone needs to come in well under 500ms to keep the total below the perceptible threshold.
Is Microsoft's MAI-Transcribe-2-Streaming ready for production?
Not yet officially. Microsoft Learn lists the features as public previews without an SLA and not recommended for production. The introductory pricing of $0.54 per audio hour runs through December 31, 2026, suggesting GA is planned but not confirmed. It is worth evaluating now and building with a swappable abstraction.
What does Mercury Voice cost per minute of conversation?
Mercury Voice lists at $0.40 per million input tokens and $1.50 per million output, 50% off at launch, which works out to roughly $0.009 per minute of conversation by Inception's calculation. That math assumes a specific average prompt and response length for customer-service calls; your cost per minute will vary with prompt depth and tool-call frequency.
Can these models be used together in one pipeline?
Yes. Mercury Voice is served through an OpenAI-compatible chat-completions endpoint so it plugs into LiveKit, Pipecat, Vapi, Retell, or a custom voice stack.
MAI-Transcribe-2-Streaming is available through the MAI Playground, Vercel, and Azure Voice Live, with LiveKit support listed as coming soon. The two vendors have not co-tested or certified a joint stack - integration testing is on you.