GPT-Live-1 arrived in the OpenAI API on September 10, 2026, after powering ChatGPT Voice since July. The headline rate is $0.05 per minute, billed by the second. But that number only covers the voice layer. Every tool call, every bit of reasoning, every database lookup your agent needs mid-conversation gets billed separately at whatever backend model you pair it with. A two-minute support call where the agent checks an order status and rewrites a return label is not a $0.10 call.
That gap between the sticker price and the real call cost is the most important thing to understand about this release.
What full-duplex voice actually means for a working agent
GPT-Live-1 is positioned as a full-duplex voice model for applications that need live speech rather than the slow sequence of record, transcribe, reason, synthesize, and play. The release formalizes a two-tier production pattern: GPT-Live-1 handles continuous voice interaction at the front end, while deeper reasoning, tool calls, business logic, and durable state remain in a backend model or agent harness controlled by the developer.
Full duplex is the defining property: the model keeps listening while it is still speaking, so a caller can interrupt mid-sentence and the reply adapts instead of finishing its turn. That removes the stiff, waiting-for-your-turn quality that has made most voice agents feel like a worse phone tree. Voice agents have felt stiff because turn-taking models make a caller wait for a whole reply before interrupting.
Voice agents have spent 2026 stuck between two bad options: stitch together separate speech-to-text, reasoning, and text-to-speech models and eat the latency of every handoff, or use a consumer voice product that locks you out of choosing your own backend model. GPT-Live-1 in the API is the third option: a production voice layer you can pair with whatever reasoning model fits your latency and cost budget.
The real cost of a voice agent call
GPT-Live-1 voice sessions cost $0.05 per minute and are billed by the second. Whatever backend model and tools a session calls on for reasoning or actions are charged separately at their normal API rates, so the total price of a call depends on how much work GPT-Live-1 delegates while the conversation runs.
Here is what that looks like in practice for a support team handling five-minute average calls:
| Cost component | Rate | 5-min call estimate |
|---|---|---|
| GPT-Live-1 voice layer | $0.05 / min | $0.25 |
| Backend model (e.g. GPT-6 Astra) | $10 / 1M input tokens | varies by context |
| Tool calls (CRM lookup, write) | per-call pricing | depends on call volume |
| Realistic total | - | $0.40-$1.20+ |
That back-of-envelope puts a high-volume support line - say, 2,000 calls a day - somewhere between $800 and $2,400 in daily AI spend before any infrastructure. Not ruinous, but worth modeling before you commit to a voice-first redesign.
Every performance number OpenAI published describes a pair: the voice model plus whichever backend model ran behind it, at whichever reasoning effort. Change the pairing and the number changes. If you benchmark GPT-Live-1 against a task with GPT-6 Astra behind it, then swap in a cheaper backend to reduce cost, your production numbers will not match the launch figures.
How the delegation model changes what your agent needs to own
OpenAI's Live API documentation separates the live conversation role from the delegated-task role: the voice model manages the spoken exchange, while a backend can reason over larger context, call tools, retrieve records, run policies, or execute application-specific workflows. That separation is not cosmetic; it determines where permissions are checked, where confirmations are stored, where retries are reconciled, and where a team proves that an external action actually happened.
In simpler terms: GPT-Live-1 will not save your application from sloppy agent design. The voice layer is deliberately thin. Heavy thinking does not happen in the voice layer: GPT-Live-1 delegates reasoning and actions to the backend models and tools it is paired with, and supports streaming plus function calling so a live conversation can trigger real work.
This means your backend agent still needs:
- A clear tool boundary (what the voice agent is allowed to call without confirmation)
- A durable log of what action was taken per session, in case the call drops
- A human approval path for anything irreversible - refunds, account changes, order cancellations
Without those, you have a voice interface on top of an unreliable agent, and the naturalness of the conversation makes errors feel more surprising, not less.
What this means for teams building on Slack and Teams
If your team already uses Slack or Teams as the primary interface for internal agents, GPT-Live-1 opens a narrow but real path to voice-triggered workflows - think a channel bot that a field engineer can talk to while hands-on with equipment, or a meeting assistant that listens and surfaces context without requiring anyone to type.
Beta access to advanced features like multi-speaker detection and emotion recognition is rolling out to select enterprise partners in Q4 2026, with broader availability expected in early 2027. Multi-speaker detection matters specifically for meeting contexts, where the agent needs to attribute who said what before it can route a task correctly.
gpt-realtime is being retired on January 20, 2027 in favor of gpt-realtime-2.1, and the docs describe GPT-Live as a different design: Realtime uses one model for speech, reasoning and tool selection; GPT-Live splits the conversation from the backend. If you have anything built on gpt-realtime today, that is a forced migration decision by January - and it is worth asking whether gpt-realtime-2.1 or GPT-Live-1 is the right target depending on whether you want a simpler stack or backend flexibility.
GPT-Live-1 API voice agents: common questions
What is GPT-Live-1 and how does it differ from GPT-Realtime?
GPT-Live-1 is OpenAI's full-duplex voice model, available in the API from September 10, 2026. Unlike GPT-Realtime, which uses one model for speech, reasoning, and tool selection, GPT-Live-1 splits the conversation layer from the backend. You choose and pay for the reasoning model separately, giving you more control over cost and capability.
How much does a GPT-Live-1 voice agent call actually cost?
The voice layer is $0.05 per minute, billed by the second. Backend model inference and tool calls are charged separately at their own rates. A five-minute call with moderate reasoning work typically costs $0.40-$1.20 in total, depending on which backend model you pair and how many tools fire per call.
Can I use GPT-Live-1 with Slack or Microsoft Teams?
Not natively - GPT-Live-1 is an API-level model you integrate yourself. You can build a voice agent that routes actions into Slack or Teams channels, but the audio interface itself runs outside those platforms. A teammate like Beagle can handle the in-channel approval and logging steps while GPT-Live-1 manages the spoken conversation.
Is GPT-Realtime being replaced by GPT-Live-1?
They are separate designs. GPT-Realtime is being retired on January 20, 2027 and replaced by gpt-realtime-2.1, which is its own successor. GPT-Live-1 is a different architecture that splits the voice layer from the backend reasoning model. Teams currently on gpt-realtime need to decide which path fits their use case before the January deadline.
What does GPT-Live-1 not handle on its own?
The voice layer is deliberately thin. It does not run complex reasoning, manage tool permissions, store session state durably, or enforce confirmation steps on irreversible actions. All of that still lives in the backend agent you pair it with - which is where most of the real engineering work sits.