OpenAI on Wednesday launched GPT‑Live, a set of new voice models that redefine how people converse with ChatGPT by replacing Advanced Voice Mode with an architecture that can listen and speak at the same time. The company says this design makes interactions feel more like a real human conversation.
The rollout begins today globally on iOS, Android and ChatGPT.com. Two models ship: GPT‑Live‑1, which becomes the default voice model for paid ChatGPT tiers (Go, Plus and Pro), and GPT‑Live‑1 mini, which is provided to free‑tier users. OpenAI plans to offer the models via API later and allows developers to sign up for notifications.
Why full‑duplex matters
GPT‑Live’s central technical advance is a full‑duplex architecture: instead of requiring discrete turn‑taking, the model continuously processes incoming audio while generating its own output. In telecom terms full‑duplex means both parties can speak and listen simultaneously; applied to AI, it removes the need for a clear silence gap to detect turn ends.
According to OpenAI’s research post, the model can make interaction decisions many times per second — whether to speak, continue listening, pause, interrupt or call a tool. In practical terms this enables conversational acknowledgments (for example, “mhmm,” “yeah,” “got it”) while a user is speaking, better handling of natural pauses without interrupting prematurely, and resilience to rapid interruptions.
Advanced Voice Mode, introduced to paid users in September 2024 after a limited July rollout, processed audio natively in a single model but still relied on rigid turn‑based exchanges; silence‑based turn detection could cause awkward interruptions in noisy or pausing situations. GPT‑Live is designed to overcome those brittleness issues.
Separating voice and reasoning
GPT‑Live also introduces an architectural split between the voice interaction layer and the reasoning layer. For straightforward queries GPT‑Live responds directly; for requests that need web search, deeper reasoning, or complex agentic work, GPT‑Live delegates the computation to a background frontier model — at launch this is GPT‑5.5, released by OpenAI in April — while continuing the conversation asynchronously.
This modular approach means the voice‑native model can be optimized for real‑time interaction and remain stable while the reasoning engine is upgraded independently. For enterprises and developers this allows a voice agent to maintain natural conversation with a user while simultaneously querying databases, searching the web, or running multi‑step workflows without the several seconds of dead air a monolithic pipeline would introduce.
The three generations of ChatGPT voice
Tracing the voice evolution helps show the change: the original ChatGPT Voice (2023) used a cascaded pipeline — Whisper for speech‑to‑text, GPT‑4 to generate text, and a text‑to‑speech model to produce audio — which incurred latency and information loss. As OpenHelm noted in October 2024, that old pipeline accumulated roughly 1,700 milliseconds of latency before the first spoken word.
Advanced Voice Mode (rolled out in July–September 2024) collapsed that three‑model pipeline into a single audio‑native model and added five voices and improved accent handling, but it still operated in discrete alternating turns. GPT‑Live moves beyond that to a continuous stream interaction.
The Scarlett Johansson controversy and voice imitation risks
Advanced Voice Mode’s launch followed a damaging controversy in May 2024 when OpenAI showcased a voice called “Sky” that many listeners found strikingly similar to Scarlett Johansson’s voice, after Johansson said she had declined an offer from CEO Sam Altman to voice the system. OpenAI removed the voice and apologized; the incident drew scrutiny from SAG‑AFTRA and members of Congress and intensified concerns about unauthorized replication of performers’ voices.
OpenAI says it has remastered the nine distinct ChatGPT voices for GPT‑Live and emphasizes that the system is designed for conversation, not for imitating real people, with safeguards to prevent voice impersonation.
What users will notice now — numbers and features
OpenAI reported that more than 150 million people use ChatGPT’s voice and dictation features each week, out of roughly 900 million weekly active users on the platform. Voice use cases include language practice, bedtime stories, commute chats and hands‑free assistance.
Key GPT‑Live features include:
- Rich visual cards that surface during voice conversations (weather, stock data, sports scores and maps), letting users glance at information without breaking the flow.
- Choice of three reasoning levels for responses: Instant, Medium and High.
- Improved patience and focus: ChatGPT will wait rather than interrupt if you pause, can obey a request to stay quiet and is better at focusing on the user’s voice amid background noise.
Early preview users gave cautiously positive feedback, praising improved front‑end feel and longer context knowledge work. Observers noted that the increased “smarts” often come from handing hard questions to GPT‑5.5; the new user experience comes from full‑duplex listening while talking.
Voice‑specific safety testing and results
OpenAI expanded safety evaluations to include audio‑native tests using both real opt‑in user voice samples and synthetically generated adversarial prompts across categories such as self‑harm, sexual content, illicit behavior, emotional reliance, mental health and hate speech.
On synthetic adversarial evaluations GPT‑Live‑1 showed substantial improvements over Advanced Voice Mode: illicit behavior safety rose from 0.63 to 0.97, self‑harm from 0.72 to 0.98, and hate speech from 0.87 to a perfect 1.00. Production‑prompt evaluations using real user audio were more mixed: GPT‑Live‑1 matched or improved on most categories but showed a small regression on emotional reliance (0.88 to 0.82), which OpenAI said was not statistically significant.
OpenAI implemented real‑time safeguards that can intervene while the model is speaking — steering toward safer replies, surfacing crisis resources, or ending the voice session in high‑risk cases — and added protections for teen users and voice‑adapted self‑harm support flows. The company also says it will roll out longer‑term measurement and post‑launch monitoring focused on emotional reliance.
Competitors and current limitations
Rivals are already shipping full‑duplex systems: Google’s Gemini Live supports full‑duplex conversation and also offers camera and screen sharing (capabilities GPT‑Live lacks at launch); Google released Gemini 3.1 Flash Live in March as a high‑quality real‑time audio model. ByteDance launched Seeduplex in April, claiming about a 50 percent reduction in false responses and false interruptions compared with its previous half‑duplex system. Nvidia’s PersonaPlex, released in January, added customizable voice and role control for full‑duplex models.
OpenAI’s strengths include its large user base, GPT‑5.5 integration and the breadth of the ChatGPT ecosystem, but gaps remain: GPT‑Live initially does not support voice with video or screen sharing, language support may be imperfect in some languages, and API access is not available on day one — all factors that will affect enterprise and developer adoption compared with competitors who already offer developer‑facing products.
Where this could lead
GPT‑Live represents OpenAI’s biggest bet yet on voice as a primary AI interface: a dedicated interaction layer that sits between users and the company’s most powerful models. OpenAI says the research could unlock use of voice for increasingly complex, longer‑running and more agentic tasks — for example telling a phone to book a flight, negotiate with an insurer or debug a server through a natural conversation while the system handles background work.
Two years ago voice replies were slow and stilted; one year ago they felt like turn‑taking phone calls. Today GPT‑Live brings conversations closer to natural speech — still imperfect and constrained in some respects, but clearly a step toward voice as a dominant interface. The launch also raises continuing technical, ethical and regulatory questions, particularly around emotional reliance and impersonation risks, that will shape how the technology is adopted and governed.



