On July 8, 2026, OpenAI released two new voice models, GPT‑Live‑1 and GPT‑Live‑1 mini, which power a rebuilt ChatGPT Voice. The main innovations are continuous, full‑duplex audio processing — handling incoming and outgoing speech at the same time — and delegating harder reasoning tasks to a background model, GPT‑5.5, while the voice model continues to speak.
How they differ from the predecessor
GPT‑Live models process audio continuously rather than waiting for a speaker’s turn to end. In telecom terms this is full‑duplex: transmission and reception occur simultaneously, like a telephone call, instead of alternating. The earlier Advanced Voice Mode (AVM) used a single model that listened, reasoned and replied in discrete turns. GPT‑Live separates the conversational "talker" function from the background "thinker."
Each GPT‑Live model decides its next action many times per second (talk, listen, wait, break in, produce backchannel sounds such as “hmm” or “yeah,” or trigger a tool). When a request needs web search, deeper reasoning, or multi‑step work that spans turns, the voice model hands the task to GPT‑5.5, keeps talking while that model runs, and weaves the result back into the conversation. The two models share the conversation context but are orchestrated and served separately.
Features and limits
- Both models are speech‑in, speech‑out, with overlapping input and output. In ChatGPT the conversations are accompanied by on‑screen visual cards (weather, stocks, sports, maps). The app accepts images and file uploads; live video and screen sharing are not available in GPT‑Live but remain in legacy Standard and AVM modes.
- Live translation is supported, and nine remastered voices are provided; voices are predefined and safeguards prevent mimicking real people.
- GPT‑Live‑1 offers user‑selectable reasoning effort: Instant, Medium, High. Instant runs GPT‑5.5 Instant in the background; Medium and High run GPT‑5.5 Thinking at corresponding effort levels. GPT‑Live‑1 mini only calls GPT‑5.5 Instant.
OpenAI did not disclose parameter counts, architecture details, training data, knowledge cutoff, latency measurements, or usage‑based pricing for GPT‑Live.
Performance and benchmarks
OpenAI’s internal evaluations show substantial gains versus AVM, especially on tasks routed to GPT‑5.5:
- GPQA (graduate‑level science: biology, chemistry, physics): GPT‑Live‑1 at high reasoning scored 84.2% compared with AVM’s 45.3%.
- BrowseComp (agentic web search for hard‑to‑find facts): GPT‑Live‑1 at high reasoning answered 75.2% of questions correctly versus 0.7% for AVM.
- Human rater preference: in matched 5–10 minute conversations, raters preferred GPT‑Live‑1 over AVM 75.7% of the time; GPT‑Live‑1 mini was preferred 69.2% of the time.
Safety detection also improved in OpenAI’s tests. Examples reported by the company:
- Illicit behavior flagged: GPT‑Live‑1 97% vs. AVM 74%.
- Self‑harm detection: GPT‑Live‑1 96% vs. AVM 89%.
- Adversarial mental‑health prompts flagged: GPT‑Live‑1 84% vs. AVM 57%.
- Adversarial self‑harm prompts: GPT‑Live‑1 98% vs. AVM 72%.
These comparisons are internally focused—OpenAI’s published benchmarks contrast GPT‑Live with its own prior AVM, not with competing vendors’ voice models.
Availability and developer access
GPT‑Live is available globally on iOS, Android and ChatGPT.com. GPT‑Live‑1 is the default voice for Go, Plus and Pro plans at no extra cost; GPT‑Live‑1 mini is the default on the free plan. No developer API for GPT‑Live has shipped yet; for now OpenAI’s developer voice option remains GPT‑Realtime‑2, which reached the Realtime API in May.
Context and significance
Both parts of GPT‑Live’s approach — full‑duplex processing and orchestration of a separate reasoning model — have precedents in research and other projects. Examples include Alibaba’s Qwen2.5‑Omni Thinker‑Talker, Thinking Machines Lab’s TML‑Interaction‑Small, Kyutai’s Moshi, Nvidia’s PersonaPlex (built on Moshi), and Google’s Gemini Live, which offers continuous conversation along with camera and screen sharing. OpenAI’s claim to novelty is shipping this combination to a mass audience; the company says more than 150 million people use ChatGPT’s voice and dictation features each week.
According to Bloomberg, OpenAI plans to unveil a portable, screenless smart speaker that relies entirely on GPT‑Live voice interactions before the end of this year, with devices expected to be available in early 2027. The reported device would learn about its user over time and access more powerful models than current smart speakers.
Conclusion
Full‑duplex voice makes conversations feel more natural by removing turn‑detection heuristics, and delegating complex work to a separate reasoning model avoids the tradeoff between quick replies and deep thought. OpenAI’s internal benchmarks show notable improvements over AVM in reasoning, conversational quality and safety, but many technical details remain undisclosed and independent comparisons are not provided. The long‑term impact will depend on whether OpenAI can make voice a practical, secure primary interface for everyday productive tasks.



