Model launches

OpenAI launches GPT‑Live‑1 models to enable more natural, interruptible voice conversations

OpenAI introduced GPT‑Live‑1 and GPT‑Live‑1 mini, full‑duplex conversational speech models designed to listen and speak simultaneously, enabling natural interruptions, live translation and longer hands‑free sessions.

OpenAI launches GPT‑Live‑1 models to enable more natural, interruptible voice conversations

OpenAI announced two new conversational speech models, GPT‑Live‑1 and GPT‑Live‑1 mini, intended to make real‑time voice interactions feel more natural. The company says these models support full‑duplex operation — they can listen and speak simultaneously — which allows users to interrupt naturally and enables features such as live translation.

Changes to ChatGPT's voice mode

OpenAI said it will replace the existing Advanced Voice Mode in ChatGPT with GPT‑Live‑1 mini by default. Paid‑tier users will be able to access the larger GPT‑Live‑1 model. Previously, the voice capability combined a speech‑to‑text model for transcription, a large language model to generate responses, and a text‑to‑speech model for audio output; the new models aim to provide a smoother, integrated experience.

Integration with newer text models

In a press briefing, the company said the new models address issues such as interrupting users mid‑speech and insufficiently capable answers. GPT‑Live‑1 models will forward queries to OpenAI’s latest text models, such as GPT‑5.5, for search, reasoning or agentic tasks while maintaining the live spoken conversation.

During a demo, OpenAI demonstrated that the model can remain silent for extended periods to absorb conversational context until it is called upon. Because the new voice mode can access newer GPT models, it can also display some information visually.

Competitors and market context

Other companies are also moving toward more visual and interactive assistants. The startup Monogram, which raised $40 million in seed funding from DST and Lux Capital, is developing visual responses to make assistants more engaging. Major players like Apple and Amazon have updated their assistants to handle context better and be more conversational, and startups such as Sesame, founded by Brendan Iribe and Ankit Kumar, have launched assistants that carry out tasks in the background while maintaining more natural dialogue.

Longer hands‑free conversations and voice as a primary interface

Atty Eleti, product lead for ChatGPT Voice, said during the briefing that he has had 30‑ to 40‑minute conversations with the voice feature while walking. Eleti suggested that over time voice could become the primary interface for computing and for managing increasingly complex, long‑running agentic work. The company also acknowledged reports that it might launch a pair of earbuds with AI capabilities this year, but provided no hardware details.

Eleti said: “Over time, we think this will also unlock the ability to use voice as a kind of primary interface to computing, and to manage increasingly complex long‑running agentic work. The kind of amazing use cases that we see people using Codex and ChatGPT to accomplish, we think voice can be the future interface to all kinds of work.”

Safety, limitations and language issues

OpenAI stressed that although the goal is a more natural‑sounding voice mode, it is not trying to position the assistant as an AI companion. The company said the new models include safeguards to provide age‑appropriate responses to teens and to offer resources if conversations turn to topics such as self‑harm.

The voice mode still requires improvement. In a demo of the live translation feature into Hindi, the assistant spoke with a heavy American accent and produced Hindi that sounded unnatural and somewhat bookish. OpenAI said the mode is optimized for “most spoken languages” but did not specify which languages those are.

Adoption and outlook

OpenAI has been enhancing voice features over the past few years to make ChatGPT’s voice mode sound more natural. The company said more than 150 million people use ChatGPT’s Voice and Dictation features for conversation. The GPT‑Live‑1 family represents OpenAI’s push toward longer, interruptible, hands‑free voice interactions supported by stronger text‑model reasoning and occasional visual output, while acknowledging remaining language and accent limitations and ongoing safety considerations.