On July 23, 2026, OpenAI announced an update to its ChatGPT desktop application that adds ChatGPT Voice, a voice-driven interface enabling users to control AI agents and perform tasks on their computers by speaking.
Capabilities of the voice interface
ChatGPT Voice is powered by OpenAI’s new family of voice models called ChatGPT-Live, introduced earlier in the month. The feature works with both ChatGPT Work and Codex, and can leverage computer-use skills to look up websites and apps. On macOS, Appshots allow the app to access what’s on the user’s screen, including alt-text descriptions.
The smartphone version of ChatGPT Voice launched earlier with improved conversational smoothness and better interruption handling, but it was not designed to take direct actions on phones. The July 23 desktop update expands on that: users can now dictate complex, multi-step commands and respond when ChatGPT requests input.
Demonstration and rollout
In a demo video, OpenAI showed a developer issuing a single command that instructed ChatGPT to create a new thread, make a pull request, and investigate the root cause of a bug. According to OpenAI’s July 23 Twitter post, ChatGPT Voice began rolling out globally that day; the company noted the feature can speak, listen, and coordinate work within the app simultaneously.
OpenAI also said ChatGPT Voice can be used with Codex via remote access from the iOS app.
Competitor developments
Anthropic has also updated Claude’s voice mode. That update lets Claude tap Anthropic’s Opus, Sonnet, and Haiku models to complete tasks inside applications such as Gmail, Calendar, Slack, Notion, and Canva.
Why this matters
The desktop integration broadens practical voice interaction capabilities: users can not only converse with AI but also initiate and direct multi-step workflows by voice, with the system able to access certain elements of the computer environment. This marks another step toward voice-enabled, action-capable AI agents.



