OpenAI has made a major upgrade to the ChatGPT desktop app today—the client, which has already merged with Codex, now officially introduces the ChatGPT Voice mode powered by the GPT-Live model. In the new version, users can directly start, check, or adjust tasks using voice in three sections: Chat, Work, and Codex, moving multi-threaded collaboration from typing on a keyboard to speaking.

image.png

This feature has a long history. OpenAI first added speech recognition to ChatGPT in 2024, then aggressively developed speech technology until the release of the latest GPT-Live model, which finally achieved full-duplex capabilities—it can listen and speak at the same time, making the experience of AI conversation more like real human interaction. In their statement, OpenAI clearly outlined their vision: in the future, users will only need to speak their needs, and ChatGPT will quickly advance multiple tasks according to your thoughts.

The most important feature of the new voice mode is that it allows users to start multiple workflows from a single conversation, while the intelligent agent silently completes the tasks in the background, enabling you to continue chatting with it. OpenAI gave an example of planning a business trip: you can ask ChatGPT to check your calendar for conflicts, scan your inbox for flight changes, and prepare meeting notes—everything happens while you are making coffee or doing something else. Compared to simply using voice as an input box, this "you speak, it handles the task, and both proceed without hindrance" interaction is exactly what GPT-Live aims to differentiate itself with.