OpenAI announced the introduction of GPT-Live-1 into its API, allowing developers to build applications and business processes that support real-time voice interaction. The model supports full-duplex conversation, able to listen and output voice simultaneously, and can handle interruptions, pauses, and background noise during the conversation.
GPT-Live-1 also supports phone scenarios, such as restaurant reservations and customer service, for full-duplex voice agents. OpenAI stated that developers can connect it with tools like Codex or use OpenAI Presence to develop voice work that can query enterprise systems, perform approved actions, and transfer to human agents when necessary.

Speech understanding and output combined, lower latency
According to the introduction, GPT-Live-1 integrates speech understanding and speech output within the same model, reducing the multiple connections of the traditional "speech-to-text - large language model - text-to-speech" process, thus reducing latency; the model also supports complex reasoning and tool calls being handled by backend text models. Developers can adjust tone, speed, and conversation style through system prompts, and can also configure backend models, tools, and agent frameworks. OpenAI stated that GPT-Live-1 supports native speech recognition transcription and response text, and has capabilities including letter and number understanding, keyword bias, and round detection.
Early evaluations show that after using GPT-Live-1, the number of mis-interruptions during learners' thinking pauses on the language learning platform Speak decreased by nearly 80% compared to the previous round-based system; Tony Stoyanov, co-founder and CTO of a medical service platform, said his team reduced the code volume by 80%, removing about 23,000 lines of code. OpenAI's evaluation results showed that GPT-Live-1 outperformed GPT-Realtime-2.1 by 30 percentage points in the Full Duplex Bench test, and ranked first in the Tau3 end-to-end voice agent task test when paired with GPT-6 Astra at medium reasoning intensity.
Front-end speech layer costs $0.05 per minute, added 12 voices
In terms of pricing, the front-end speech layer of GPT-Live-1 costs $0.05 per minute, which is approximately $3 (about 20.2 RMB) per hour. The cost of the back-end model and agent tools is calculated separately. OpenAI also expanded real-time voice options, adding 12 new voices including Quartz, Ripple, Vesper, Willow, Stone, Gleam, Meridian, Bossa, Tempo, Beacon, Delta, and Cinder, and plans to continue adding more voices and language support in the coming months.


