back to top
HomeTechOpenAI's New Voice Models Want to Do More Than Talk Back

OpenAI’s New Voice Models Want to Do More Than Talk Back

- Advertisement -

OpenAI is pushing deeper into voice.

The company just launched three new realtime audio models in its API. GPT-Realtime-2 for conversational reasoning, GPT-Realtime-Translate for live multilingual translation, and GPT-Realtime-Whisper for streaming speech transcription.

GPT-Realtime-2 can now handle longer conversations, recover from interruptions more naturally, use tools while someone is still talking, and respond with different reasoning levels depending on the task. OpenAI says the model is designed for things like customer support, scheduling, travel assistance, and other workflows where the AI actually has to keep track of context instead of just replying quickly.

OpenAI is no longer treating voice as a side feature attached to chatbots. It’s starting to position voice as the interface itself. That means live translation during conversations. Real time transcription while meetings are still happening. AI agents that can check your calendar, pull information from apps, or complete actions while the conversation keeps moving.

Voice models are starting to behave more like agents

Three audio models in the OpenAI API

The most interesting part of this launch is not the voices themselves. It’s the fact that OpenAI keeps framing these systems around actions and workflows instead of conversations.

The company highlighted examples like Zillow building voice agents that can search for homes and schedule tours, Deutsche Telekom testing multilingual customer support, and Priceline exploring trip planning that happens conversationally from start to finish.

That points to a shift happening across AI right now. Voice assistants used to exist mainly to answer questions. These new systems are being designed to stay active while tasks are unfolding like checking calendars, updating bookings, pulling information from apps, translating conversations live, or handling interruptions without restarting the interaction.

That’s also why OpenAI focused heavily on realtime reasoning and tool use in this launch. A voice assistant that simply sounds natural is no longer enough. The hard part is making the system useful while the conversation is still moving.

Voice changes how people use software

Typing naturally creates pauses. People send a prompt, wait for a response, then move on.

Voice interactions work differently. Conversations keep moving even when requests change halfway through or multiple things happen at once.

That creates a much harder problem for AI systems. The model has to listen continuously, decide when to respond, remember context across longer sessions, and sometimes use tools without interrupting the flow of the conversation itself.

And that’s probably why companies like OpenAI are suddenly investing so heavily in realtime infrastructure.

You May Like: Open-Source TTS Models That Can Clone Voices and Actually Sound Human

GPT-Realtime-Translate may end up being the sleeper feature

The reasoning upgrades will get most of the attention, but the realtime translation model could end up having the bigger commercial impact.

OpenAI says GPT-Realtime-Translate can handle more than 70 input languages and translate into 13 output languages while keeping pace with live conversations. That opens the door for customer support, meetings, events, travel assistance, and sales calls where people no longer need to speak the same language fluently to communicate smoothly.

And unlike older translation systems, OpenAI is clearly pushing for conversations that continue naturally while the translation happens in the background.

The company also highlighted testing from BolnaAI, which said the model handled regional Indian languages like Hindi, Tamil, and Telugu with lower word error rates and fewer fallback failures compared to other systems they tested.

Vimeo is experimenting with the model as well. The company says it’s using GPT-Realtime-Translate for live translation during broadcasts so creators can reach global audiences while streaming in real time. According to Vimeo, one of the biggest improvements was how well the system handled multilingual conversations without breaking flow mid-interaction.

Multilingual voice AI becomes much more useful once it starts handling accents, interruptions, and regional speech patterns reliably in real time.

You May Like: SubQ’s 12M Token Model Could Change How AI Handles Long Context. If It’s Real.

Pricing

All three models are available through OpenAI’s Realtime API.

GPT-Realtime-2 is priced at $32 per million audio input tokens and $64 per million audio output tokens, while GPT-Realtime-Translate costs $0.034 per minute and GPT-Realtime-Whisper costs $0.017 per minute.

Developers can also test the models through OpenAI’s Playground before integrating them into apps and workflows.

Don’t miss any Tech Story

Subscribe To Firethering NewsLetter

You Can Unsubscribe Anytime! Read more in our privacy policy

LEAVE A REPLY

Please enter your comment!
Please enter your name here

YOU MAY ALSO LIKE
Zuckerberg Wrote 14 Pages About Open AI. His Best AI Model Is Still Closed

Zuckerberg Wrote 14 Pages About Open AI. His Best AI Model Is Still Closed.

0
Mark Zuckerberg published a 14-page essay today about why open-source AI is the path forward for humanity. Distribute intelligence rather than centralize it. Put the power in everyone's hands. A new era of personal empowerment. On the same day, Meta released Muse Glimmer, an open-source version of its most powerful model, Muse Spark, that anyone can download, modify, and build on for free. But the interesting part is, Muse Spark itself stays closed. You still pay to access it. The open version is nearly identical, Meta says, but the model that actually competes at the frontier, the one Zuckerberg's essay is implicitly defending remains behind a paywall. That gap between the philosophy and the product decision is what makes today's announcement interesting.

The Biggest AI Companies Are All Building Their Own Chips. That’s Not a Coincidence.

0
Anthropic confirmed this week it's hiring a custom silicon team to design chips for running Claude. The announcement was quiet a job listing, a spokesperson confirmation, no big launch event. Easy to file under "interesting but expected" and move on. But zoom out for a second. OpenAI shipped its first custom inference chip in June. Google has been running models on its own TPUs for years. Meta has designed and deployed its own silicon. Mistral is reportedly exploring the same path. And now Anthropic. Five of the most important AI labs in the world, all arriving at the same decision, within roughly the same window. None of them are copying each other. All of them looked at the same competitive landscape and reached the same conclusion independently. That kind of convergence doesn't happen by accident. It happens when an entire industry agrees that the thing everyone assumed was someone else's problem is actually the problem and that whoever solves it first has an advantage that's very hard to close later.
AI Was Supposed to Stop Cheating. Instead, 58,000 Students Must Retake Their Exams

AI Was Supposed to Stop Cheating. Instead, 58,000 Students Must Retake Their Exams.

0
UNAM runs the largest university in Mexico. Every year, hundreds of thousands of students take an entrance exam that determines whether they get in. This year, for the first time, the whole thing went remote. They deployed a lockdown browser, AI webcam monitoring, and one human supervisor per 150 applicants. The kind of setup that sounds serious on paper. Then the scores came in. Students hitting 100 or above jumped from 3.5 percent in previous years to 16.3 percent this year. At the very top end, scores of 110 or higher went from 0.9 percent to 5.5 percent. Not a small shift. Not noise. A roughly fivefold increase in top scores, in one year, under one new format. An expert commission investigated. Their conclusion: administer the entire exam again, in person, to around 58,000 people. The rector apologized to students who hadn't cheated. They now have to prepare for and sit another exam anyway.