Monday, August 3, 2026

AI & Models

OpenAI expands Realtime API with reasoning and translation

OpenAI has updated its Realtime API with new voice intelligence features, including GPT-5-class reasoning and real-time translation, though the company notes these tools could potentially be misused.

OpenAI expands Realtime API with reasoning and translation
Photo: OpenAI

OpenAI announced on Thursday that its API will now include a suite of new voice intelligence features. The launch is designed to help developers build applications that can talk, transcribe, and translate conversations with users in real-time. This expansion aims to assist developers across multiple sectors, including customer service, education, media, events, and creator platforms.

The update, which is housed within OpenAI’s Realtime API, introduces three distinct capabilities to handle complex user requests:

  • GPT-Realtime-2: A voice model built to create realistic vocal simulations that can converse with users. Unlike its predecessor, this model is built with GPT-5-class reasoning—a higher performance tier designed to deal with more complicated requests from users.
  • GPT-Realtime-Translate: A conversational translation service designed to provide real-time translation. The feature supports more than 70 input languages that the model can comprehend, and 13 output languages that it relays back to the speaker.
  • GPT-Realtime-Whisper: A live speech-to-text transcription capability that gives users live transcription captured as interactions occur.

According to OpenAI, the newly launched models are designed to transition real-time audio from simple call-and-response systems to active voice interfaces that can listen, reason, translate, transcribe, and take action during a conversation.

While these tools offer utility for enterprises looking to expand customer service, education, media, events, and creator platforms, the company acknowledges that the new features could be misused. To address these risks, OpenAI has built guardrails to stop its new features from being abused to create spam, fraud, or other forms of online abuse. The company has embedded specific triggers in the system so that “conversations can be halted if they are detected as violating our harmful content guidelines,” OpenAI said.

Why it matters

These models shift real-time audio from simple call-and-response to active voice interfaces capable of reasoning and taking action, offering significant utility for developers in customer service, education, and media.