Inworld releases Realtime TTS-2 for responsive voice agents
Inworld has launched Realtime TTS-2, a voice model built for realtime conversation that can process the full audio context of an exchange, including a user’s tone, pacing and emotional state. The company says the model is ranked #1 on Artificial Analysis and is designed to make generated speech feel more responsive and natural.
Developers can steer delivery with natural-language prompts rather than preset emotion controls, using instructions such as warm, soothing, urgent or tired. Realtime TTS-2 also supports voice cloning, advanced voice design, inline non-verbal markers and crosslingual switching across 200+ languages while preserving a single voice identity.
The model is part of Inworld’s broader realtime stack, which combines speech transcription, voice profiling, model routing and speech generation over a persistent connection. Realtime TTS-2 is available through the Inworld API and Inworld Realtime API, with Node and Python SDKs, REST access and WebSocket support. Customers using Realtime TTS 1.5 can upgrade by changing the model identifier.