Inworld Realtime TTS 2 API Usage Guide
Overview
Inworld Realtime TTS 2 is Inworld’s next-generation text-to-speech model, delivering higher quality audio with lower latency. It supports the same 65 voices across 16 languages as TTS 1.5, with enhanced naturalness and expressiveness. Includes phoneme-level timing and viseme symbols for lip-sync animation.Key Features:
- 65 Voices across 16 languages
- Higher Quality: Improved naturalness and expressiveness over TTS 1.5
- Lower Latency: Optimized for real-time applications
- Word & Character Timestamps: Optional alignment metadata for captions and highlights
- Multiple Formats: MP3, WAV, OGG_OPUS, FLAC, ALAW, MULAW
- Text Normalization: Automatic expansion of numbers, dates, abbreviations
Available Voices (65 total)
English (25 voices)
Chinese 中文 (4 voices)
Japanese 日本語 (2 voices)
Korean 한국어 (4 voices)
Other Languages
- French (4): Alain, Hélène, Mathieu, Étienne
- German (2): Johanna, Josef
- Spanish (4): Diego, Lupita, Miguel, Rafael
- Italian (2): Gianni, Orietta
- Portuguese-BR (2): Heitor, Maitê
- Russian (4): Svetlana, Elena, Dmitry, Nikolai
- Dutch (4): Erik, Katrien, Lennart, Lore
- Polish (2): Szymon, Wojciech
- Hindi (2): Riya, Manoj
- Hebrew (2): Yael, Oren
- Arabic (2): Nour, Omar
Authentication
All API requests require Bearer authentication:Submit TTS Request
Endpoint
Request Format
Response
Inworld TTS 2 is synchronous and returns results immediately.Pricing
- Pricing Type: Per character
- Price: Contact for pricing
- Unit: Characters