Text To Speech Updates & Release Notes
5 updates curated from 1 source by the Releasebot Team. Last updated: Aug 11, 2026
- Aug 9, 2026
- Date parsed from source:Aug 9, 2026
- First seen by Releasebot:Aug 11, 2026
Realtime TTS-2 Flash
Text To Speech launches Realtime TTS-2 Flash, its fastest and most cost-efficient TTS-2 model, delivering 20 ms time to first audio, 200+ languages and locales, instant voice cloning, timestamp alignment, and support for non-verbal tags like [laugh].
Launched Realtime TTS-2 Flash (
inworld-tts-2-flash), the fastest member of the TTS-2 family — see Models:- Our lowest latency: 20 ms time to first audio (server-side P90 TTFB, excluding network latency) — 5× faster than
inworld-tts-2at 100 ms, making it the best choice for latency-critical real-time agents. - Our lowest cost: The most cost-efficient model per character, ideal for high-volume workloads.
- Full TTS-2 language coverage: The same 200+ languages and locales as
inworld-tts-2, plus instant voice cloning and timestamp alignment.
Steering instructions and Professional Voice Cloning are supported on
Original sourceinworld-tts-2only — use it when you need directed, contextually aware delivery. Non-verbal tags like[laugh]work on both models. - Aug 6, 2026
- Date parsed from source:Aug 6, 2026
- First seen by Releasebot:Aug 11, 2026
Steering instructions now persist
Text To Speech adds clearer steering for inworld-tts-2 with tags that stay active until changed, a new [reset] tag to return to the voice’s natural delivery, pause-proof instructions, and a request-level instruction field for whole-request control.
Steering on
inworld-tts-2follows one rule: a[tag]applies from where you write it until you change it. See the Steering guide.Behavior change
An inline
[tag]previously affected only the text immediately after it and delivery could revert on its own partway through longer text. A tag now stays in force until you change it. If you relied on an instruction wearing off — for example[shout] Hi. Normal text.expecting the second sentence unstyled — add[reset]where normal delivery should resume. Requests that use no inline tags are unaffected, as areinworld-tts-1.5-maxandinworld-tts-1.5-mini.[reset]New reserved tag that ends a styled passage and returns the voice to its own character for the rest of the text.
[shouting] We need to leave now! [reset] Do you understand me?shouts only the first sentence.Instructions survive pauses
A
<break/>no longer clears the active instruction. A pause is a pause and never changes delivery.Request-level
instructionfieldSet one instruction for the whole request without putting tags in your text. See
Original sourceinstruction. Use either this field or inline tags, not both. All of your release notes in one feed
Join Releasebot and get updates from Inworld and hundreds of other software products.
- May 5, 2026
- Date parsed from source:May 5, 2026
- First seen by Releasebot:May 5, 2026
Realtime TTS-2
Text To Speech launches Realtime TTS-2, its most expressive TTS model, with natural language steering, stronger multilingual synthesis across 15 languages, cross-lingual voice reuse, voice localization, a new deliveryMode control, and an updated Voice Design with improved generations.
Launched Realtime TTS-2 (
inworld-tts-2), our most powerful and expressive TTS model:- Natural Language Steering: Direct any voice with bracketed instructions like
[say excitedly],[whisper in a hushed style], or free-form directions like[speak as if barely holding back rage]. Covers articulation, intonation, volume, pitch, range, speed, vocal style, and non-verbals ([laugh],[sigh], etc.). See the Steering guide. - Stronger Multilingual Support: Production-quality synthesis across 15 languages, plus experimental support for 90+ additional languages. See Languages.
- Cross-Lingual Voice Synthesis: Reuse the same voice across multiple languages. For best results, specify the
languagefield. - Voice Localization: Localize your voice for the most consistent, native-sounding speech in a target language. See Voice Localization.
- Delivery Mode: New
deliveryModefield (STABLE,BALANCED,EXPRESSIVE) controls the trade-off between consistency and emotional range. - Updated Voice Design: Released an updated version of Voice Design with improved generations. See Voice Design.
- Jan 21, 2026
- Date parsed from source:Jan 21, 2026
- First seen by Releasebot:Jan 21, 2026
- Modified by Releasebot:Apr 20, 2026
Inworld TTS 1.5
Text To Speech launches Inworld TTS 1.5, a new realtime model generation with faster first-audio latency, more expressive and stable speech, and support for additional languages including Hindi, Arabic, and Hebrew.
Launched Inworld TTS 1.5, our newest generation of realtime TTS models featuring:
- Two New Models: Our flagship model
inworld-tts-1.5-maxis ideal for most use cases, with the best balance of quality and speed. For use cases where latency is the top priority, we also offerinworld-tts-1.5-mini. - Latency Improvements: Our new TTS-1.5 models achieve P90 latency for first audio chunk delivery under 250ms for our Max model and under 130ms for our Mini model, a 4x improvement compared to TTS-1.
- More Expressive and More Stable: TTS-1.5 is 30% more expressive than prior generations and demonstrates a 40% reduction in word error rates.
- Additional Languages: We've added support for additional languages, including Hindi, Arabic, and Hebrew, bringing total languages supported to 15.
- Aug 22, 2025
- Date parsed from source:Aug 22, 2025
- First seen by Releasebot:Dec 23, 2025
- Modified by Releasebot:Apr 20, 2026
Updates to Inworld TTS
Text To Speech releases upgraded Inworld TTS models with clearer speech, better voice similarity, stronger multilingual output and inline IPA.
Released an upgraded version of the Inworld TTS models with higher overall quality.
- Speech Quality: Clearer, more natural speech with smoother pacing and more accurate pronunciation.
- Voice Similarity: Cloned voices sound closer to the originals, preserving each voice’s unique style.
- Non-English Languages: More consistent, reliable output across supported non-English languages.
- Custom Pronunciation: New support for inline IPA, giving you control over exact word pronunciations. See the Key Features for details.
Similar to Text To Speech with recent updates:
- Fish Audio updates15 release notes · Latest Mar 10, 2026
- Hume updates43 release notes · Latest May 16, 2026
- ChatGPT updates216 release notes · Latest Sep 3, 2026
- OpenAI Models updates49 release notes · Latest Aug 18, 2026
- Claude updates136 release notes · Latest Sep 2, 2026
- Eleven Labs updates90 release notes · Latest Aug 31, 2026
This is the end. You've seen all the release notes in this feed!
Curated by the Releasebot team
Releasebot is an aggregator of official product update announcements from hundreds of software vendors and thousands of sources.
Our editorial process involves the manual review and audit of release notes procured with the help of automated systems.