Inworld TTS

Description
️ Tool Name: 🖼
Inworld Realtime TTS
Categories: 🔖
Audio, Speech, and Music
Integrations and APIs
Text-to-Speech / Speech-to-Text
Voice Cloning
️ What does this tool offer? ✏
Inworld Realtime TTS is Inworld’s real-time text-to-speech technology, providing voice models tailored for applications that require fast, realistic, and interactive speech generation. The platform includes Realtime TTS-2, the current flagship model, and Realtime TTS-2 Flash, designed for applications requiring low latency and lower cost, in addition to previous Realtime TTS models such as Realtime TTS 1 and Realtime TTS 1-Max. Realtime TTS-2 offers advanced vocal expression and the ability to control delivery style using natural language instructions; it also supports nonverbal cues such as laughter, sighs, breathing, coughing, yawning, and throat-clearing, and a single voice can be used across multiple languages. The platform offers instant voice cloning using a short voice sample, as well as Voice Design for creating custom voices, and supports control over speech rate, pitch, custom pronunciation, timestamps, and alignment data. Realtime TTS-2 offers a latency of less than 100 ms based on server-side TTFB measurements, while Realtime TTS-2 Flash achieves 25 ms, according to the official website. The current models also support more than 200 languages with multilingual capabilities and the ability to switch between languages using BCP-47 codes.
Realtime TTS can be used via the API and Realtime API, and you can try out the voices and features directly through the TTS Playground. Inworld provides integration interfaces for real-time voice applications, including WebSocket, with WebRTC and SIP support in early access, according to the official website. The technology can be used in applications including voice agents, consumer apps, educational apps, games, virtual companions, and other applications that require interactive real-time voice.
Realtime TTS 1, announced by Inworld in 2025, provided realistic, context-aware voice synthesis and voice cloning from short samples without prior training, and was available via API and TTS Playground, while Realtime TTS 1-Max was available in beta. Currently, Inworld offers Realtime TTS-2 as its flagship model, with Realtime TTS-2 Flash as a faster option focused on reducing latency and cost.
Inworld also provides developer tools, including an API Reference, Documentation, and TTS Playground, as well as an open-source training framework linked to Realtime TTS research, allowing developers to rebuild TTS models according to the methodology published by the company. The platform also offers integration capabilities with voice application systems, and its models can be used directly through APIs.
What does it actually offer based on user experience? ⭐
• It provides fast speech synthesis suitable for applications that require immediate response, with a TTFB of less than 100 ms in Realtime TTS-2 and 25 ms in Realtime TTS-2 Flash, according to measurements published by Inworld.
• Realtime TTS-2 offers advanced voice expression with the ability to control tone and emotions using natural language instructions directly within the text.
• It allows for the addition of nonverbal cues such as laughter, sighs, breathing, coughing, yawning, and throat clearing to the generated speech.
• It enables instant voice cloning using a short voice sample; Inworld notes that voice cloning can be performed using a clean sample from a single speaker lasting 5 to 15 seconds.
• Voice Design allows you to create a custom voice from a written description, without the need for a reference audio sample.
• It supports multilingual use; the current website states that it supports over 200 languages, with the ability to use a single voice across multiple languages.
• According to Inworld, the quality of Realtime TTS is evaluated through blind listening tests involving thousands of users, The company also notes that Realtime TTS-2 has been fine-tuned to achieve greater expressiveness while reducing issues such as hallucination and word dropouts and improving conversational naturalness. These are published results and descriptions from Inworld and do not constitute an independent evaluation by me.
• The platform allows you to test voices directly through the TTS Playground before integrating them into apps, then transition to the API once you begin development.
Does it include automation? 🤖
The officially available information does not mention a standalone automation system within Realtime TTS in the traditional sense of task automation tools. However, the tool provides an API and a Realtime API, allowing developers to integrate speech synthesis into their applications and automate programming workflows. The Realtime API also allows you to run speech layers, text-to-speech, and other features within a single session, with the option to use WebSocket, while WebRTC and SIP are available in early access.
Pricing Model: 💰
Freemium / Paid / Subscription / Pay-as-you-go. Inworld offers an On-Demand plan that you can start for free, as well as monthly subscription plans—including Creator, Builder, Developer, and Growth—in addition to a custom-priced Enterprise plan. The cost per character decreases at higher tiers, and the company offers an annual payment option with two free months. TTS usage is billed by character count, STT is billed by the hour, and LLM models are billed based on the provider’s rates.
🆓 Free Plan Details:
| Item | Details |
|---|---|
| Plan | On-Demand |
| Price | Start for free |
| TTS-2 | $25 per 1 million characters |
| TTS-2 Flash | $15 per 1 million characters |
| STT 1 | $0.15 per hour |
| LLMs | At cost |
| Included TTS | Up to 70 minutes |
| Custom voices | 100 |
| Voice cloning and voice design | Available |
| Real-time API | Available |
| LLM Models via Router | Over 220 models |
| Commercial License | Available |
| Support | Community support |
| Access | On-Demand can be selected from the billing page |
The pricing page notes that On-Demand is suitable for evaluation and prototyping, and that it provides up to 70 minutes of TTS for free. The official cost calculator also shows that the recommended plan when selecting Text-to-Speech with 0 minutes of usage is On-Demand, and the total cost of the plan is Free.
Paid plan details: 💳
| Plan | Monthly Price | Realtime TTS-2 | Realtime TTS-2 Flash | Monthly Credit | Key Features |
|---|---|---|---|---|---|
| Creator | $25/mo | $20 per 1M characters | $10 per 1M characters | $25 | Up to 33% off, 500 custom voices, 40K characters per request in TTS Playground, and the ability to create and share Workspaces and manage your team |
| Builder | $100/mo | $17.50 per 1M characters | $9 per 1M characters | $100 | Up to 40% off, 3,000 custom voices, higher synchronization limits, and the ability to create and share Workspaces and manage your team |
| Developer | $300/mo | $15 per 1M characters | $8 per 1M characters | $300 | Up to 47% off, 10,000 custom voices, higher synchronization limits, professional voice cloning as an add-on, and priority email support |
| Growth | $1,500/mo | $12.50 per 1M characters | $7 per 1M characters | $1,500 | Up to 53% off, 30,000 custom voices, higher API limits, one professional voice cloning, and ZDR, HIPAA, and BAA add-ons |
| Enterprise | Custom | Starting at $5 per 1M characters | Less than $5 per 1M characters | Custom | Custom pricing and limits, price match, SLA and DPA, data residency in the EU and India, and a dedicated account manager and Slack channel |
The official pricing page shows that TTS-2 prices range from $25 per million characters for On-Demand to $20 for Creator, $17.50 for Builder, $15 in Developer, $12.50 in Growth, and as low as $5 in Enterprise. As for TTS-2 Flash, prices start at $15 and go down to $10, $9, $8, and $7, respectively, and then to less than $5 in the Enterprise tier. Additional usage is billed at the same rate as the plan.
| Feature | On-Demand | Creator | Builder | Developer | Growth | Enterprise |
|---|---|---|---|---|---|---|
| API Access | Available | Available | Available | Available | Available | Available |
| Audio downloads | Available | Available | Available | Available | Available | Available |
| Character limit in TTS Playground | 2,000 | 40,000 | 40,000 | 40,000 | 40,000 | Custom |
| Custom voices | 100 | 500 | 3,000 | 10,000 | 30,000 | Custom |
| Instant voice cloning | Available | Available | Available | Available | Available | Available |
| Professional voice cloning | — | — | — | Add-on | 1 | Custom |
| Voice design | Available | Available | Available | Available | Available | Available |
| Steering | Available | Available | Available | Available | Available | Available |
| Speaking rate control | Available | Available | Available | Available | Available | Available |
| Temperature control | Available | Available | Available | Available | Available | Available |
| 200+ languages | Available | Available | Available | Available | Available | Available |
| Custom pronunciation | Available | Available | Available | Available | Available | Available |
| Timestamps | Available | Available | Available | Available | Available | Available |
| Concurrent requests | 5 | 10 | 50 | 150 | 500 | Custom |
| Estimated concurrent user sessions | 20 | 40 | 200 | 600 | 2,000 | Custom |
| GDPR & SOC 2 Type II | Available | Available | Available | Available | Available | Available |
| Commercial license | Available | Available | Available | Available | Available | Available |
| Credit rollover | Available | Available | Available | Available | Available | Available |
| HIPAA & BAA | — | — | — | — | Add-on | Available |
| Zero data retention | — | — | — | — | Add-on | Available |
| SLA & DPA | — | — | — | — | — | Available |
| Data residency | — | — | — | — | — | EU, India |
| Payment | — | Credit card | Credit card | Credit card | Credit card | Invoicing / PO |
| Support | Community | Community | Community | Priority Email | Priority email | AM + Slack |
The official pricing page explains that the estimated number of concurrent user sessions varies depending on usage, but is typically at least four times the concurrent request limit. It also explains that paid credits are subject to specific expiration terms, and that rollover or cancellation details depend on the subscription status and billing cycle.
How to access the tool: 🧭
| Access Method | Details |
|---|---|
| Web | You can visit the official Inworld website and try the service via TTS Playground |
| TTS Playground | Try out the voices, settings, and voice synthesis right away |
| API | Integrate Realtime TTS into your applications via the API |
| Realtime API | Build real-time voice applications and run voice layers and models within a single session |
| WebSocket | Available in the Realtime API |
| WebRTC | Available in Early Access |
| SIP | Available in Early Access |
| Documentation | Documentation and API Reference are available to developers |
| Enterprise Plans | Contact the sales team for pricing and custom limits |
Inworld offers the ability to try out Realtime TTS through the TTS Playground. You can also create an API key and start making API requests. According to the official website, the service currently supports audio outputs including MP3, Linear PCM/WAV, and Opus.
Demo link or official website: 🔗
Inworld Realtime TTS – Official Website