Description
Pricing Details
Soniox offers fair and flexible pricing designed to scale with user needs, where users pay only for the Speech-to-Text and Text-to-Speech APIs they use, whether transcribing in real time or in batches across 60+ languages. The pricing calculator allows users to compare the costs of Speech-to-Text, Text-to-Speech, and Speech Translation, with options for Real-time, Async, and STS. The comparison is based on 1,000 hours of audio per month, with the clarification that pricing assumptions are based on public pay-as-you-go pricing, that enterprise discounts and committed-use contracts may differ, and that some service providers charge separately for certain features, while the calculator uses the general pricing structure of the provider that most closely matches Soniox. In the comparison, Soniox uses the stt-rt-v5 model at a price of $120, and the price was updated in June 2026 to $0.12/hr, while OpenAI uses the gpt-realtime-whisper model priced at $1,020, and the price was updated in June 2026 to $1.02/hr, The savings amount to $900; the page explains that at this volume, Soniox saves approximately $900 per month compared to OpenAI—that is, 88% less—with the option to “Compare on live audio” via https://soniox.com/compare-stt?providers=soniox,openai#try. As for the Speech-to-Text API, pricing is based on a token-based system, where all API costs are calculated based on tokens, with the cost amounting to approximately $0.10/hour for async (file) and approximately $0.12/hour for real-time (streaming) transcription. You can get started with STT at https://soniox.com/docs/stt/get-started. For async (file) transcription, the price for input audio tokens—which represent the duration of the audio or streaming session—is $1.50 per 1M tokens, the price for input text tokens—which represent the custom instructions or context you provide—is $3.50 per 1M tokens, while the price for output text tokens—which represent the transcription and, optionally, translation or other text returned by the model—is $3.50 per 1M tokens, As for Real-time (streaming), the price for Input audio tokens is $2.00 per 1M tokens, the price for Input text tokens is $4.00 per 1M tokens, and the price for Output text tokens is $4.00 per 1M tokens. The Usage Reference indicates that 1 hour of audio is approximately 30,000 input audio tokens, 1 hour of speech is approximately 15,000 output text tokens, and 1 character of output is approximately 0.3 tokens. The Text-to-Speech API also uses token-based pricing, where all API costs are calculated based on tokens; the cost is approximately $0.70 per hour of generated speech. You can get started with TTS at https://soniox.com/docs/tts/get-started, The price for input text tokens—which represent the text input used to generate speech—is $4.00 per 1 million tokens, while the price for output audio tokens—which represent the duration of generated audio—is $21.50 per 1 million tokens. The usage reference indicates that 1 character ≈ 0.3 input text tokens, that 15,000 input text tokens ≈ 1 hour of generated speech, and that 1 hour of generated speech ≈ 30,000 output audio tokens. Soniox explains that its low cost stems from the efficiency of its technology, as it has built a full speech AI stack in-house—from models to inference and real-time cloud infrastructure— with each layer optimized to process a larger volume of audio with shorter response times and less wasted computing resources. The company explains that this efficiency allows it to offer production-grade speech AI at a fraction of the cost of traditional providers. It also explains that Soniox’s models are built from scratch for real-time speech understanding and generation, rather than being adapted from general-purpose models that waste computational resources, and that its custom inference engine is built for low-latency audio streaming, batching, scheduling, and GPU utilization, enabling the devices themselves to process a larger volume of audio at a lower cost. The Soniox platform is designed to efficiently handle hundreds of thousands of concurrent streams, which translates to lower prices per customer. The Frequently Asked Questions section includes questions such as “How much does the Soniox API cost?” “How does Soniox compare to Google, Azure, and OpenAI in terms of price?” “Is translation included, or billed separately?” “Is Soniox cheaper than Deepgram?” “Do I pay extra for diarization, language detection, or formatting?” “What is the difference between real-time and asynchronous pricing?” and “How much does Soniox Text-to-Speech cost?” You can get started with the platform by creating an account right away or reaching out to design a custom business package, with the “Build with API” option available at https://console.soniox.com/signup?ref=019a58ea-51b6-7a99-947a-6d01b4e0a2d4. Documentation is also available to help you get started in minutes and spend your time building instead of dealing with the API, You can explore the documentation at https://soniox.com/docs. There’s also a “See what you’ll pay” section that lets you pay only for what you use through flexible pricing designed for scalability, with pricing details available at https://soniox.com/pricing. The APIs include the Platform at https://soniox.com/platform, the Speech-to-Text API at https://soniox.com/speech-to-text, the Text-to-Speech API at https://soniox.com/text-to-speech, the Speech Translation API at https://soniox.com/speech-translation, Documentation at https://soniox.com/docs, Benchmarks at https://soniox.com/benchmarks, Compare at https://soniox.com/compare-stt, and Pricing at https://soniox.com/pricing, while Use Cases include Call Center at https://soniox.com/speech-to-text/use-cases/call-center, Medical Transcription at https://soniox.com/speech-to-text/use-cases/medical-transcription, Media Transcription at https://soniox.com/speech-to-text/use-cases/media-transcription, Speech analytics at https://soniox.com/speech-to-text/use-cases/speech-analytics, Speech translation at https://soniox.com/speech-to-text/use-cases/speech-translation, Voice Agents via https://soniox.com/speech-to-text/use-cases/voice-agents, and Wearables via https://soniox.com/speech-to-text/use-cases/wearables, The App category includes Home at https://soniox.com/soniox-app, Smart Scribe at https://soniox.com/soniox-app/smart-scribe, Translator at https://soniox.com/soniox-app/translator, Voice Typing at https://soniox.com/soniox-app/voice-typing, Business at https://soniox.com/soniox-app/business, and Pricing at https://soniox.com/soniox-app/pricing, and Help Center at https://app.soniox.com/help-center; Resources includes Documentation at https://soniox.com/docs, the Blog at https://soniox.com/blog, Videos at https://www.youtube.com/@soniox_ai/videos. The “Company” section includes “About Us” at https://soniox.com/about, “Contact” at https://soniox.com/contact, "Careers" at https://soniox.com/careers, "Brand" at https://soniox.com/brand, and "Policies" at https://soniox.com/policies. Soniox describes itself as “The voice platform for every language,” It also includes Subscribe and links to LinkedIn at https://www.linkedin.com/company/soniox/, GitHub at https://github.com/soniox, X at https://x.com/soniox_ai, Discord at https://discord.gg/rWfnk9uM5j, Soniox in the App Store at https://apps.apple.com/us/app/soniox/id1560199731, and Soniox on Google Play at https://play.google.com/store/apps/details?id=com.soniox.sonioxmobileapp. Copyright © 2026 Sonio.
