MiniMax Speech Turbo is a low-latency, multilingual text-to-speech model optimized for high-throughput and real-time conversational workloads. Reach for it when you need fast, cost-efficient audio generation for voice agents or bulk narration.
Modalities
Text → Audio
Price
$0.09 / 1K char
Speed
fast
Performance
Live production data from real requests on Kyma — not synthetic benchmarks.
Rank
#82
of 87 active models
Tokens served
56
all-time
Top apps using this model
Public apps sending the most traffic to this model — a signal of what real workloads it fits.
Pricing
Per 1,000 characters of input text. Billed on successful synthesis.
$0.09 / 1K charWhen to use MiniMax Speech Turbo
Where this model earns its cost — and where it doesn't.
MiniMax Speech Turbo converts text input into audio output with a focus on speed and efficiency. It supports multiple languages and operates in the cheap cost tier and fast speed tier, making it suitable for applications that prioritize low latency and high request volume over studio-grade fidelity.
On Kyma, the model is accessible via an OpenAI-compatible endpoint using a single API key. It supports prompt caching, which applies a 90% discount to repeated input prefixes, and includes automatic failover to reroute requests if a serving path degrades. Responses return exact billing metrics in the usage.cost field and report the active model in the X-Kyma-Model header.
The endpoint does not support reasoning, vision, or structured outputs, and it is strictly a speech generation interface. It is designed for scalable deployments where consistent throughput and fast delivery are the primary requirements.
Real-time voice agents
Generates rapid audio responses for interactive conversational systems.
Bulk narration pipelines
Processes large volumes of text into speech efficiently.
Conversational AI backends
Delivers fast, multilingual audio output for chat and voice assistants.
Scalable text to speech
Handles concurrent requests efficiently for high-volume audio generation.
Not ideal for: Do not use this model when you require studio-quality voice fidelity, complex audio editing, or structured JSON outputs.
Quick start
Up and running in under two minutes.
- 1
Create an API key
Sign up and grab a key from the dashboard — $0.50 free credit, no card required.
Get API key → - 2
Make your first request
Drop in your key and send a chat completion — fully OpenAI-compatible.
curl https://kymaapi.com/v1/chat/completions \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "minimax-speech-turbo", "messages": [ {"role": "user", "content": "Explain prompt caching in one paragraph."} ] }' - 3
Stream responses
Add
"stream": trueto receive tokens as they arrive.curl https://kymaapi.com/v1/chat/completions \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "minimax-speech-turbo", "stream": true, "messages": [{"role": "user", "content": "Hello!"}] }'
FAQ
Common questions about this model.
How much does MiniMax Speech Turbo cost?
How do I use MiniMax Speech Turbo?
Does this model support structured outputs or JSON formatting?
How does prompt caching work for this endpoint?
What happens if the serving path experiences latency?
More models by MiniMax
See all 13 →| Model | Context | Input | Output |
|---|---|---|---|
MiniMax M3 | 1M | $0.3852 | $1.54 |
MiniMax Music Pro | — | $0.21 / song | |
MiniMax M2.7 | 205K | $0.405 | $1.62 |
MiniMax M2.5 | 197K | $0.3826 | $1.346 |
MiniMax Music | — | $0.045 / song | |
MiniMax Voice Design | — | $4.20 / call | |
Hailuo 02 (1080p) | — | $0.78 / video | |
Hailuo 02 (768p) | — | $0.42 / video | |
