MiniMax Speech HD is a cost-effective, multilingual text-to-speech model built for production audio generation. Reach for it when you need expressive voice synthesis at a cheap tier without sacrificing output quality.
Modalities
Text → Audio
Price
$0.14 / 1K char
Speed
medium
Performance
Live production data from real requests on Kyma — not synthetic benchmarks.
Rank
#76
of 87 active models
Tokens served
56
all-time
Top apps using this model
Public apps sending the most traffic to this model — a signal of what real workloads it fits.
Pricing
Per 1,000 characters of input text. Billed on successful synthesis.
$0.14 / 1K charWhen to use MiniMax Speech HD
Where this model earns its cost — and where it doesn't.
MiniMax Speech HD converts text into high-quality audio across multiple languages. It delivers expressive voice synthesis with a 5,000-token context window and is optimized for strong audio quality at a cheap tier.
On Kyma, the model runs through an OpenAI-compatible endpoint using a single API key. Every request includes automatic failover if a serving path degrades, and responses return exact billing data in usage.cost alongside an X-Kyma-Model header. Prompt caching is supported, billing repeated prefixes at 10% of the standard input rate.
The model does not support reasoning, vision, or structured outputs. It operates at a medium speed tier, making it suitable for batch processing and asynchronous audio pipelines rather than low-latency conversational applications.
Multilingual Content Narration
Generate expressive voiceovers for videos and podcasts across multiple languages.
Budget Brand Voiceovers
Produce consistent audio assets for marketing campaigns at a lower cost tier.
Audio Translation Workflows
Convert localized text into natural-sounding speech for global distribution.
Extended Audiobook Generation
Process long text passages into continuous audio using the full context window.
Not ideal for: It is not suitable for real-time conversational voice agents or applications requiring sub-second audio latency.
Quick start
Up and running in under two minutes.
- 1
Create an API key
Sign up and grab a key from the dashboard — $0.50 free credit, no card required.
Get API key → - 2
Make your first request
Drop in your key and send a chat completion — fully OpenAI-compatible.
curl https://kymaapi.com/v1/chat/completions \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "minimax-speech-hd", "messages": [ {"role": "user", "content": "Explain prompt caching in one paragraph."} ] }' - 3
Stream responses
Add
"stream": trueto receive tokens as they arrive.curl https://kymaapi.com/v1/chat/completions \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "minimax-speech-hd", "stream": true, "messages": [{"role": "user", "content": "Hello!"}] }'
FAQ
Common questions about this model.
How much does MiniMax Speech HD cost?
How do I use MiniMax Speech HD?
Does this model support prompt caching?
What happens if the serving path fails during generation?
Can I use this for real-time voice conversations?
More models by MiniMax
See all 13 →| Model | Context | Input | Output |
|---|---|---|---|
MiniMax M3 | 1M | $0.3852 | $1.54 |
MiniMax Music Pro | — | $0.21 / song | |
MiniMax M2.7 | 205K | $0.405 | $1.62 |
MiniMax M2.5 | 197K | $0.3826 | $1.346 |
MiniMax Music | — | $0.045 / song | |
MiniMax Voice Design | — | $4.20 / call | |
Hailuo 02 (1080p) | — | $0.78 / video | |
Hailuo 02 (768p) | — | $0.42 / video | |
