MiniMax

MiniMax

MiniMax Speech Turbo

MiniMax Speech Turbo is a low-latency, multilingual text-to-speech model optimized for high-throughput and real-time conversational workloads. Reach for it when you need fast, cost-efficient audio generation for voice agents or bulk narration.

Modalities

Text → Audio

Price

$0.04 / 1K char

Context

5K

Usage

How much this model is actually called here.

Rank

#92

of 101 active models

Tokens served

56

all-time

Platform share

0.0%

of all tokens

Pricing

Per 1,000 characters of input text. Billed on successful synthesis.

$0.04 / 1K char

When to use MiniMax Speech Turbo

Updated 2026-07-31

Where this model earns its cost — and where it doesn't.

MiniMax Speech Turbo converts text input into audio output with a focus on speed and efficiency. It supports multiple languages and operates in the cheap cost tier and fast speed tier, making it suitable for applications that prioritize low latency and high request volume over studio-grade fidelity.

On Kyma, the model is accessible via an OpenAI-compatible endpoint using a single API key. It does not support prompt caching, and it includes automatic failover to reroute requests if a serving path degrades. Responses return exact billing metrics in the usage.cost field and report the active model in the X-Kyma-Model header.

The endpoint does not support reasoning, vision, or structured outputs, and it is strictly a speech generation interface. It is designed for scalable deployments where consistent throughput and fast delivery are the primary requirements.

Real-time voice agents

Generates rapid audio responses for interactive conversational systems.

Bulk narration pipelines

Processes large volumes of text into speech efficiently.

Conversational AI backends

Delivers fast, multilingual audio output for chat and voice assistants.

Scalable text to speech

Handles concurrent requests efficiently for high-volume audio generation.

Not ideal for: Do not use this model when you require studio-quality voice fidelity, complex audio editing, or structured JSON outputs.

Quick start

Up and running in under two minutes.

  1. 1

    Create an API key

    Sign up and grab a key from the dashboard. MiniMax Speech Turbo needs a top-up — the signup credit covers the free tier.

    Get API key →
  2. 2

    Make your first request

    Drop in your key and send a chat completion — fully OpenAI-compatible.

    curl https://kymaapi.com/v1/chat/completions \
      -H "Authorization: Bearer YOUR_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "model": "minimax-speech-turbo",
        "messages": [
          {"role": "user", "content": "Explain prompt caching in one paragraph."}
        ]
      }'
  3. 3

    Stream responses

    Add "stream": true to receive tokens as they arrive.

    curl https://kymaapi.com/v1/chat/completions \
      -H "Authorization: Bearer YOUR_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "model": "minimax-speech-turbo",
        "stream": true,
        "messages": [{"role": "user", "content": "Hello!"}]
      }'

FAQ

Common questions about this model.

How much does MiniMax Speech Turbo cost?

$0.04 per 1K char. Per 1,000 characters of input text. Billed on successful synthesis.

How do I use MiniMax Speech Turbo?

Kyma is OpenAI-compatible: point your SDK's base URL at https://kymaapi.com/v1, use your Kyma API key, and set the model to minimax-speech-turbo. Signing up is free, but this model needs a top-up: the $0.50 signup credit covers the free tier only.

Does this model support structured outputs or JSON formatting?

No, it only accepts text input and returns audio output.

How does prompt caching work for this endpoint?

This model does not support prompt caching.

What happens if the serving path experiences latency?

Kyma automatically reroutes the request to a healthy serving path to maintain delivery.

Needs a top-up — the signup credit covers the free tier.Create account →

More models by MiniMax

ModelContextInputOutput
MiniMaxMiniMax M31M$0.405$1.62
MiniMaxMiniMax M2.7205K$0.405$1.62
MiniMaxMiniMax M2.5197K$0.405$1.62
MiniMaxMiniMax Speech HD—$0.07 / 1K char
MiniMaxMiniMax Image 01—$0.005 / image