Kuaishou

Kuaishou

Kling 3 Pro (Audio)

Kling 3 Pro (Audio) generates video clips with synchronized native audio, including ambient sound and dialogue. Use it when your pipeline requires diegetic soundtracks or talking-head footage without post-production audio syncing.

Modalities

Text → Audio

Price

$0.2268 / sec

Speed

slow

Performance

Live production data from real requests on Kyma — not synthetic benchmarks.

Rank

#40

of 87 active models

Tokens served

160.8K

all-time

Total requests21
Platform share0.0%

Pricing

Per second of generated video. Failed jobs are refunded in full.

$0.2268 / sec

When to use Kling 3 Pro (Audio)

Updated 2026-07-31

Where this model earns its cost — and where it doesn't.

Created by Kuaishou, this model extends the visual output of Kling 3 Pro by generating synchronized audio tracks alongside video. It accepts text prompts and reference images, returning video and audio outputs in a single request.

Kyma serves this model through an OpenAI-compatible endpoint with automatic failover if a serving path degrades. Prompt caching is supported, billing repeated prompt prefixes at 10% of the standard input rate. New accounts receive a $0.50 free credit to test the endpoint. Every response includes exact billing in the usage.cost field and identifies the executed model via the X-Kyma-Model header.

The model operates on a premium cost tier and generates output at a slower speed. It does not support structured outputs, reasoning chains, or vision analysis tasks. Input is limited to text and images, while output consists exclusively of video and synchronized audio tracks.

Cinematic clips with ambient sound

Generate short film scenes that include synchronized environmental audio and background noise.

Talking head video generation

Produce character or spokesperson footage with matching dialogue audio tracks.

Product showcase with sound effects

Create promotional clips where visual actions align with generated sound effects.

Animated scene prototyping

Draft early video concepts with placeholder audio to evaluate pacing and mood.

Not ideal for: Do not use this model when you only need silent video generation or require fast, low-cost batch processing.

Quick start

Up and running in under two minutes.

  1. 1

    Create an API key

    Sign up and grab a key from the dashboard — $0.50 free credit, no card required.

    Get API key →
  2. 2

    Make your first request

    Submit a generation job and poll until it succeeds.

    curl https://kymaapi.com/v1/videos/generations \
      -H "Authorization: Bearer YOUR_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "model": "kling-3-pro-audio",
        "prompt": "A wave breaking on a rocky shore at golden hour, cinematic",
        "duration": 5
      }'
    # Poll: GET /v1/jobs/{id} until status="succeeded"

FAQ

Common questions about this model.

How much does Kling 3 Pro (Audio) cost?

$0.2268 per sec. Per second of generated video. Failed jobs are refunded in full.

How do I use Kling 3 Pro (Audio)?

Kyma is OpenAI-compatible: point your SDK's base URL at https://kymaapi.com/v1, use your Kyma API key, and set the model to kling-3-pro-audio. Signing up is free and includes $0.50 of credit — no card required.

How does prompt caching work with this model?

Kyma caches repeated prompt prefixes and bills them at 10% of the standard input rate, reducing costs for iterative generation workflows.

Can I use image inputs to guide the video generation?

Yes, the model accepts both text prompts and reference images to control visual composition and style.

What happens if the generation endpoint fails?

Kyma automatically reroutes the request to an alternative serving path, ensuring the request completes without manual retry logic.

Start with $0.50 free credit — no card required.Create account →

More models by Kuaishou

ModelContextInputOutput
KuaishouKling 3 Pro$0.1512 / sec
KuaishouKling 2.5 Pro$0.0945 / sec