Models

106 models, one API key. Compare them by what they do and what they cost.

CompareModelPrecisionCapabilities
500K$2.16$6.48$0.54
100%
32K$0.0567$0.00—
100%
1M$0.1048$0.3141—fp435tok/s1.24s
100%
—$0.019 / image8.9s/img8.87s
—$0.108 / song
1M$13.50$67.50$1.3525tok/s1.52s
40.5%
1M$1.013
$2.026
$5.063
$10.126
—31tok/s2.19s
99.2%
1M$1.688$5.738—
1M$13.50$67.50—14tok/s3.06s
100%
—$0.006075 / min4.8×rt2.31s
100%
—$0.00675 / min2.3×rt4.45s
85.9%
1M$0.2045$0.641—26tok/s5.11s
92.1%
1M$0.1791$0.5972—32tok/sfrom traffic, 7d6.70sfrom traffic, 7d
99.9%
1M$0.2514$0.7542—49tok/s1.96s
97.1%
1M$1.178$3.702—32tok/s3.02s
100%
1M$0.3576$2.554—36tok/s2.46s
98.9%
1M$1.013
$2.026
$5.063
$10.126
$0.05062578tok/s3.16s
99.0%
500K$2.758$8.274—61tok/s5.21s
100%
131K$0.405$1.62—67tok/s4.57s
100%
1M$1.773$6.025—114tok/s4.14s
100%
1M$2.228$6.684$0.278127tok/s2.72s
100%
1M$0.1859$0.3718—27tok/s3.02s
100%
1M$0.0504$0.2187$0.008168tok/s4.95s
96.6%
1M$6.75$33.75$0.67522tok/s1.76s
100%
1M$1.013$5.063$0.1012524tok/s2.71s
90.9%
1M$0.405$3.375$0.040545tok/s1.38s
99.3%
1M$1.542$5.243—178tok/s1.93s
100%
1M$3.768$18.837$0.40552tok/sfrom traffic, 7d2.79sfrom traffic, 7d
98.3%
1M$2.70
$5.40
$13.50
$27.00
$0.337554tok/s3.70s
100%
1M$3.151
$5.40
$15.753
$27.00
$0.39387527tok/s1.78s
96.6%
1M$2.70$16.20$0.2768tok/s3.08s
100%
1M$2.70$16.20$0.2736tok/s1.26s
86.2%
1M$0.3127$1.877$0.02766tok/sfrom traffic, 7d3.73sfrom traffic, 7d
98.3%
1M$0.1548$0.9288$0.02775tok/s3.25s
89.6%
500K$2.73$8.189—48tok/s2.36s
100%
262K$0.108$0.4455—63tok/s4.68s
100%
1M$2.70$13.50$0.2720tok/s3.75s
91.3%
1M$1.101$3.76$0.29835fp4 · fp873tok/s2.66s
100%
262K$1.009$4.774—33tok/s3.07s
100%
1M$13.50$67.50—15tok/s3.15s
100%
1M$0.675$3.375$0.135fp4 · fp835tok/s3.63s
100%
1M$0.4431$1.773$0.086481tok/s14.72s
100%
1M$0.405$1.62$0.0756fp4 · fp838tok/s2.06s
99.6%
256K$0.27$1.552$0.054fp853tok/s3.04s
100%
1M$1.863$5.586$0.1687576tok/s11.27s
100%
256K$1.389$2.777$0.2797tok/s3.15s
100%
1M$1.928$11.563$0.2025114tok/s2.50s
97.7%
1M$1.92$3.838$0.2763tok/s2.96s
100%
1M$0.1264$0.2526$0.024326tok/s1.64s
100%
—$0.072 / image35s/img35.29s
97.9%
262K$0.7856$3.667—29tok/s1.57s
100%
1M$6.75$33.75$0.67524tok/s1.69s
100%
203K$1.89$5.94$0.27675fp4 · fp828tok/s13.86s
100%
1M$0.454$2.724—53tok/s10.97s
99.9%
128K$0.0763$0.218$0.135fp4 · fp88tok/s6.85s
99.7%
2M$1.767$3.534—365tok/s5.67s
100%
2M$1.763$3.528—34tok/s1.58s
100%
205K$0.405$1.62—51tok/s5.66s
100%
—$0.061 / image11s/img10.88s
1M$2.70$16.20$0.2796tok/s4.71s
100%
—$0.108 / image
1M$4.05$20.25$0.40529tok/s2.05s
99.0%
—$0.054 / image13s/img13.11s
—$0.3375 / image
197K$0.405$1.62—48tok/s3.27s
97.7%
—$0.1512 / sec
203K$0.081$0.54$0.0135fp849tok/s5.17s
100%
1M$0.675$4.05$0.067536tok/s1.65s
99.4%
160K$0.518$0.7571—fp412tok/s1.18s
97.5%
—$0.15525 / sec
—$0.0405 / image12s/img12.24s
—$0.135 / sec
200K$1.35$6.75$0.13534tok/s1.75s
99.1%
2K$0.0027$0.00—1096tok/s0.37s
100%
128K$0.0494$0.2406—27tok/s5.04s
98.3%
131K$0.1945$1.272$0.03375fp841tok/s3.17s
100%
131K$0.3586$1.629$0.135fp4 · fp847tok/s1.13s
99.9%
33K$0.0135$0.00—299tok/s1.37s
100%
33K$0.0135$0.00—1228tok/s0.33s
33K$0.027$0.00—628tok/s0.65s
41K$0.135$0.00—
—$0.054 / image
33K$0.108$0.378—26tok/s20.70s
73.8%
1M$0.27$1.08—fp834tok/s1.08s
100%
—$0.07 / 1K char39ch/s3.25s
—$0.04 / 1K char40ch/s3.15s
—$0.081 / image
200K$4.05$20.25—23tok/s1.82s
100%
127K$1.35$1.35—29tok/s1.56s
100%
—$0.005 / image31s/img31.35s
97.8%
64K$0.7425$2.957—55tok/s1.60s
100%
128K$0.135$0.432—14tok/s1.15s
88.6%
—$0.2025 / 1K char
—$0.0009 / min7.7×rt1.31s
99.4%
131K$1.35$1.35—fp817tok/s0.89s
100%
131K$0.945$0.945—fp821tok/s0.96s
100%
—$0.2025 / 1K char
—$0.00405 / min0.41sfrom traffic, 7d
100%
1K$0.0135$0.00—592tok/s0.87s
8K$0.0135$0.00—1311tok/s0.40s
100%
1K$0.0135$0.00—1067tok/s0.38s
1K$0.00675$0.00—1393tok/s0.29s
—$0.405 / 1K char
100%
1K$0.00675$0.00—475tok/s0.85s
1K$0.00675$0.00—92tok/s1.40s
1K$0.00675$0.00—265tok/s0.97s

Token models priced per 1M tokens. The cached column is what a repeated prompt prefix costs on that model — a rate of its own, read from the same source as the input rate beside it, not a fixed fraction of it. It runs from a tenth of the input rate to the full input rate depending on the model, and a dash means nobody publishes one, which is a different statement from “free”. Image, video, and audio models priced per unit (image, second, video, minute, song, call, or 1K characters). Throughput and latency are a probe with an identical input when Kyma has that sample. If a row is labeled “from traffic, 7d”, that number is production traffic instead — real prompts, not comparable across models. Image throughput is seconds per succeeded job (30-day median). Video, music and realtime stay empty until a job exists. An empty throughput, latency or uptime cell means Kyma does not yet have a number for that cell. Uptime is a percentage only after 20 observations, over 30 days, probes and real requests together, including Kyma's own accounts. Rankings demand is the same total. Precision names the numeric formats a model may be served at that are BELOW the one its creator released it in, and it is filled on the 14 models where that is true. A low format is not a downgrade when it is how the model shipped, so this is a comparison against the creator’s own model card rather than a list of small numbers. An empty precision cell means there is nothing to flag: either no route reports going under the release, or nobody could source what the release was. Those two are different answers and the model’s own page prints which one it is.

Start building with Kyma API

$0.50 signup credit for eligible free-tier models after email verification.

Create account