Models

106 models, one API key. Compare them by what they do and what they cost.

CompareModelPrecisionCapabilities
500K$2.16$6.48$0.54
100%
32K$0.0567$0.00—
100%
1M$0.1048$0.3141—fp435tok/s1.24s
100%
—$0.019 / image8.9s/img8.87s
—$0.108 / song
1M$13.50$67.50$1.3525tok/s1.52s
40.5%
1M$1.013
$2.026
$5.063
$10.126
—32tok/s2.22s
99.2%
1M$1.688$5.738—
1M$13.50$67.50—15tok/s2.81s
100%
—$0.006075 / min4.7×rt2.33s
100%
—$0.00675 / min2.3×rt4.45s
92.9%
1M$0.2045$0.641—26tok/s5.11s
92.3%
1M$0.1791$0.5972—32tok/sfrom traffic, 7d6.58sfrom traffic, 7d
99.9%
1M$0.2514$0.7542—51tok/s1.91s
97.0%
1M$1.178$3.702—32tok/s2.90s
100%
1M$0.3576$2.554—36tok/s2.46s
98.8%
1M$1.013
$2.026
$5.063
$10.126
$0.05062577tok/s2.96s
98.9%
500K$2.758$8.274—61tok/s5.23s
131K$0.405$1.62—61tok/s5.55s
100%
1M$1.773$6.025—122tok/s3.63s
100%
1M$2.228$6.684$0.278129tok/s2.50s
100%
1M$0.1859$0.3718—28tok/s3.04s
100%
1M$0.0504$0.2187$0.008190tok/sfrom traffic, 7d34.58sfrom traffic, 7d
96.9%
1M$6.75$33.75$0.67522tok/s1.76s
1M$1.013$5.063$0.1012526tok/s2.54s
88.0%
1M$0.405$3.375$0.040545tok/s1.38s
99.3%
1M$1.542$5.243—178tok/s1.95s
1M$3.768$18.837$0.40552tok/sfrom traffic, 7d2.79sfrom traffic, 7d
98.9%
1M$2.70
$5.40
$13.50
$27.00
$0.337555tok/s3.65s
1M$3.151
$5.40
$15.753
$27.00
$0.39387536tok/sfrom traffic, 7d4.87sfrom traffic, 7d
96.7%
1M$2.70$16.20$0.2767tok/s3.07s
1M$2.70$16.20$0.2740tok/s1.20s
1M$0.3127$1.877$0.02766tok/sfrom traffic, 7d3.73sfrom traffic, 7d
98.3%
1M$0.1548$0.9288$0.02775tok/s3.03s
85.7%
500K$2.73$8.189—48tok/s2.39s
100%
262K$0.108$0.4455—67tok/s4.67s
100%
1M$2.70$13.50$0.2747tok/sfrom traffic, 7d7.77sfrom traffic, 7d
93.3%
1M$1.101$3.76$0.29835fp4 · fp873tok/s2.67s
100%
262K$1.009$4.774—31tok/s3.12s
1M$13.50$67.50—15tok/s3.19s
100%
1M$0.675$3.375$0.135fp4 · fp834tok/s3.94s
1M$0.4431$1.773$0.086480tok/s14.79s
100%
1M$0.405$1.62$0.0756fp4 · fp838tok/s1.99s
99.6%
256K$0.27$1.552$0.054fp862tok/s2.92s
100%
1M$1.863$5.586$0.1687574tok/s11.66s
100%
256K$1.389$2.777$0.2789tok/s3.18s
1M$1.928$11.563$0.2025116tok/s2.51s
97.1%
1M$1.92$3.838$0.2762tok/s3.03s
1M$0.1264$0.2526$0.024326tok/s1.64s
100%
—$0.072 / image35s/img35.29s
97.9%
262K$0.7856$3.667—29tok/s1.57s
100%
1M$6.75$33.75$0.67527tok/s1.65s
203K$1.89$5.94$0.27675fp4 · fp828tok/s13.89s
1M$0.454$2.724—53tok/s10.94s
99.9%
128K$0.0763$0.218$0.135fp4 · fp88tok/s6.85s
99.7%
2M$1.767$3.534—414tok/s5.25s
2M$1.763$3.528—31tok/s1.62s
205K$0.405$1.62—50tok/s5.68s
100%
—$0.061 / image11s/img10.88s
1M$2.70$16.20$0.2796tok/s4.64s
—$0.108 / image
1M$4.05$20.25$0.40530tok/s2.04s
100%
—$0.054 / image13s/img13.11s
—$0.3375 / image
197K$0.405$1.62—25tok/s3.09s
97.7%
—$0.1512 / sec
203K$0.081$0.54$0.0135fp847tok/s5.60s
100%
1M$0.675$4.05$0.067536tok/s1.65s
99.4%
160K$0.518$0.7571—fp413tok/s1.19s
97.9%
—$0.15525 / sec
—$0.0405 / image12s/img12.24s
—$0.135 / sec
200K$1.35$6.75$0.13534tok/s1.78s
99.3%
2K$0.0027$0.00—1096tok/s0.37s
100%
128K$0.0494$0.2406—27tok/s4.65s
97.7%
131K$0.1945$1.272$0.03375fp847tok/s3.21s
100%
131K$0.3586$1.629$0.135fp4 · fp847tok/s1.12s
99.9%
33K$0.0135$0.00—299tok/s1.37s
100%
33K$0.0135$0.00—1228tok/s0.33s
33K$0.027$0.00—628tok/s0.65s
41K$0.135$0.00—
—$0.054 / image
33K$0.108$0.378—26tok/s24.27s
71.8%
1M$0.27$1.08—fp834tok/s1.07s
100%
—$0.07 / 1K char39ch/s3.23s
—$0.04 / 1K char40ch/s3.20s
—$0.081 / image
200K$4.05$20.25—23tok/s1.82s
127K$1.35$1.35—29tok/s1.54s
100%
—$0.005 / image31s/img31.35s
97.8%
64K$0.7425$2.957—54tok/s1.60s
100%
128K$0.135$0.432—14tok/s1.15s
88.5%
—$0.2025 / 1K char
—$0.0009 / min7.7×rt1.31s
99.3%
131K$1.35$1.35—fp817tok/s0.89s
100%
131K$0.945$0.945—fp822tok/s0.96s
100%
—$0.2025 / 1K char
—$0.00405 / min0.37sfrom traffic, 7d
100%
1K$0.0135$0.00—592tok/s0.87s
8K$0.0135$0.00—1311tok/s0.40s
100%
1K$0.0135$0.00—1067tok/s0.38s
1K$0.00675$0.00—1393tok/s0.29s
—$0.405 / 1K char
100%
1K$0.00675$0.00—475tok/s0.85s
1K$0.00675$0.00—92tok/s1.40s
1K$0.00675$0.00—265tok/s0.97s

Token models priced per 1M tokens. The cached column is what a repeated prompt prefix costs on that model — a rate of its own, read from the same source as the input rate beside it, not a fixed fraction of it. It runs from a tenth of the input rate to the full input rate depending on the model, and a dash means nobody publishes one, which is a different statement from “free”. Image, video, and audio models priced per unit (image, second, video, minute, song, call, or 1K characters). Throughput and latency are a probe with an identical input when Kyma has that sample. If a row is labeled “from traffic, 7d”, that number is production traffic instead — real prompts, not comparable across models. Image throughput is seconds per succeeded job (30-day median). Video, music and realtime stay empty until a job exists. An empty throughput, latency or uptime cell means Kyma does not yet have a number for that cell. Uptime is a percentage only after 20 observations, over 30 days, probes and real requests together, including Kyma's own accounts. Rankings demand is the same total. Precision names the numeric formats a model may be served at that are BELOW the one its creator released it in, and it is filled on the 14 models where that is true. A low format is not a downgrade when it is how the model shipped, so this is a comparison against the creator’s own model card rather than a list of small numbers. An empty precision cell means there is nothing to flag: either no route reports going under the release, or nobody could source what the release was. Those two are different answers and the model’s own page prints which one it is.

Start building with Kyma API

$0.50 signup credit for eligible free-tier models after email verification.

Create account