4 models · 476 observations · under measurement since 2 May

Claude Fable 5.1 vs Claude Opus 5 vs Claude Sonnet 4.6 vs Gemini 3.8 Flash

Kyma serves all of these, so the availability, throughput and latency below are its own readings rather than anybody's marketing. What each creator publishes about its own model is further down, kept separate, and clearly theirs.

Measured by Kyma476 observationssince 2 May
AnthropicClaude Fable 5.1

Anthropic

22 obs

since 3 Sept

15 from Kyma's own probe

AnthropicClaude Opus 5

Anthropic

12 obs

since 1 Aug

15 from Kyma's own probe

AnthropicClaude Sonnet 4.6

Anthropic

89 obs

since 2 May

15 from Kyma's own probe

GoogleGemini 3.8 Flash

Google

353 obs

since 3 Sept

9 from Kyma's own probe

Availability, and how much of it we can actually see

Successful observations divided by total observations over the last 30 days, counting Kyma's own probe and real customer requests the same way. Same method and same figures as /models.

kymaapi.com

No winner here. Claude Fable 5.1, Claude Sonnet 4.6, Gemini 3.8 Flash sit inside the band, so this measurement cannot separate them. The honest claim is that they are all up between 99.2 and 100.0%.

AnthropicClaude Fable 5.1
AnthropicClaude Opus 5
AnthropicClaude Sonnet 4.6
GoogleGemini 3.8 Flash
100.0%
not measured
100.0%
99.2%
80%86.4%100%

The tinted band is the measurement's own grain. It is 13.6 points wide, which is what three failed requests are worth in the thinnest sample on this page. One or two is what luck looks like, so a lead has to clear all three before this page will call it. Scale starts at 80%, not at zero, and dot size is the size of the sample behind the figure.

Anything inside the band is level with the leader as far as this measurement can tell. Anything left of it is genuinely behind.

Throughput

Median tokens per second, from one identical prompt sent to every model every six hours, so the only variable left between two rows is the model.

AnthropicClaude Fable 5.1
14 tok/s
AnthropicClaude Opus 5
22 tok/s
AnthropicClaude Sonnet 4.6
29 tok/s
GoogleGemini 3.8 Flash
31 tok/s
Longer is better.

Latency

Median time to a complete response on that same fixed prompt. It moves with output length, so it is a reading on one prompt rather than a promise about yours.

AnthropicClaude Fable 5.1
3.06 s
AnthropicClaude Opus 5
1.76 s
AnthropicClaude Sonnet 4.6
2.05 s
GoogleGemini 3.8 Flash
2.19 s
Shorter is better.

At least one model here has been measured for days rather than months. Its figures are real readings, not estimates, but they rest on a sample thin enough that a single failure would move them by a full point or more, which is why this page will not rank on them.

Price on Kyma

What a request costs today, per 1M tokens. A model on offer shows the rate you are actually charged and the date it returns to list.

FactAnthropicClaude Fable 5.1AnthropicClaude Opus 5AnthropicClaude Sonnet 4.6GoogleGemini 3.8 Flash

Input

per 1M tokens

$13.50$6.75$4.05$1.013

Output

per 1M tokens

$67.50$33.75$20.25$5.063

Cached input

repeated prefixes

not supported$0.675$0.405not supported
Published by the creator

The specification each creator publishes for its own model, read straight from the catalogue Kyma serves from. Nothing in this block is a Kyma opinion.

FactAnthropicClaude Fable 5.1AnthropicClaude Opus 5AnthropicClaude Sonnet 4.6GoogleGemini 3.8 Flash

Context window

tokens in one request

1M1M1M1.05M

Max output

ceiling on one response

128K128K128K66K

Accepts

Text, image, fileText, image, fileText, imageText, image, video, file, audio

Tool calling

YesYesYesYes

Reasoning

YesYesYesYes

Vision input

YesYesYesYes

Weights

Closed
Closed
Closed
Closed

Below release precision

narrower than the creator shipped

Not establishedNot establishedNot establishedNot established

What this instrument does not measure

This page carries no benchmark scores, so it will not tell you which model is smarter. It tells you which is cheaper, which is faster, which stays up, and what each one accepts. For the quality question, run both on your own prompts: one key, one line changed, and the answer is about your work rather than someone else's test set.

Every model above runs on one key and one balance. $0.50 of free-tier credit on signup, no card.Get API key

More comparisons