4 models · 4,981 observations · under measurement since 3 Sept
Claude Fable 5.1 vs DeepSeek V4.1 Flash vs Gemini 3.8 Flash vs GPT-6 Astra
Kyma serves all of these, so the availability, throughput and latency below are its own readings rather than anybody's marketing. What each creator publishes about its own model is further down, kept separate, and clearly theirs.
Anthropic
22 obs
since 3 Sept
15 from Kyma's own probe
DeepSeek V4.1 FlashDeepSeek
4,532 obs
since 12 Sept
210 from Kyma's own probe
353 obs
since 3 Sept
9 from Kyma's own probe
GPT-6 AstraOpenAI
74 obs
since 12 Sept
21 from Kyma's own probe
Availability, and how much of it we can actually see
Successful observations divided by total observations over the last 30 days, counting Kyma's own probe and real customer requests the same way. Same method and same figures as /models.
No winner here. Claude Fable 5.1, DeepSeek V4.1 Flash, Gemini 3.8 Flash sit inside the band, so this measurement cannot separate them. The honest claim is that they are all up between 40.5 and 100.0%.
DeepSeek V4.1 Flash
GPT-6 AstraThe tinted band is the measurement's own grain. It is 13.6 points wide, which is what three failed requests are worth in the thinnest sample on this page. One or two is what luck looks like, so a lead has to clear all three before this page will call it. Scale starts at 16%, not at zero, and dot size is the size of the sample behind the figure.
Throughput
Median tokens per second, from one identical prompt sent to every model every six hours, so the only variable left between two rows is the model.
DeepSeek V4.1 Flash
GPT-6 AstraLatency
Median time to a complete response on that same fixed prompt. It moves with output length, so it is a reading on one prompt rather than a promise about yours.
DeepSeek V4.1 Flash
GPT-6 AstraAt least one model here has been measured for days rather than months. Its figures are real readings, not estimates, but they rest on a sample thin enough that a single failure would move them by a full point or more, which is why this page will not rank on them.
What a request costs today, per 1M tokens. A model on offer shows the rate you are actually charged and the date it returns to list.
| Fact | DeepSeek V4.1 Flash | GPT-6 Astra | ||
|---|---|---|---|---|
Input per 1M tokens | $13.50 | $0.1048 | $1.013 | $13.50 |
Output per 1M tokens | $67.50 | $0.3141 | $5.063 | $67.50 |
Cached input repeated prefixes | not supported | not supported | not supported | $1.35 |
The specification each creator publishes for its own model, read straight from the catalogue Kyma serves from. Nothing in this block is a Kyma opinion.
| Fact | DeepSeek V4.1 Flash | GPT-6 Astra | ||
|---|---|---|---|---|
Context window tokens in one request | 1M | 1.05M | 1.05M | 1.05M |
Max output ceiling on one response | 128K | 131K | 66K | 128K |
Accepts | Text, image, file | Text, image | Text, image, video, file, audio | Text, image |
Tool calling | Yes | Yes | Yes | Yes |
Reasoning | Yes | Yes | Yes | Yes |
Vision input | Yes | Yes | Yes | Yes |
Weights | Closed | Opendeepseek-ai/DeepSeek-V4.1-Flash | Closed | Closed |
Below release precision narrower than the creator shipped | Not established | fp4 | Not established | Not established |
What this instrument does not measure
This page carries no benchmark scores, so it will not tell you which model is smarter. It tells you which is cheaper, which is faster, which stays up, and what each one accepts. For the quality question, run both on your own prompts: one key, one line changed, and the answer is about your work rather than someone else's test set.