Changelog

What changed in Kyma API, newest first. Each change with its detail is in the full release notes.

October 9, 2026

  • Generate images with GPT Image 2.5 Flare, which keeps a character consistent across reference images
  • Generate cinematic video clips with Kling 3 Pro, from a prompt or from a first-frame image
  • Turn one portrait and one voice or song track into a lip-synced talking or singing clip
  • Generate a full song with vocals from a prompt and your own lyrics, in Vietnamese or any other language, with Lyria 3.5
  • Generate video again, including clips with native speech and sound, with Veo 3.1 Fast
  • A failed image, video or audio job on the Logs page now shows the same error code as GET on the job

October 8, 2026

  • Three media model ids that can no longer be served are retired, so no page, alias or example offers a model that cannot answer

October 7, 2026

  • v1/messages honours the Anthropic thinking parameter and returns the model's reasoning as thinking content blocks, streamed and not

October 6, 2026

  • Your unverified accounts and API keys are now strictly protected from unauthorized changes.
  • You will no longer accidentally accumulate debt when running embeddings or reranking with an empty balance.

October 4, 2026

  • 12 prices fall and 7 rise, every one computed from what the request actually costs to serve
  • The default API key can now be disabled and deleted like any other, and the key list comes back newest first instead of in whatever order Postgres felt like
  • A kyma-gen run that takes longer than the sync wait now keeps going to its result instead of stopping with a timeout
  • A refusal is never answered by another creator's model under the name you asked for
  • When Kyma falls back to another model, the bill follows that model's price and never exceeds the price of the model you asked for

September 29, 2026

  • Creating, disabling and deleting API keys now requires a dashboard session, so a leaked API key cannot mint new keys that outlive it

September 28, 2026

  • Audio understanding is billed from the audio itself, so the price no longer depends on a duration you type

September 24, 2026

  • Every API error now carries a stable string code, a category, whether to retry, a link to its docs entry and a request id you can quote to support, and /v1/messages answers every error in Anthropic's own format
  • A stream that fails after it has started now ends with a clear, retryable error instead of passing through the upstream's own message, and /v1/messages clients see a proper error event

September 22, 2026

  • Billing settings and admin now require a dashboard session, so an API key used for inference cannot change auto top-up, saved cards or billing details
  • Tier rows in GET /v1/auth/limits report audio concurrency per capability, the same shape GET /v1/limits/tiers already uses, and four error answers now tell you what to do instead of carrying internal routing details
  • Request logs show which key made each call by its name and last six characters; logs no longer return full key strings
  • cancelling tells you plainly when a job can no longer be stopped, instead of reporting it cancelled
  • Sending an embedding or reranking model to /v1/messages returns a clear 400 naming the right endpoint instead of a chat answer
  • GET /v1/auth/limits now shows how many requests and tokens you have used this minute, ?tools=true on GET /v1/models lists only models that can call tools, and chat responses carry usage.cached_tokens on every call

September 21, 2026

  • Grok 4.7, xAI's newest flagship for coding, agentic tasks and knowledge work, is callable as grok-4.7 with a 500K context window and image and file input
  • An image or video job now shows the same cost on every endpoint that returns it, and says how much of the hold came back
  • The catalogue, pricing, docs and llms.txt stop offering two models Google is shutting down, so nobody starts new work on them, while existing calls keep working until each retirement date

September 20, 2026

  • When every provider for a model is failing, you get 503 with Retry-After instead of a 500 with no guidance, and the failure shows up in your request log
  • Every kyma-gen generation now tells you what it held, what it charged and what it gave back, including when you cancel it
  • POST /v1/messages now tells you what a non-streamed call cost, and browser code can finally read Kyma's response headers
  • A request the fallback cannot serve now gets a clear 400 telling you to call the primary endpoint, instead of an unexplained 500
  • Kyma stops advertising video models it no longer serves, the description you read now lists exactly what the catalogue can do
  • A public Muse share link no longer carries the account id of whoever shared it, and neither do the URLs of files you upload
  • A streamed turn that answers with a tool call or with reasoning is now charged and logged like any other, instead of silently costing nothing and leaving no record
  • The welcome email, the repo's front page and the fallback backend's own description stop naming a kind of model nobody can buy

September 19, 2026

  • A model that answers a question with a choice and a confidence instead of prose, at POST /v1/decisions
  • Muse Spark 1.3, Meta's release after 1.2, is callable as muse-spark-1.3 with a 1M context window and multimodal input
  • Transcription checks your balance against the file you sent, before the audio leaves Kyma, and a transcription Kyma has already paid for is always charged instead of being refused afterwards

September 13, 2026

  • A text document sent to /v1/messages now reaches the model, so the answer is about the document you sent instead of one the model never saw
  • Images sent to /v1/messages now reach vision models, so the answer describes the picture you sent instead of one the model never saw
  • A streamed /v1/messages response now ends with exactly one message_delta and one message_stop, and that message_delta reports input_tokens and output_tokens, so clients that track usage from the stream see the real numbers
  • tool_choice none and disable_parallel_tool_use now reach the model, so a request that says no tool calls, or one tool call at a time, gets that
  • during a promotion both endpoints now quote what the model costs today, with the list price beside it, instead of the list price alone
  • A request the model rejects as invalid now returns that status (400, 413 or 422) instead of 503, so SDKs stop retrying a request that cannot succeed and the error says to change the request
  • turning parallel tool calls off now means at most one tool call per turn on every model, including models that ignored the setting
  • an oversized request is refused at once with a 413 that names its size, before Kyma reads the body, on every chat endpoint

September 12, 2026

  • Two new models, DeepSeek V4.1 Flash, the lab's successor to both V4 Flash and V4 Pro with a 1M context window and image input at the Flash price band, and GPT-6 Astra, OpenAI's GPT-6 flagship with a 1.05M context window. deepseek-v4-pro retires on 2026-09-14 in favour of deepseek-v4.1-flash, and from now on a model with an announced retirement date returns the retired-model error naming its replacement from that date
  • The catalogue now lists only the media models Kyma actually carries, so a model you can pick is a model that works; paused SKUs say so up front instead of failing at the provider

September 11, 2026

  • Chat completions on the Bun runtime answer again instead of 500ing on every request
  • Two model ids that were being answered by a different model are retired with a named replacement, one id is pinned to the build its route actually serves, and two retirement dates are corrected
  • An upload no longer fails on a generic Content-Type or a WebM/OGG/FLAC file; a segment format on a model without timestamps is refused up front instead of returning an empty file that was billed; a silent file is a 422, not a 502; every transcription error names a model that can serve the request
  • A transcription call that omits response_format works again on every model; the default follows the model instead of assuming segment timestamps

September 8, 2026

  • Saving or updating a chat over 512 KB is refused up front with 413 instead of being stored, and an error while sharing or cancelling a job no longer carries internal database text

September 7, 2026

  • A comparison you build yourself now becomes a real, listed page. Pick any two to four models on a /compare page and the combination shows up on /compare under "Recently compared", and once it has been opened on three separate days and every model in it has enough measurement behind it, it goes into the sitemap and can be found in search. Three and four model comparisons are no longer excluded from search on principle. The conclusion sentence on each page ("X is ahead at 99.4%") now sits above the chart instead of under it, so the answer is the first thing read.
  • Every model on GET /v1/models now says whether an account that has never added credits can call it, so you can tell before you send a request instead of finding out from a 402
  • Signing in now also sets a session cookie, so the dashboard can render your sidebar and email on the server instead of waiting for JavaScript to read localStorage and ask who you are
  • Your per-model usage breakdown answers in a fraction of a second instead of seconds when you do not name a time window, because it now defaults to the last 30 days rather than scanning your whole account history
  • Every place that mentions the $0.50 signup credit now says which models it buys, the free tier, instead of implying it covers the whole catalogue, so you find out before a request that a video, image or frontier model needs a top-up
  • Every page now shows the price you are actually charged, to the digit
  • The rankings Performance board stops going blank when one leaderboard query fails
  • The rankings page stops showing a blank Performance board for five minutes after the API has recovered
  • The rankings board ranks on reliability, prices every model in the unit it is sold by, and reports the same uptime as the rest of the site

September 6, 2026

  • The dashboard shell loads a 190-byte identity instead of a 14 KB key dump on every page, per-key spend lives on the keys endpoint, and the Usage and Logs pages answer in well under a second
  • The MCP sign-in screen now has Continue with Google, so accounts created with Google can connect Claude, Cursor, Codex and ChatGPT without a password

September 5, 2026

  • qwen-3.6-plus gets cheaper, the published rate follows what it actually costs to serve
  • The Usage page loads in well under a second, and 7d / 30d / 90d show the range they say

September 4, 2026

  • When the measurement query behind model uptime is unavailable, the endpoint now says so instead of reporting every model as unmeasured, and any unexpected failure at the edge comes back as readable JSON a browser can actually display
  • Saved chats, one-click unsubscribe, sharing a generated image or video, and cancelling a kyma-gen job are now answered at the Cloudflare edge instead of being relayed to the origin
  • Rankings, model uptime, cost per session, model recommendations, public stats and badges, the media registry, generation pricing and public Muse shares are now answered at the Cloudflare edge instead of being relayed to the origin

September 3, 2026

  • A request for a retired Imagen 4 model now names its replacement instead of failing upstream with nothing to act on
  • Gemini 3.8 Flash and Claude Fable 5.1 are available by name on the chat endpoints
  • Two open-weight multimodal models are available by name, Qwen 3.8 27B and DeepSeek V4 Flash Vision (experimental)
  • Promotional prices on five models now show as promotions with an end date, and GPT-5.6 Sol is listed at the price you were already being billed
  • qwen-3.7-max is listed at the price of the route that now serves it, so the page and the bill agree again
  • The 43 image, video, speech, music and transcription model pages now show the price in the unit the model bills in, and only the features it has

August 14, 2026

  • Gemini 3.1 Pro is live at $2.70/$16.20 per million tokens, the first Gemini Pro tier on Kyma, and Nano Banana 3 Flash is repaired after its upstream preview endpoint stopped answering
  • Gemini 3.7 Flash is live at $1.01/$5.06 per million tokens, and half that, $0.51/$2.53, while the launch promotion runs
  • 7 prices fall and 14 rise, every one computed from what the request actually costs to serve
  • v1/models now quotes the price you actually pay today, so a model on promotion no longer reads as twice its real cost

August 13, 2026

  • Claude Fable 5 is live at $13.50 in and $67.50 out per million tokens
  • Grok 4.6 is live at $2.70 in and $8.10 out per million tokens
  • Qwen 3.8 Max is live at $2.2275 in and $6.684 out per million tokens
  • Qwen 3.8 Max is listed once, at the $2.2275 / $6.684 rate

August 3, 2026

  • You are now billed at the exact rates published by the suppliers, eliminating incorrect pricing on several popular models.
  • The model uptime dashboard now accurately reflects supplier availability instead of showing false outages caused by our own probe limits.

August 2, 2026

  • 1 model got cheaper to call, effective immediately

August 1, 2026

  • You can list everything one lab makes in a single request
  • One click unsubscribes you from Kyma email, and it sticks
  • You can add live web results to 25 models by appending one suffix to the model id
  • Grok, Meta's newest line and Claude are all callable with the same key as everything else
  • Published prices follow the cheapest route that can serve you, not whichever one happened to carry the last request
  • What a model says it can do now matches what its provider says it can do
  • Searching the models page no longer throws you somewhere else on the page

July 31, 2026

  • You can compare models on measured speed and uptime instead of taking a claim on trust
  • When a supplier cuts its price, yours falls too, automatically, and within a day
  • Your account is harder to attack, and a database read no longer yields anything replayable
  • The cost in your streaming response is what you were charged, not what your request cost Kyma

July 30, 2026

  • Older clients that speak the legacy completions shape work without changing your code
  • You can build retrieval on Kyma without a second vendor for the embedding half
  • Every model has its own page, with the whole price and a plain-markdown mirror an agent can read
  • You are billed for the route that actually served you, and never above the published price
  • Pages load faster, and the site stops rendering dark on a light theme

July 29, 2026

  • You have until 20 October to move off nano-banana, and Kyma stopped recommending it today
  • Six models joined the catalogue, including the cheapest one Kyma sells
  • Five models were repriced against what they actually cost to serve, two of them had been sold below that
  • Multi-turn reasoning with Gemini 3 keeps its train of thought across turns

July 28, 2026

  • A long conversation is never rescued by a model too small to hold it

July 27, 2026

  • You can ask the catalogue for exactly the models you need instead of reading all of them
  • An error now tells you whose problem it is, so you know whether retrying will help

June 2026

  • Every model got a page with live numbers instead of a specification copied from a launch post
  • Six models joined, including two flagship coding models and the strongest open-weight tier Kyma had carried
  • Speech and transcription stopped failing outright when one supplier did

June 19, 2026

  • 2 new flagship models: GLM 5.2 and Kimi K2.7 Code

June 10, 2026

  • 4 new models: Qwen 3.7 Plus, MiniMax M3, Nemotron 3 Ultra, Step 3.7 Flash

June 4, 2026

  • ElevenLabs v3, most expressive TTS + low-latency streaming

May 2026

  • Speech, music and sound effects arrived, so an audio app no longer needs a second vendor and a second bill
  • Image and video generation grew a real catalogue, and pricing that matches how each one actually charges
  • Web-grounded answers became callable models rather than a separate search product to integrate
  • Signing up got harder to abuse and easier for real users on shared connections

May 17, 2026 (later)

  • Google media models, 7 new SKUs + public pricing catalog

May 17, 2026

  • Audio infrastructure refresh

May 1, 2026

  • Several updates, listed in the full release notes

April 30, 2026

  • MiniMax bundle, 9 new SKUs across audio, image, video
  • Sharing a cloned voice_id with another account is rejected
  • Image catalog refresh, 5 new SKUs

April 29, 2026

  • Audio - 2 new endpoints + 2 SKUs

April 26, 2026

  • Video Generation - 5 new models

April 25, 2026

  • DeepSeek V4, Pro and Flash
  • Image Generation, Week 1

April 23, 2026

  • API Reliability and Platform Changes

April 21, 2026

  • Product and Dashboard Updates

April 19, 2026

  • Product and Dashboard Updates

April 17, 2026

  • Agent, Install, and Runtime Improvements

April 16, 2026

  • Kyma Agent v0.1.12, KYMA.md context + MCP servers
  • GLM family from Z.AI

April 15, 2026

  • Kyma Agent v0.1.8

April 11, 2026

  • Kyma CLI v0.3

April 10, 2026

  • Higher Limits, Better Emails

April 8, 2026

  • Models & Pricing

April 7, 2026

  • Reliability & Performance

April 4, 2026

  • Launch