Google

Google

Gemma 4 31B

#3 on KymaCheapest vision

The model the vision alias points to — Google's newest open model, with image understanding at cheap-tier pricing.

Modalities

Text+Image → Text

Below release

fp4 · fp8

Weights

Published

google/gemma-4-31B-it

Where it sits in the catalogue Kyma measures

Every number against every text model Kyma prices per token — a stated rule, not a chosen line-up.

Price · input#4/71

$0.0763

output $0.218 · cached $0.135

median $1.013 · best $0.0494

Price · output#1/70

$0.218

per 1M tokens generated

median $3.531 · best $0.218

Availability · 30d#31/54

99.71%

683 observations since 2026-05-02

median 99.9% · best 100.0%

Throughput#65/65

8 tok/s

probe median, one fixed prompt

median 35.7 tok/s · best 365.1 tok/s

Response time#60/65

6.85 s

probe median, to a complete answer

median 2.96 s · best 0.89 s

Context

128K

max output 8,192

published by Google, not measured here

Tick above each rail is this model, below it the other 70. Dashed rule is the field median, solid is its best. Better is left; the axis stops at the 90th percentile, so a few models sit past its right edge.

Usage

How much this model is actually called here.

Rank

#3

of 101 active models

Tokens served

395.5M

all-time

Platform share

9.0%

of all tokens

Tokens · last 14 daysSep 26 → Oct 9

How it behaves under real clients

The same model answering different prompt shapes — measured, not benchmarked.

ClientRequestsTokensThroughputTo first tokenp95 totalCompleted
Node.js App113878.2K3 tok/s12 s41 s87.6%

Grouped by the client that sent the request. Not a benchmark: the same model on different prompt shapes. Throughput and time-to-first-token are medians; p95 is the slowest response in twenty. 4 clients under 100 requests not shown.

Two clocks, and why they disagree

Kyma measures this model twice. Both are real; they answer different questions.

Probe · every 6h · 30 days

6.85s to answer

One fixed prompt, on a schedule, to every model. Comparable, because the model is the only thing that changes.

Observations
683
Answered by a substitute
1

Real traffic · last 7 days

7.90s to answer

Your requests, at the lengths clients actually send. Not comparable between models, but it is what running this one feels like.

Requests
157
Completed
100%
p95
21.0 s

Caching, as realised · last 7 days

1.5% of input cached

The share of input that actually hit cache, so the effective rate below is what was charged, not a best case.

Cached input
8.4K
Fresh input
569.2K
List input
$0.0763 /1M
Effective input
$0.077153 /1M

The gap is prompt length, not the model degrading. Use the probe figure to choose between models, the traffic figure to budget for your own.

Pricing

Pay per token. Cached input bills at this model’s own cached rate, listed below.

$0.0763 /1M input$0.218 /1M output
Hobby10 req/day · 2K in / 500 out
~$0.08/mo
Production1,000 req/day · 2K in / 500 out
~$7.85/mo
Scale20,000 req/day · 2K in / 500 out
~$157/mo
+ Estimate your workload
1,000
2,000
500
30%

Estimated monthly cost

$8.90

$0.2968 / day on Gemma 4 31B

Same workload on:

Gemini 3.8 Flash$137+1435%
Gemini 3.5 Flash Lite$74.92+741%

Estimates use list pricing with cached input at this model's own cached rate. Actual bills depend on real token counts, and every response includes its exact cost.

When to use Gemma 4 31B

Updated 2026-06-10

Where this model earns its cost — and where it doesn't.

Gemma 4 31B is Google's newest open model and a cheap-tier model on Kyma that accepts images, not just text. It sits in the strong quality tier and is built for multimodal and general-purpose work: send it screenshots, photos, or document scans alongside your prompt and get text back.

Coding agents drive most of its traffic here. Requests sent with the `vision` alias resolve to it. Every call gets automatic failover if a serving path degrades, and prompt caching bills repeated prompt prefixes at this model's cached input rate, which matters for agents that resend the same long system prompt.

The 128K-token context window combines with function calling and structured outputs, so it can read an image and return clean JSON in a single call — a complete loop for vision-driven pipelines.

Image understanding

Describe, classify, or answer questions about screenshots, photos, and charts — the core workload the `vision` alias exists for.

Visual data extraction

Vision input plus structured outputs means it can turn receipts, forms, or UI screenshots into validated JSON in one request.

Agent tool use

It supports function calling and already runs real agent traffic in production, from OpenClaw to Claude Code to Hermes Agent.

High-volume general tasks

Cheap-tier pricing with strong-tier quality fits summarization, classification, and chat workloads where cost per call dominates.

Long-context review

The 128K window fits large documents or long agent histories — with or without images attached.

Not ideal for: Extended reasoning problems (it has no reasoning mode) or very long single generations — output is capped at 8K tokens per request.

How it compares

Against the peers people actually weigh it against.

SpecGemma 4 31BGemini 3.8 FlashGemini 3.5 Flash Lite
Input /1M$0.0763$1.013$0.405
Output /1M$0.218$5.063$3.375
Context128K1M1M
ToolsYesYesYes
ReasoningYesYesYes
Throughput8 tok/s31 tok/s45 tok/s

Quick start

Up and running in under two minutes.

  1. 1

    Create an API key

    Sign up and grab a key from the dashboard — $0.50 free credit on the free tier, which covers Gemma 4 31B. No card required.

    Get API key →
  2. 2

    Make your first request

    Drop in your key and send a chat completion — fully OpenAI-compatible.

    curl https://kymaapi.com/v1/chat/completions \
      -H "Authorization: Bearer YOUR_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "model": "gemma-4-31b",
        "messages": [
          {"role": "user", "content": "Explain prompt caching in one paragraph."}
        ]
      }'
  3. 3

    Stream responses

    Add "stream": true to receive tokens as they arrive.

    curl https://kymaapi.com/v1/chat/completions \
      -H "Authorization: Bearer YOUR_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "model": "gemma-4-31b",
        "stream": true,
        "messages": [{"role": "user", "content": "Hello!"}]
      }'

FAQ

Common questions about this model.

What is the context window of Gemma 4 31B?

Gemma 4 31B has a 128K-token context window — roughly 188 pages of text in a single request.

How much does the Gemma 4 31B API cost?

$0.0763 per 1M input tokens and $0.218 per 1M output tokens, with cached input at $0.135/1M for repeated prompt prefixes, which this model does not discount. No subscription; you pay only for what you use.

Does Gemma 4 31B support function calling?

Yes — Gemma 4 31B supports tool/function calling and structured outputs (JSON mode), so it works with agent frameworks out of the box.

Are the weights for Gemma 4 31B publicly available?

Yes. Google publishes Gemma 4 31B's weights as google/gemma-4-31B-it, so you can download and run the model yourself (https://huggingface.co/google/gemma-4-31B-it, read 2026-08-14). Kyma serves it because it is convenient and has failover behind it, not because it is the only way to reach it.

Is Gemma 4 31B ever served below the precision its creator released it at?

Sometimes. Gemma 4 31B was released by Google at bf16 (https://huggingface.co/google/gemma-4-31b-it, read 2026-08-14), and at least one route serving it here reports fp4 or fp8 — narrower than that. Weights compressed below the release usually answer close to it, but it is not the identical artefact, and nothing in the response tells you which one answered. Kyma checked the routes on 2026-09-21, so it is your call rather than a silent one.

How do I use Gemma 4 31B?

Kyma is OpenAI-compatible: point your SDK's base URL at https://kymaapi.com/v1, use your Kyma API key, and set the model to gemma-4-31b. Signing up is free and includes $0.50 of credit on the free tier, which covers this model — no card required.

Is Gemma 4 31B good for vision tasks?

Yes — it's the model Kyma's vision alias resolves to, and image understanding is what it's recommended for. It accepts text and images in, returns text out, and pairs vision with structured outputs, so you can go from a screenshot to clean JSON in a single call.

When should I pick Gemma 4 31B over a bigger model?

When the task involves images, or when volume makes cost the deciding factor. It sits in the cheap cost tier with strong quality, which is why it carries nearly a fifth of all production tokens on Kyma. For problems that need an extended reasoning mode or outputs longer than 8K tokens, reach for a model built for that instead.

Why run Gemma 4 31B through Kyma?

One OpenAI-compatible endpoint and one API key cover this and every other model on the platform. Every request gets automatic failover, prompt caching bills repeated prefixes at this model's cached input rate, and each response reports its exact cost in usage.cost.

Start with $0.50 free credit on the free tier — no card required.Create account →

More models by Google

See all 13 →
ModelContextInputOutput
GoogleLyria 3.5—$0.108 / song
GoogleGemini 3.8 Flash1M$1.013$5.063
GoogleGemini 3.5 Transcribe—$0.00675 / min
GoogleGemini 3.7 Flash1M$1.013$5.063
GoogleGemini 3.6 Flash1M$1.013$5.063
GoogleGemini 3.5 Flash Lite1M$0.405$3.375
GoogleGemini 3.5 Flash1M$1.928$11.563
GoogleNano Banana 3 Flash—$0.061 / image