Alibaba

Alibaba

Qwen3 Embedding 8B

Qwen3 Embedding 8B generates 4096-dimensional text embeddings optimized for multilingual retrieval and long-context RAG pipelines. Reach for it when document length or cross-language recall outweighs the need for the lowest possible per-token rate.

Modalities

Text → Text

Usage

How much this model is actually called here.

Rank

#66

of 101 active models

Tokens served

150.3K

all-time

Platform share

0.0%

of all tokens

Tokens · last 14 daysSep 26 → Oct 9

Pricing

Pay per token. Cached input bills at this model’s own cached rate, listed below.

$0.0135 /1M input$0.00 /1M output
Hobby10 req/day · 2K in / 500 out
~$0.01/mo
Production1,000 req/day · 2K in / 500 out
~$0.81/mo
Scale20,000 req/day · 2K in / 500 out
~$16.20/mo
+ Estimate your workload
1,000
2,000
500

Estimated monthly cost

$0.8100

$0.0270 / day on Qwen3 Embedding 8B

Same workload on:

Qwen 3.8 27B$59.77+7279%
Qwen 3.7 Flash$6.30+678%

Estimates use list pricing. Actual bills depend on real token counts, and every response includes its exact cost.

When to use Qwen3 Embedding 8B

Updated 2026-07-31

Where this model earns its cost — and where it doesn't.

This model produces dense vector representations with a 32,768-token context window. It is designed for text-only embedding tasks and does not support reasoning, vision, or structured output generation.

On Kyma, it runs as an OpenAI-compatible endpoint behind a single API key. Requests benefit from automatic failover if a serving path degrades, and responses return the exact cost in usage.cost alongside an X-Kyma-Model header. Prompt caching does not apply to embeddings: every input token bills at the input rate.

Because it is strictly an embedding model, it returns zero output tokens and cannot generate conversational text. It operates at a medium speed tier, making it better suited for batch indexing or retrieval-heavy workflows than for real-time, low-latency interactive search.

Multilingual Document Search

Finds relevant passages across different languages without translation overhead.

Long-Form Context Indexing

Embeds full documents up to 32K tokens to avoid aggressive chunking.

High-Recall Retrieval Pipelines

Prioritizes semantic match quality over minimal compute cost for RAG systems.

Not ideal for: Do not use this model for real-time chat, text generation, or tasks requiring sub-100ms latency, as it only outputs vectors and runs at a medium speed tier.

How it compares

Against the peers people actually weigh it against.

SpecQwen3 Embedding 8BQwen 3.8 27BQwen 3.7 Flash
Input /1M$0.0135$0.3576$0.0504
Output /1M$0.00$2.554$0.2187
Context33K1M1M
ToolsNoYesYes
ReasoningNoYesYes
Throughput299 tok/s36 tok/s68 tok/s

Quick start

Up and running in under two minutes.

  1. 1

    Create an API key

    Sign up and grab a key from the dashboard — $0.50 free credit on the free tier, which covers Qwen3 Embedding 8B. No card required.

    Get API key →
  2. 2

    Make your first request

    Submit a generation job and poll until it succeeds.

    curl https://kymaapi.com/v1/embeddings \
      -H "Authorization: Bearer YOUR_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "model": "qwen3-embedding-8b",
        "input": ["first document", "second document"]
      }'

FAQ

Common questions about this model.

What is the context window of Qwen3 Embedding 8B?

Qwen3 Embedding 8B has a 33K-token context window — roughly 48 pages of text in a single request.

How much does the Qwen3 Embedding 8B API cost?

$0.0135 per 1M input tokens and $0.00 per 1M output tokens. No subscription; you pay only for what you use.

Does Qwen3 Embedding 8B support function calling?

No — Qwen3 Embedding 8B does not support tool calling. For agent workloads, choose a tools-enabled model from the catalog.

How do I use Qwen3 Embedding 8B?

Kyma is OpenAI-compatible: point your SDK's base URL at https://kymaapi.com/v1, use your Kyma API key, and set the model to qwen3-embedding-8b. Signing up is free and includes $0.50 of credit on the free tier, which covers this model — no card required.

Does this model generate text responses?

No, it is strictly an embedding model that outputs 4096-dimensional vectors and returns zero output tokens.

How does Kyma handle prompt caching for this endpoint?

Prompt caching does not apply to embedding requests: there is no cached rate, and every input token bills at the input rate.

Can I use the same API key for this model as I do for others?

Yes, Kyma is OpenAI-compatible and uses a single API key across all models at the standard base URL.

Start with $0.50 free credit on the free tier — no card required.Create account →

More models by Alibaba

See all 14 →
ModelContextInputOutput
AlibabaQwen 3.8 Flash1M$0.2045$0.641
AlibabaQwen 3.8 27B1M$0.3576$2.554
AlibabaQwen 3.8 Max1M$2.228$6.684
AlibabaQwen 3.7 Flash1M$0.0504$0.2187
AlibabaQwen 3.7 Plus1M$0.4431$1.773
AlibabaQwen 3.7 Max1M$1.863$5.586
AlibabaQwen 3.6 Plus1M$0.454$2.724
AlibabaQwen 3 Coder131K$0.3586$1.629