Skip to navigationSkip to loginSkip to main contentSkip to footer section

Model-as-a-service Pricing

Serve Generative AI models and pay for a dedicated infrastructure or for thousands of tokens

Generative APIs - Serverless

Use the latest AI models via API, pay by thousand tokens.
Try out new models with our free tier: 1 million tokens and 60 minutes of audio transcription.
All requests performed using Batches API are priced with a -50% discount.

Generative API
glm-5.2
Chat and code
€1.80 /million tokens
€5.50 /million tokens
deepseek-v4-flash-0731
Chat and code
€0.40 /million tokens
€0.08 /million tokens cached
€0.80 /million tokens
qwen3.8-27b
Chat, code, and Vision
€0.60 /million tokens
€0.12 /million tokens cached
€3.30 /million tokens
gemma-4-26b-a4b-it
Chat, code, and Vision
€0.25 /million tokens
€0.50 /million tokens
mistral-medium-3.5-128b
Chat, code, and Vision
€1.50 /million tokens
€7.50 /million tokens
whisper-large-v3
Audio transcription
€0.003 /Audio minute
Free
llama-3.3-70b-instruct
Chat
€0.90 /million tokens
€0.90 /million tokens
qwen3-235b-a22b-instruct-2507
Chat
€0.75 /million tokens
€2.25 /million tokens
qwen3-coder-30b-a3b-instruct
Chat and code
€0.20 /million tokens
€0.80 /million tokens
qwen3-embedding-8b
Embeddings
€0.10 /million tokens
Free
pixtral-12b-2409
Chat and Vision
€0.20 /million tokens
€0.20 /million tokens
mistral-small-3.2-24b-instruct-2506
Chat and Vision
€0.15 /million tokens
€0.35 /million tokens
gpt-oss-120b
Chat
€0.15 /million tokens
€0.60 /million tokens
bge-multilingual-gemma2
Embeddings
€0.10 /million tokens
Free
qwen3.6-35b-a3b
Chat, code, and Vision
€0.25 /million tokens
€1.50 /million tokens
qwen3.5-397b-a17b
Chat, code, and Vision
€0.60 /million tokens
€3.60 /million tokens
InformationOutlineIconLegal noticeArrowDownIcon

Prices before tax.
You benefit from a free tier on the first 1,000,000 tokens. You'll be charged from token number 1,000,001.

Generative APIs - Dedicated Deployment

Deploy your managed AI infrastructure with dedicated GPUs and optimized models. You are charged for usage of the GPU type you choose. Billing only starts once the model is deployed

Managed Inference (sorted by column hourlyPrice ascending)
L4-1-24G
€0.93
€678.90
L40S-1-48G
€1.72
€1,255.31
H100-1-80G
€3.40
€2,481.71
H100-2-80G
€6.68
€4,876.25
H100-SXM-2-80G
€7.95
€5,803.50
H100-SXM-4-80G
€15.22
€11,110.57
H100-SXM-8-80G
€30.06
€21,943.80
InformationOutlineIconLegal noticeArrowDownIcon

Prices before tax