llmcloud.ai
Hosted · Fleet

What we serve ourselves

Open weights on llmcloud accelerators. Served precision is published per model, prices are per 1M tokens, and every row is reachable through the same gateway API.

10 models
Model IDParamsPrecision servedContext$/1M in$/1M outtok/sRegionsStatus
llmcloud/deepseek-v3.2-exp
Served at the weights the lab published — no community quant, no silent swap.
685B MoEFP8 (native)164k$0.22$0.85240us-east, us-west, eu-westLive
llmcloud/llama-4-maverick
Long-context workhorse; 1M window served without a context downgrade at peak.
400B MoEBF161000k$0.28$0.90210us-east, us-west, eu-westLive
llmcloud/qwen3-235b-a22b
Strong multilingual default; the cheap alternate on several workload rows.
235B MoEBF16256k$0.20$0.70265us-west, eu-westLive
llmcloud/qwen3-coder-480b
Cheap alternate for the coding-agent workload at a fraction of frontier price.
480B MoEFP8256k$0.35$1.20190us-west, eu-westLive
llmcloud/gpt-oss-120b
High-throughput default for classification, extraction and batch jobs.
120B MoEMXFP4128k$0.10$0.40320us-east, eu-westLive
llmcloud/mistral-large-3
EU-only capacity while we finish burn-in on the Frankfurt cluster.
123B denseBF16256k$0.50$1.50140eu-westBeta
llmcloud/kimi-k2-0905
Agentic tool-use candidate; conformance probes still running.
1T MoEFP8256k$0.45$1.80120us-westBeta
llmcloud/qwen3-vl-235b
Vision and document parsing on our own GPUs. Waitlist open.
235B MoEBF16128k$0.40$1.40110us-westComing soon
llmcloud/whisper-v4
Speech-to-text on hosted capacity, billed per audio-minute. Pricing TBA.
1.6BFP16TBATBAus-east, eu-westComing soon
llmcloud/embed-3
Retrieval embeddings served next to the models that consume them.
7BBF1632k$0.02TBAus-east, us-west, eu-westComing soon

These models also appear in the main catalog alongside every other provider serving the same weights — that is where you compare our price and latency against theirs. Nothing here is ranked differently because we host it.