Hosted · Fleet
What we serve ourselves
Open weights on llmcloud accelerators. Served precision is published per model, prices are per 1M tokens, and every row is reachable through the same gateway API.
10 models
| Model ID | Params | Precision served | Context | $/1M in | $/1M out | tok/s | Regions | Status |
|---|---|---|---|---|---|---|---|---|
llmcloud/deepseek-v3.2-exp Served at the weights the lab published — no community quant, no silent swap. | 685B MoE | FP8 (native) | 164k | $0.22 | $0.85 | 240 | us-east, us-west, eu-west | Live |
llmcloud/llama-4-maverick Long-context workhorse; 1M window served without a context downgrade at peak. | 400B MoE | BF16 | 1000k | $0.28 | $0.90 | 210 | us-east, us-west, eu-west | Live |
llmcloud/qwen3-235b-a22b Strong multilingual default; the cheap alternate on several workload rows. | 235B MoE | BF16 | 256k | $0.20 | $0.70 | 265 | us-west, eu-west | Live |
llmcloud/qwen3-coder-480b Cheap alternate for the coding-agent workload at a fraction of frontier price. | 480B MoE | FP8 | 256k | $0.35 | $1.20 | 190 | us-west, eu-west | Live |
llmcloud/gpt-oss-120b High-throughput default for classification, extraction and batch jobs. | 120B MoE | MXFP4 | 128k | $0.10 | $0.40 | 320 | us-east, eu-west | Live |
llmcloud/mistral-large-3 EU-only capacity while we finish burn-in on the Frankfurt cluster. | 123B dense | BF16 | 256k | $0.50 | $1.50 | 140 | eu-west | Beta |
llmcloud/kimi-k2-0905 Agentic tool-use candidate; conformance probes still running. | 1T MoE | FP8 | 256k | $0.45 | $1.80 | 120 | us-west | Beta |
llmcloud/qwen3-vl-235b Vision and document parsing on our own GPUs. Waitlist open. | 235B MoE | BF16 | 128k | $0.40 | $1.40 | 110 | us-west | Coming soon |
llmcloud/whisper-v4 Speech-to-text on hosted capacity, billed per audio-minute. Pricing TBA. | 1.6B | FP16 | — | TBA | TBA | — | us-east, eu-west | Coming soon |
llmcloud/embed-3 Retrieval embeddings served next to the models that consume them. | 7B | BF16 | 32k | $0.02 | TBA | — | us-east, us-west, eu-west | Coming soon |
These models also appear in the main catalog alongside every other provider serving the same weights — that is where you compare our price and latency against theirs. Nothing here is ranked differently because we host it.