What we serve ourselves
Open weights on llmcloud accelerators. Served precision is published per model, prices are per 1M tokens, and every row is reachable through the same gateway API.
| Model ID | Params | Precision served | Context | $/1M in | $/1M out | tok/s | Regions | Status |
|---|---|---|---|---|---|---|---|---|
llmcloud/deepseek-v3.2-exp Served at the weights the lab published — no community quant, no silent swap. | 685B MoE | FP8 (native) | 164k | $0.22 | $0.85 | 240 | us-east, us-west, eu-west | Live |
llmcloud/llama-4-maverick Long-context workhorse; 1M window served without a context downgrade at peak. | 400B MoE | BF16 | 1000k | $0.28 | $0.90 | 210 | us-east, us-west, eu-west | Live |
llmcloud/qwen3-235b-a22b Strong multilingual default; the cheap alternate on several workload rows. | 235B MoE | BF16 | 256k | $0.20 | $0.70 | 265 | us-west, eu-west | Live |
llmcloud/qwen3-coder-480b Cheap alternate for the coding-agent workload at a fraction of frontier price. | 480B MoE | FP8 | 256k | $0.35 | $1.20 | 190 | us-west, eu-west | Live |
llmcloud/gpt-oss-120b High-throughput default for classification, extraction and batch jobs. | 120B MoE | MXFP4 | 128k | $0.10 | $0.40 | 320 | us-east, eu-west | Live |
llmcloud/mistral-large-3 EU-only capacity while we finish burn-in on the Frankfurt cluster. | 123B dense | BF16 | 256k | $0.50 | $1.50 | 140 | eu-west | Beta |
llmcloud/kimi-k2-0905 Agentic tool-use candidate; conformance probes still running. | 1T MoE | FP8 | 256k | $0.45 | $1.80 | 120 | us-west | Beta |
llmcloud/qwen3-vl-235b Vision and document parsing on our own GPUs. Waitlist open. | 235B MoE | BF16 | 128k | $0.40 | $1.40 | 110 | us-west | Coming soon |
llmcloud/whisper-v4 Speech-to-text on hosted capacity, billed per audio-minute. Pricing TBA. | 1.6B | FP16 | — | TBA | TBA | — | us-east, eu-west | Coming soon |
llmcloud/embed-3 Retrieval embeddings served next to the models that consume them. | 7B | BF16 | 32k | $0.02 | TBA | — | us-east, us-west, eu-west | Coming soon |
These models also appear in the main catalog alongside every other provider serving the same weights — that is where you compare our price and latency against theirs. Nothing here is ranked differently because we host it.
The rates above are pay-as-you-go. The same models are also included unlimited on a flat monthly plan — $5 for everything under 50B parameters, $10 for everything under 200B, and $20 for the whole fleet plus free overflow when your Claude or OpenAI plan hits its daily limit. See pricing.
How to read this table
The precision actually served, not the precision the weights were released in. Where we serve a quantized build we say so and publish the eval delta against the reference build.
The window that works end to end on our serving config, including KV cache headroom under concurrency — not the architectural maximum.
Median decode tokens per second per stream at rated concurrency, measured on our own probes rather than taken from a launch blog post.
Jurisdictions where the weights are resident and inference executes. Pinned requests never leave the listed region, including cache and logs.
USD per million tokens at list. There is no platform fee on top and no credit float — you are paying for compute.
Live means in the router pool today. Coming soon means capacity is reserved and the entry exists so you can plan against it, not that it is bookable.
What a hosted token actually costs
Per-token price on open weights is not a margin decision, it is an arithmetic one: accelerator lease cost divided by realized throughput, adjusted for the output-to-input ratio of real traffic. A model that decodes at 900 tokens per second on a node costing a fixed hourly rate is roughly six times cheaper per token than the same weights decoding at 150 tokens per second, which is why serving configuration matters more than hardware branding. We publish the throughput number the price is derived from in the table above, so the arithmetic is checkable rather than asserted.
Output tokens are priced higher than input tokens on every model here because prefill is compute-bound and parallel while decode is memory-bandwidth-bound and sequential. For a typical chat workload with a 6:1 input-to-output ratio, blended cost sits close to the input rate; for reasoning workloads that emit long hidden traces the blend moves sharply toward the output rate. Model the blend before comparing headline prices across providers.
Hosted fleet FAQ
Why host models at all if the gateway is provider-neutral?
Because someone has to set the price floor. Owning capacity on the most-requested open weights means there is always a published reference price the rest of the pool has to beat, and it gives us first-hand serving data to grade other providers with. It is also the product we charge for, which is what keeps routing free.
Do hosted models get routing preference?
No. The fleet is registered as an ordinary provider with the same published grade and the same scoring inputs as everyone else. When a third party is cheaper or faster for your policy, the router sends the request there and the route object shows you why.
Will you ever quantize a model without saying so?
No. Served precision is a published field per model and a change to it is announced before it ships. This is the same weight-integrity rule third-party providers agree to during certification.
What happens to a request pinned to a region with no capacity?
It fails with an explicit error rather than silently executing elsewhere. Residency is a hard constraint, not a routing preference.
Can we get reserved throughput instead of shared capacity?
Yes — that is what dedicated endpoints are for. You get isolated accelerators, a fixed rate limit, and predictable tail latency instead of pooled best-effort throughput.