llmcloud.ai
Hosted · Fleet

What we serve ourselves

Open weights on llmcloud accelerators. Served precision is published per model, prices are per 1M tokens, and every row is reachable through the same gateway API.

10 models
Model IDParamsPrecision servedContext$/1M in$/1M outtok/sRegionsStatus
llmcloud/deepseek-v3.2-exp
Served at the weights the lab published — no community quant, no silent swap.
685B MoEFP8 (native)164k$0.22$0.85240us-east, us-west, eu-westLive
llmcloud/llama-4-maverick
Long-context workhorse; 1M window served without a context downgrade at peak.
400B MoEBF161000k$0.28$0.90210us-east, us-west, eu-westLive
llmcloud/qwen3-235b-a22b
Strong multilingual default; the cheap alternate on several workload rows.
235B MoEBF16256k$0.20$0.70265us-west, eu-westLive
llmcloud/qwen3-coder-480b
Cheap alternate for the coding-agent workload at a fraction of frontier price.
480B MoEFP8256k$0.35$1.20190us-west, eu-westLive
llmcloud/gpt-oss-120b
High-throughput default for classification, extraction and batch jobs.
120B MoEMXFP4128k$0.10$0.40320us-east, eu-westLive
llmcloud/mistral-large-3
EU-only capacity while we finish burn-in on the Frankfurt cluster.
123B denseBF16256k$0.50$1.50140eu-westBeta
llmcloud/kimi-k2-0905
Agentic tool-use candidate; conformance probes still running.
1T MoEFP8256k$0.45$1.80120us-westBeta
llmcloud/qwen3-vl-235b
Vision and document parsing on our own GPUs. Waitlist open.
235B MoEBF16128k$0.40$1.40110us-westComing soon
llmcloud/whisper-v4
Speech-to-text on hosted capacity, billed per audio-minute. Pricing TBA.
1.6BFP16—TBATBA—us-east, eu-westComing soon
llmcloud/embed-3
Retrieval embeddings served next to the models that consume them.
7BBF1632k$0.02TBA—us-east, us-west, eu-westComing soon

These models also appear in the main catalog alongside every other provider serving the same weights — that is where you compare our price and latency against theirs. Nothing here is ranked differently because we host it.

The rates above are pay-as-you-go. The same models are also included unlimited on a flat monthly plan — $5 for everything under 50B parameters, $10 for everything under 200B, and $20 for the whole fleet plus free overflow when your Claude or OpenAI plan hits its daily limit. See pricing.

How to read this table

Precision

The precision actually served, not the precision the weights were released in. Where we serve a quantized build we say so and publish the eval delta against the reference build.

Context

The window that works end to end on our serving config, including KV cache headroom under concurrency — not the architectural maximum.

Throughput

Median decode tokens per second per stream at rated concurrency, measured on our own probes rather than taken from a launch blog post.

Regions

Jurisdictions where the weights are resident and inference executes. Pinned requests never leave the listed region, including cache and logs.

Price

USD per million tokens at list. There is no platform fee on top and no credit float — you are paying for compute.

Status

Live means in the router pool today. Coming soon means capacity is reserved and the entry exists so you can plan against it, not that it is bookable.

What a hosted token actually costs

Per-token price on open weights is not a margin decision, it is an arithmetic one: accelerator lease cost divided by realized throughput, adjusted for the output-to-input ratio of real traffic. A model that decodes at 900 tokens per second on a node costing a fixed hourly rate is roughly six times cheaper per token than the same weights decoding at 150 tokens per second, which is why serving configuration matters more than hardware branding. We publish the throughput number the price is derived from in the table above, so the arithmetic is checkable rather than asserted.

Output tokens are priced higher than input tokens on every model here because prefill is compute-bound and parallel while decode is memory-bandwidth-bound and sequential. For a typical chat workload with a 6:1 input-to-output ratio, blended cost sits close to the input rate; for reasoning workloads that emit long hidden traces the blend moves sharply toward the output rate. Model the blend before comparing headline prices across providers.

Hosted fleet FAQ

Why host models at all if the gateway is provider-neutral?

Because someone has to set the price floor. Owning capacity on the most-requested open weights means there is always a published reference price the rest of the pool has to beat, and it gives us first-hand serving data to grade other providers with. It is also the product we charge for, which is what keeps routing free.

Do hosted models get routing preference?

No. The fleet is registered as an ordinary provider with the same published grade and the same scoring inputs as everyone else. When a third party is cheaper or faster for your policy, the router sends the request there and the route object shows you why.

Will you ever quantize a model without saying so?

No. Served precision is a published field per model and a change to it is announced before it ships. This is the same weight-integrity rule third-party providers agree to during certification.

What happens to a request pinned to a region with no capacity?

It fails with an explicit error rather than silently executing elsewhere. Residency is a hard constraint, not a routing preference.

Can we get reserved throughput instead of shared capacity?

Yes — that is what dedicated endpoints are for. You get isolated accelerators, a fixed rate limit, and predictable tail latency instead of pooled best-effort throughput.