Route anywhere. Or run it on our metal.
We host open frontier weights and sovereign deployments in our own data centers. Same API key, same model IDs, published per-token prices — and no preferential routing just because we own the GPUs.
Zero margin routing
Route to any vetted provider at their list price. We take no cut, ever. Our own fleet sits in that same registry and wins traffic only when it wins on price, latency or quality.
Our own inference
Billed per token, or included unlimited on a flat plan from $5 a month. 5 open-weight models live today on our accelerators, with sovereign regions, dedicated endpoints and fine-tuning rolling out. This is the product we charge for — which is exactly why routing stays free.
Why run it here
The price floor
We publish per-token rates on our own hardware and let the router undercut us when someone else is cheaper. No credit float, no markup.
Weights you asked for
Served precision is published per model. No community quant swapped in under load, no silent model substitution.
Region pinning
Pin a request to a jurisdiction and it stays there — inference, cache and logs. Sovereign deployments extend this to operator-of-record.
Capacity we control
Reserved accelerators mean rate limits we can actually promise, and burn-in data we measure ourselves.
Your weights, our GPUs
Fine-tune an open model and serve the adapter behind the same model ID your code already uses.
No home-field advantage
Our fleet passes the same published admission bar as everyone else and carries a public trust score. Routing never prefers it.
Serving modes
| Mode | Billing | What it is | Status |
|---|---|---|---|
| Shared serverless | Per token, published rates | Pay per token on shared capacity. No minimum, no reservation, scales to zero. | Live |
| Dedicated endpoint | Per GPU-hour | Isolated replicas pinned to a region and a precision. Your own rate limits and SLA. | Beta |
| Reserved cluster | Annual commit | Reserved accelerators with capacity guarantees, custom models and support SLA. | Coming soon |
curl https://api.llmcloud.ai/v1/chat/completions \
-H "Authorization: Bearer $LLMCLOUD_KEY" \
-d '{
"model": "llmcloud/deepseek-v3.2-exp",
"messages": [{"role":"user","content":"hi"}]
}'{
"model": "auto:cost",
"provider": { "region": "eu-west" },
"messages": [{"role":"user","content":"hi"}]
}Why a neutral gateway also owns hardware
A router that owns no capacity has no independent view of what inference should cost. By serving the most-requested open weights ourselves we establish a published reference price the rest of the pool has to beat, and we generate first-hand serving data — real decode throughput, real tail latency under concurrency, real failure behaviour during capacity events — that we use to grade third-party providers. Running the workload is the only honest way to audit the people you route to.
It is also what pays for the gateway. Routing carries no margin and never will, so the business has to earn its revenue somewhere that is not a tax on your tokens: flat Team seats and compute we actually operate. That separation is deliberate — the moment routing revenue depends on which provider wins a request, the routing stops being trustworthy.
What sovereign hosting means here
Sovereign is not a marketing region label. For a deployment to qualify, the weights must be resident in the jurisdiction, inference and KV cache must execute inside it, operational access must be held by personnel subject to local law, and no telemetry containing prompt or completion content may leave it. Requests pinned to a sovereign region fail rather than spill over when capacity is short. Where a jurisdiction is on the roadmap instead of live, it is marked coming soon rather than quietly served from elsewhere.
Hosted inference FAQ
Does the router favour your own models?+
No. The hosted fleet is registered as an ordinary provider with the same trust score and the same scoring inputs as everyone else. When a third party is cheaper or faster for your policy the router sends the request there, and the route object in the response shows the decision.
How is this consistent with zero margin?+
Routing is free forever and tokens are billed at provider list price. Hosted inference is a separate product with its own published price — you are paying for compute we operate, not for the privilege of routing.
What happens if your capacity runs out?+
The gateway fails over to another qualified provider serving the same model and records the substitution in the response metadata, unless a residency or provider policy forbids it, in which case the request fails explicitly instead of moving.
Can I get a dedicated endpoint today?+
Dedicated endpoints are in beta with design-partner customers. Reserved clusters and fine-tuning are on the roadmap; the waitlist gates access rather than a sales cycle.
Do you quantize models to cut costs?+
Only when it is published. Served precision is a field on every hosted model and a change ships with the eval delta against the reference build. Silent quantization is the single failure that removes a provider from our pool, so we hold the fleet to it as well.
Which models do you host and how are they chosen?+
Open-weight frontier models with sustained request volume through the gateway, plus models required for sovereign deployments where no compliant third-party endpoint exists. Demand data from routing decides the fleet, not vendor partnerships.
Prices and performance figures on hosted pages are illustrative until published measurement lands. See provider standards for how we measure ourselves.