llmcloud.ai
Hosted · llmcloud Inference

Route anywhere. Or run it on our metal.

We host open frontier weights and sovereign deployments in our own data centers. Same API key, same model IDs, published per-token prices — and no preferential routing just because we own the GPUs.

Gateway

Zero margin routing

Route to any vetted provider at their list price. We take no cut, ever. Our own fleet sits in that same registry and wins traffic only when it wins on price, latency or quality.

Hosted

Our own inference

Billed per token, or included unlimited on a flat plan from $5 a month. 5 open-weight models live today on our accelerators, with sovereign regions, dedicated endpoints and fine-tuning rolling out. This is the product we charge for — which is exactly why routing stays free.

Why run it here

price

The price floor

We publish per-token rates on our own hardware and let the router undercut us when someone else is cheaper. No credit float, no markup.

integrity

Weights you asked for

Served precision is published per model. No community quant swapped in under load, no silent model substitution.

residency

Region pinning

Pin a request to a jurisdiction and it stays there — inference, cache and logs. Sovereign deployments extend this to operator-of-record.

capacity

Capacity we control

Reserved accelerators mean rate limits we can actually promise, and burn-in data we measure ourselves.

tuning

Your weights, our GPUs

Fine-tune an open model and serve the adapter behind the same model ID your code already uses.

neutrality

No home-field advantage

Our fleet passes the same published admission bar as everyone else and carries a public trust score. Routing never prefers it.

Serving modes

ModeBillingWhat it isStatus
Shared serverlessPer token, published ratesPay per token on shared capacity. No minimum, no reservation, scales to zero.Live
Dedicated endpointPer GPU-hourIsolated replicas pinned to a region and a precision. Your own rate limits and SLA.Beta
Reserved clusterAnnual commitReserved accelerators with capacity guarantees, custom models and support SLA.Coming soon
Ask for our fleet explicitly
curl https://api.llmcloud.ai/v1/chat/completions \
  -H "Authorization: Bearer $LLMCLOUD_KEY" \
  -d '{
    "model": "llmcloud/deepseek-v3.2-exp",
    "messages": [{"role":"user","content":"hi"}]
  }'
Or let the router decide (it may pick us, or not)
{
  "model": "auto:cost",
  "provider": { "region": "eu-west" },
  "messages": [{"role":"user","content":"hi"}]
}

Why a neutral gateway also owns hardware

A router that owns no capacity has no independent view of what inference should cost. By serving the most-requested open weights ourselves we establish a published reference price the rest of the pool has to beat, and we generate first-hand serving data — real decode throughput, real tail latency under concurrency, real failure behaviour during capacity events — that we use to grade third-party providers. Running the workload is the only honest way to audit the people you route to.

It is also what pays for the gateway. Routing carries no margin and never will, so the business has to earn its revenue somewhere that is not a tax on your tokens: flat Team seats and compute we actually operate. That separation is deliberate — the moment routing revenue depends on which provider wins a request, the routing stops being trustworthy.

What sovereign hosting means here

Sovereign is not a marketing region label. For a deployment to qualify, the weights must be resident in the jurisdiction, inference and KV cache must execute inside it, operational access must be held by personnel subject to local law, and no telemetry containing prompt or completion content may leave it. Requests pinned to a sovereign region fail rather than spill over when capacity is short. Where a jurisdiction is on the roadmap instead of live, it is marked coming soon rather than quietly served from elsewhere.

Hosted inference FAQ

Does the router favour your own models?+

No. The hosted fleet is registered as an ordinary provider with the same trust score and the same scoring inputs as everyone else. When a third party is cheaper or faster for your policy the router sends the request there, and the route object in the response shows the decision.

How is this consistent with zero margin?+

Routing is free forever and tokens are billed at provider list price. Hosted inference is a separate product with its own published price — you are paying for compute we operate, not for the privilege of routing.

What happens if your capacity runs out?+

The gateway fails over to another qualified provider serving the same model and records the substitution in the response metadata, unless a residency or provider policy forbids it, in which case the request fails explicitly instead of moving.

Can I get a dedicated endpoint today?+

Dedicated endpoints are in beta with design-partner customers. Reserved clusters and fine-tuning are on the roadmap; the waitlist gates access rather than a sales cycle.

Do you quantize models to cut costs?+

Only when it is published. Served precision is a field on every hosted model and a change ships with the eval delta against the reference build. Silent quantization is the single failure that removes a provider from our pool, so we hold the fleet to it as well.

Which models do you host and how are they chosen?+

Open-weight frontier models with sustained request volume through the gateway, plus models required for sovereign deployments where no compliant third-party endpoint exists. Demand data from routing decides the fleet, not vendor partnerships.

Prices and performance figures on hosted pages are illustrative until published measurement lands. See provider standards for how we measure ourselves.