llmcloud.ai
Pricing

Unlimited from $5. Or tokens at cost, forever.

Two ways to buy. Flat monthly plans give you unlimited inference on the open-weight models we host ourselves — $5 under 50B, $10 under 200B, $20 to extend your Claude or OpenAI plan so a daily limit never stops you again. Everything else stays metered at provider cost with a $0 gateway fee.

Unlimited plans

flat monthly · hosted on our own GPUs

Lite

Sold out
$5per month · flat

Unlimited on every hosted open-weight model under 50B parameters.

  • Unlimited messages and tokens in the band — no per-token bill at all
  • Runs on llmcloud GPUs, served precision published per model
  • Same OpenAI-compatible API, same keys, same routing
  • Fair use: 20 requests / minute, 1 concurrent stream, personal use

e.g. Llama 3.3 8B, Qwen3 32B, Mistral Small, Gemma 27B

Sold out — join waitlist

Standard

Sold out
$10per month · flat

Everything in Lite, plus every hosted model under 200B parameters.

  • Unlimited on the mid and large open-weight fleet, including MoE builds
  • Long-context windows at full served length, no surcharge
  • Smart routing inside the band — auto:cost and auto:quality
  • Fair use: 60 requests / minute, 3 concurrent streams, personal use

e.g. DeepSeek V3.1, Qwen3 235B MoE, Llama 3.1 70B, GLM 4.6

Sold out — join waitlist

Max

Most popular
$20per month · flat

Everything in Standard, plus your Claude or OpenAI plan extended so a daily limit never stops you.

  • The $20 goes toward your frontier lab subscription — we take no margin on it
  • Hit Anthropic's or OpenAI's daily cap and we keep you going on open-weight models, free
  • Overflow is automatic: same conversation, same API call, no key swapping
  • Unlimited on the whole hosted fleet, any parameter band
  • Fair use: 120 requests / minute, 5 concurrent streams, personal use

e.g. Claude and GPT via your own plan + the full llmcloud fleet as overflow

Start Max
What each unlimited plan includes by model parameter band
What's included by parameter bandLite · $5Standard · $10Max · $20
Hosted open weights under 50BUnlimitedUnlimitedUnlimited
Hosted open weights 50B – 200BUnlimitedUnlimited
Hosted open weights above 200BUnlimited
Closed frontier models (Claude, GPT, Gemini)Per token, at costPer token, at costYour own lab plan, extended
Overflow when the lab caps you for the dayCovered free on open weights
Gateway fee on anything metered$0$0$0
BYOK routing

How Max works with your Claude or OpenAI plan

never hit a daily limit again
01

Keep the plan you already pay for

Stay on Claude Pro or ChatGPT Plus. Connect it once; we never proxy or resell it, and your $20 with us is applied against that subscription cost rather than kept as margin.

02

Work until the lab cuts you off

Every frontier plan has a daily or rolling cap. Instead of the usual 'try again in four hours', llmcloud sees the limit response and takes over mid-session.

03

Carry on, unlimited and free

Your request re-routes to the best open-weight model for that workload on our own GPUs. No extra charge, no token meter, same conversation and the same API call.

The $20 is applied to your frontier lab subscription cost — llmcloud takes no margin on it. The overflow inference you get when the lab cuts you off runs on our own GPUs and costs you nothing extra, because open weights on our metal are cheap enough for us to absorb and it is the fastest way to show you what they can do.

Team & Enterprise platform

per seat · tokens at cost

Seats are a separate purchase from the unlimited plans and you can hold both: the plan covers your own inference, the seat buys org-level routing, analytics, retained logs and compliance for the team around you. The gateway fee is $0 either way.

Developer

$0forever · tokens at cost

The whole catalog, no fee, no seat, no credit card.

  • 0% routing fee — you pay the provider's list price
  • All 150+ models across 21 vetted providers
  • Failover routing: retry and reroute when a provider errors
  • Basic analytics — last 24 hours, aggregate spend and requests
  • No log retention: requests are not stored after the response
  • One personal workspace, BYOK with no surcharge
Create account

Team

Most teams
$10per user / month · tokens still at cost

Org-level control, memory and an opinion per workload.

  • Everything in Developer, gateway fee still $0
  • Smart routing: auto:cost, auto:quality, workload-aware directives, fallback chains
  • Org analytics per member, key, model and project — latency, errors, cost anomalies
  • 90-day log retention, request inspection, CSV and warehouse export
  • Semantic cache with a private namespace and per-project hit rebates
  • Roles, per-member and per-key spend caps with hard budget enforcement
  • Scorecard-driven model recommendations per workload
Start a team

Enterprise

Customannual

Dedicated capacity and the sovereign compliance pack.

  • Everything in Team
  • Dedicated hosted models on llmcloud GPUs — reserved replicas, pinned region and precision
  • 99.9–99.99% SLA with service credits and your own rate limits
  • Sovereign pack: residency pinning, EU AI Act artefacts, signed evidence bundles, SIEM export
  • SAML SSO + SCIM, granular RBAC, custom log retention, zero-retention routing
  • Self-host in your VPC or on-prem under licence
  • Named solutions engineer, shared Slack, P1 on-call
Talk to us

Compare tiers

CapabilityDeveloper · $0Team · $10 / user / moEnterprise · custom
Routing
Gateway fee on tokens$0$0$0
Failover & retries
auto:cost / auto:quality
Workload-aware directives
Custom fallback chains & pinning
Analytics & logs
Spend analyticsLast 24h, aggregateOrg-wide, per member/key/modelOrg-wide + custom dimensions
Log retentionNone90 daysCustom / zero-retention
Request & response inspection
CSV / warehouse export✓ + SIEM push
Cost anomaly alerts
Caching
Semantic cacheShared namespacePrivate namespacePrivate + regional pinning
Cache-hit rebates on invoice✓ per project✓ per project
Org & access
Workspaces1 personalUnlimited sharedUnlimited + projects
Roles & permissionsGranular RBAC
Spend caps & budgetsPer keyPer member, key and projectPer jurisdiction and project
SAML SSO / SCIM
Hosted & compliance
Hosted models (per-token)
Dedicated endpoints & reserved GPUs
Residency pinning & sovereign regions
Signed attestations, DPA, HIPAA BAADPA✓ full pack
Self-host licence
Support
ChannelCommunityEmailNamed SE + shared Slack
SLABest effort99.9–99.99% contracted

The math on one million tokens

Provider list price (e.g. 1M in / 1M out)$3.00
llmcloud routing fee$0.00
Cache hit rebate (typical 28% hit rate)−$0.84
You pay$2.16

A 5% gateway fee on the same traffic would add $0.15 per million and $1,500 per month at $30k of spend. A ten-person team on llmcloud pays $100 — flat, whatever the volume.

And on an unlimited plan the same million tokens is $0 in the band you bought: heavy personal use that would meter at $40–$60 a month lands inside Lite at $5 or Standard at $10.

How billing works

  • · Metered per 1M input / output tokens, per provider, per model.
  • · Seats are billed monthly per active member; tokens are billed separately in arrears.
  • · Cache hits bill $0 tokens and no lookup fee.
  • · Failover retries to a backup are never double-charged.
  • · Provider discounts we negotiate appear as a credit line, not our spread.
  • · Invoices export as CSV, JSON, or push to your warehouse.

Hosted inference on our metal

where we do earn margin

Open frontier and sovereign models we serve ourselves are billed per token like any other provider — published rate, no gateway fee on top. Our fleet is listed as an ordinary provider and wins traffic only when it is genuinely the best route.

  • · Per-token, pay-as-you-go on every tier. No minimum, no reservation required.
  • · Served precision published per model — no silent quantization.
  • · Dedicated endpoints billed per GPU-hour, an Enterprise entitlement (Beta).
  • · Fine-tuned adapter serving priced at base-model rates (coming soon).

Frequently asked

What does 'unlimited' actually mean on the $5, $10 and $20 plans?+

No token meter and no monthly quota on the models in your band — you pay the flat price whatever you consume. What is bounded is concurrency, not volume: a per-minute request rate limit and a concurrent-stream limit per plan (20/1 on Lite, 60/3 on Standard, 120/5 on Max), personal single-person use, and no bulk generation of training data or resale of access.

Which models fall into each parameter band?+

Lite covers hosted open weights under 50B parameters, such as Llama 3.3 8B, Qwen3 32B and Gemma 27B. Standard adds everything under 200B, including DeepSeek V3.1 and Qwen3 235B MoE class builds where the active parameter count sits in band. Max covers the whole hosted fleet with no cap. Every model page states its band.

Does the $20 Max plan replace my Claude or OpenAI subscription?+

No — it extends it. You keep paying the lab directly for the model you like; the $20 you pay us is applied toward that frontier subscription cost, and what you get from llmcloud is unlimited open-weight coverage the moment the lab's daily limit stops you.

What exactly happens when the lab cuts me off?+

We detect the rate-limit response and re-route that request to the best open-weight model on our own GPUs for that workload, in the same session and through the same API call. It is free — no token charge, no overage. When the lab's window resets you go straight back to the frontier model.

Can I have a subscription and per-token billing at the same time?+

Yes, and most people do. Anything inside your band is covered by the flat plan; anything outside it — a closed frontier model called directly, for example — is metered at provider cost with the usual $0 gateway fee.

Do the unlimited plans apply to BYOK?+

BYOK is free on every plan, but it is a separate thing: with your own key you pay that provider directly and we add nothing. The unlimited plans cover inference we serve on our own hardware.

Where is the catch on 'zero margin'?+

There isn't one on tokens. We pass the provider's list price through and itemise it; the gateway line reads $0.00 on every tier. We earn on seats ($10 per user per month for Team), on Enterprise controls and dedicated capacity, on self-host licences, and on our own hosted fleet where we are the provider and take a normal inference margin.

What does the $10 seat actually buy?+

Intelligence and memory. Smart routing (auto:cost, auto:quality, workload-aware directives and fallback chains), org-level analytics broken down per member, key, model and project, 90-day log retention with request inspection and export, a private semantic cache namespace, and enforced spend caps. Tokens stay at provider cost.

What counts as a user?+

Any member with access to the org workspace in a given month. Flat $10 each, no tiering by usage, no per-request platform fee. API keys and service accounts are free — you only pay for humans.

What happens to my logs on Developer?+

Nothing is retained. Requests and responses are not stored after the response is returned, and analytics are limited to a rolling 24-hour aggregate. If you need history to answer 'what did we spend last month and why did that request go to that model', that is Team.

Do I need Enterprise for dedicated models?+

Yes for reserved capacity. Every tier can call our hosted fleet per token as an ordinary provider. Dedicated endpoints — isolated replicas billed per GPU-hour, with pinned region and precision and a contracted SLA — are an Enterprise entitlement.

Why no prepaid credits?+

Credit balances are float — your money sitting on our balance sheet, often non-refundable. You pay in arrears for tokens you actually consumed, in your own currency, and you can cap spend per key.

Is BYOK really free?+

Yes. Attach your own provider keys and we charge nothing to route through them on any tier. Several gateways bill a per-request BYOK fee; we do not.

What if a provider gives you a volume discount?+

It flows to you. Negotiated rates below list appear as a discount line on your invoice rather than becoming our spread.