Unlimited from $5. Or tokens at cost, forever.
Two ways to buy. Flat monthly plans give you unlimited inference on the open-weight models we host ourselves — $5 under 50B, $10 under 200B, $20 to extend your Claude or OpenAI plan so a daily limit never stops you again. Everything else stays metered at provider cost with a $0 gateway fee.
Unlimited plans
flat monthly · hosted on our own GPUsLite
Sold outUnlimited on every hosted open-weight model under 50B parameters.
- Unlimited messages and tokens in the band — no per-token bill at all
- Runs on llmcloud GPUs, served precision published per model
- Same OpenAI-compatible API, same keys, same routing
- Fair use: 20 requests / minute, 1 concurrent stream, personal use
e.g. Llama 3.3 8B, Qwen3 32B, Mistral Small, Gemma 27B
Sold out — join waitlistStandard
Sold outEverything in Lite, plus every hosted model under 200B parameters.
- Unlimited on the mid and large open-weight fleet, including MoE builds
- Long-context windows at full served length, no surcharge
- Smart routing inside the band — auto:cost and auto:quality
- Fair use: 60 requests / minute, 3 concurrent streams, personal use
e.g. DeepSeek V3.1, Qwen3 235B MoE, Llama 3.1 70B, GLM 4.6
Sold out — join waitlistMax
Most popularEverything in Standard, plus your Claude or OpenAI plan extended so a daily limit never stops you.
- The $20 goes toward your frontier lab subscription — we take no margin on it
- Hit Anthropic's or OpenAI's daily cap and we keep you going on open-weight models, free
- Overflow is automatic: same conversation, same API call, no key swapping
- Unlimited on the whole hosted fleet, any parameter band
- Fair use: 120 requests / minute, 5 concurrent streams, personal use
e.g. Claude and GPT via your own plan + the full llmcloud fleet as overflow
Start Max →| What's included by parameter band | Lite · $5 | Standard · $10 | Max · $20 |
|---|---|---|---|
| Hosted open weights under 50B | Unlimited | Unlimited | Unlimited |
| Hosted open weights 50B – 200B | — | Unlimited | Unlimited |
| Hosted open weights above 200B | — | — | Unlimited |
| Closed frontier models (Claude, GPT, Gemini) | Per token, at cost | Per token, at cost | Your own lab plan, extended |
| Overflow when the lab caps you for the day | — | — | Covered free on open weights |
| Gateway fee on anything metered | $0 | $0 | $0 |
| BYOK routing | ✓ | ✓ | ✓ |
How Max works with your Claude or OpenAI plan
never hit a daily limit againKeep the plan you already pay for
Stay on Claude Pro or ChatGPT Plus. Connect it once; we never proxy or resell it, and your $20 with us is applied against that subscription cost rather than kept as margin.
Work until the lab cuts you off
Every frontier plan has a daily or rolling cap. Instead of the usual 'try again in four hours', llmcloud sees the limit response and takes over mid-session.
Carry on, unlimited and free
Your request re-routes to the best open-weight model for that workload on our own GPUs. No extra charge, no token meter, same conversation and the same API call.
The $20 is applied to your frontier lab subscription cost — llmcloud takes no margin on it. The overflow inference you get when the lab cuts you off runs on our own GPUs and costs you nothing extra, because open weights on our metal are cheap enough for us to absorb and it is the fastest way to show you what they can do.
Team & Enterprise platform
per seat · tokens at costSeats are a separate purchase from the unlimited plans and you can hold both: the plan covers your own inference, the seat buys org-level routing, analytics, retained logs and compliance for the team around you. The gateway fee is $0 either way.
Developer
The whole catalog, no fee, no seat, no credit card.
- 0% routing fee — you pay the provider's list price
- All 150+ models across 21 vetted providers
- Failover routing: retry and reroute when a provider errors
- Basic analytics — last 24 hours, aggregate spend and requests
- No log retention: requests are not stored after the response
- One personal workspace, BYOK with no surcharge
Team
Most teamsOrg-level control, memory and an opinion per workload.
- Everything in Developer, gateway fee still $0
- Smart routing: auto:cost, auto:quality, workload-aware directives, fallback chains
- Org analytics per member, key, model and project — latency, errors, cost anomalies
- 90-day log retention, request inspection, CSV and warehouse export
- Semantic cache with a private namespace and per-project hit rebates
- Roles, per-member and per-key spend caps with hard budget enforcement
- Scorecard-driven model recommendations per workload
Enterprise
Dedicated capacity and the sovereign compliance pack.
- Everything in Team
- Dedicated hosted models on llmcloud GPUs — reserved replicas, pinned region and precision
- 99.9–99.99% SLA with service credits and your own rate limits
- Sovereign pack: residency pinning, EU AI Act artefacts, signed evidence bundles, SIEM export
- SAML SSO + SCIM, granular RBAC, custom log retention, zero-retention routing
- Self-host in your VPC or on-prem under licence
- Named solutions engineer, shared Slack, P1 on-call
Compare tiers
| Capability | Developer · $0 | Team · $10 / user / mo | Enterprise · custom |
|---|---|---|---|
| Routing | |||
| Gateway fee on tokens | $0 | $0 | $0 |
| Failover & retries | ✓ | ✓ | ✓ |
| auto:cost / auto:quality | — | ✓ | ✓ |
| Workload-aware directives | — | ✓ | ✓ |
| Custom fallback chains & pinning | — | ✓ | ✓ |
| Analytics & logs | |||
| Spend analytics | Last 24h, aggregate | Org-wide, per member/key/model | Org-wide + custom dimensions |
| Log retention | None | 90 days | Custom / zero-retention |
| Request & response inspection | — | ✓ | ✓ |
| CSV / warehouse export | — | ✓ | ✓ + SIEM push |
| Cost anomaly alerts | — | ✓ | ✓ |
| Caching | |||
| Semantic cache | Shared namespace | Private namespace | Private + regional pinning |
| Cache-hit rebates on invoice | ✓ | ✓ per project | ✓ per project |
| Org & access | |||
| Workspaces | 1 personal | Unlimited shared | Unlimited + projects |
| Roles & permissions | — | ✓ | Granular RBAC |
| Spend caps & budgets | Per key | Per member, key and project | Per jurisdiction and project |
| SAML SSO / SCIM | — | — | ✓ |
| Hosted & compliance | |||
| Hosted models (per-token) | ✓ | ✓ | ✓ |
| Dedicated endpoints & reserved GPUs | — | — | ✓ |
| Residency pinning & sovereign regions | — | — | ✓ |
| Signed attestations, DPA, HIPAA BAA | — | DPA | ✓ full pack |
| Self-host licence | — | — | ✓ |
| Support | |||
| Channel | Community | Named SE + shared Slack | |
| SLA | — | Best effort | 99.9–99.99% contracted |
The math on one million tokens
A 5% gateway fee on the same traffic would add $0.15 per million and $1,500 per month at $30k of spend. A ten-person team on llmcloud pays $100 — flat, whatever the volume.
And on an unlimited plan the same million tokens is $0 in the band you bought: heavy personal use that would meter at $40–$60 a month lands inside Lite at $5 or Standard at $10.
How billing works
- · Metered per 1M input / output tokens, per provider, per model.
- · Seats are billed monthly per active member; tokens are billed separately in arrears.
- · Cache hits bill $0 tokens and no lookup fee.
- · Failover retries to a backup are never double-charged.
- · Provider discounts we negotiate appear as a credit line, not our spread.
- · Invoices export as CSV, JSON, or push to your warehouse.
Hosted inference on our metal
where we do earn marginOpen frontier and sovereign models we serve ourselves are billed per token like any other provider — published rate, no gateway fee on top. Our fleet is listed as an ordinary provider and wins traffic only when it is genuinely the best route.
- · Per-token, pay-as-you-go on every tier. No minimum, no reservation required.
- · Served precision published per model — no silent quantization.
- · Dedicated endpoints billed per GPU-hour, an Enterprise entitlement (Beta).
- · Fine-tuned adapter serving priced at base-model rates (coming soon).
Frequently asked
What does 'unlimited' actually mean on the $5, $10 and $20 plans?+
No token meter and no monthly quota on the models in your band — you pay the flat price whatever you consume. What is bounded is concurrency, not volume: a per-minute request rate limit and a concurrent-stream limit per plan (20/1 on Lite, 60/3 on Standard, 120/5 on Max), personal single-person use, and no bulk generation of training data or resale of access.
Which models fall into each parameter band?+
Lite covers hosted open weights under 50B parameters, such as Llama 3.3 8B, Qwen3 32B and Gemma 27B. Standard adds everything under 200B, including DeepSeek V3.1 and Qwen3 235B MoE class builds where the active parameter count sits in band. Max covers the whole hosted fleet with no cap. Every model page states its band.
Does the $20 Max plan replace my Claude or OpenAI subscription?+
No — it extends it. You keep paying the lab directly for the model you like; the $20 you pay us is applied toward that frontier subscription cost, and what you get from llmcloud is unlimited open-weight coverage the moment the lab's daily limit stops you.
What exactly happens when the lab cuts me off?+
We detect the rate-limit response and re-route that request to the best open-weight model on our own GPUs for that workload, in the same session and through the same API call. It is free — no token charge, no overage. When the lab's window resets you go straight back to the frontier model.
Can I have a subscription and per-token billing at the same time?+
Yes, and most people do. Anything inside your band is covered by the flat plan; anything outside it — a closed frontier model called directly, for example — is metered at provider cost with the usual $0 gateway fee.
Do the unlimited plans apply to BYOK?+
BYOK is free on every plan, but it is a separate thing: with your own key you pay that provider directly and we add nothing. The unlimited plans cover inference we serve on our own hardware.
Where is the catch on 'zero margin'?+
There isn't one on tokens. We pass the provider's list price through and itemise it; the gateway line reads $0.00 on every tier. We earn on seats ($10 per user per month for Team), on Enterprise controls and dedicated capacity, on self-host licences, and on our own hosted fleet where we are the provider and take a normal inference margin.
What does the $10 seat actually buy?+
Intelligence and memory. Smart routing (auto:cost, auto:quality, workload-aware directives and fallback chains), org-level analytics broken down per member, key, model and project, 90-day log retention with request inspection and export, a private semantic cache namespace, and enforced spend caps. Tokens stay at provider cost.
What counts as a user?+
Any member with access to the org workspace in a given month. Flat $10 each, no tiering by usage, no per-request platform fee. API keys and service accounts are free — you only pay for humans.
What happens to my logs on Developer?+
Nothing is retained. Requests and responses are not stored after the response is returned, and analytics are limited to a rolling 24-hour aggregate. If you need history to answer 'what did we spend last month and why did that request go to that model', that is Team.
Do I need Enterprise for dedicated models?+
Yes for reserved capacity. Every tier can call our hosted fleet per token as an ordinary provider. Dedicated endpoints — isolated replicas billed per GPU-hour, with pinned region and precision and a contracted SLA — are an Enterprise entitlement.
Why no prepaid credits?+
Credit balances are float — your money sitting on our balance sheet, often non-refundable. You pay in arrears for tokens you actually consumed, in your own currency, and you can cap spend per key.
Is BYOK really free?+
Yes. Attach your own provider keys and we charge nothing to route through them on any tier. Several gateways bill a per-request BYOK fee; we do not.
What if a provider gives you a volume discount?+
It flows to you. Negotiated rates below list appear as a discount line on your invoice rather than becoming our spread.