Usage-based. No seats. No lock-in.
Pay providers at cost, plus a small routing fee. Upgrade only when you need analytics, guardrails, or enterprise controls.
Free
Kick the tires. 150+ models, one key.
- $1 in free inference credits
- Access to all 150+ models
- Smart routing (auto:cost, auto:quality)
- Semantic caching (shared cache)
- Community support
Pro
Most popularFor teams shipping to production.
- Pay provider cost + 5% routing fee
- Private semantic cache
- Cost & latency analytics
- Guardrails (PII, jailbreak, schema)
- Team API keys with per-key budgets
- Email support · 99.9% SLO
Enterprise
SSO, audit logs, BYOK, self-host.
- SAML SSO + SCIM provisioning
- SOC 2 audit logs & data residency
- BYOK across all providers
- Self-host in your VPC / on-prem
- Dedicated routing pool + 99.99% SLA
- Named solutions engineer
Add-ons
How billing works
- · Metered per 1M input / output tokens per provider, per model.
- · Cached hits are billed at $0 tokens + a $0.0001 lookup fee.
- · Failover retries to a cheaper backup are not double-charged.
- · Invoices export as CSV, JSON, or push to your data warehouse.
Frequently asked
How does the 5% routing fee work?+
You pay the underlying provider's list price per token, plus a flat 5% for routing, caching, failover, and analytics. No markup on the base tokens themselves — invoices line-item both.
Do I need a credit card to start?+
No. Sign up with email or GitHub and you get $1 in free credits — enough to run several thousand requests on smaller models.
What about BYOK (bring-your-own-key)?+
On Pro and Enterprise you can route through your own provider keys. In that mode we only charge the 5% routing fee — provider tokens are billed by them directly.
Can I self-host llmcloud?+
Yes. Enterprise includes a container image plus Terraform for AWS, GCP, and Azure. The control plane can stay managed or run fully in your VPC.