Tokens on tap — capacity is our problem
Shared, always-warm endpoints for the most-used open models. You don't deploy anything or pay for idle time. Change one base URL and pay per token at a published rate.
| Model | $/1M in | $/1M out | tok/s | TTFT | Status |
|---|---|---|---|---|---|
| llmcloud/deepseek-v3.2-exp | $0.22 | $0.85 | 240 | 190ms | Live |
| llmcloud/llama-4-maverick | $0.28 | $0.90 | 210 | 150ms | Live |
| llmcloud/qwen3-235b-a22b | $0.20 | $0.70 | 265 | 130ms | Live |
| llmcloud/qwen3-coder-480b | $0.35 | $1.20 | 190 | 210ms | Live |
| llmcloud/gpt-oss-120b | $0.10 | $0.40 | 320 | 95ms | Live |
| llmcloud/mistral-large-3 | $0.50 | $1.50 | 140 | 220ms | Beta |
| llmcloud/kimi-k2-0905 | $0.45 | $1.80 | 120 | 280ms | Beta |
| llmcloud/qwen3-vl-235b | $0.40 | $1.40 | 110 | 320ms | Coming soon |
What you get
Speculative decoding by default
Draft models and custom kernels on H200 and B200 keep decode speed high, even with long prompts.
Prompt caching
Repeated prefixes are billed at 25% of the input rate. Agent loops and RAG with shared system prompts get cheaper automatically.
Rate limits that scale
Limits start at 600 requests/min and rise with usage. Need a hard guarantee? Move to dedicated capacity.
Structured outputs
JSON schema mode, function calling and grammar-constrained decoding work on every live text model.
Published precision
Each model lists the precision it's served at. We never swap in a lighter quantization under load.
Gateway fallback
If a model is at capacity, or we don't host it, the request can route to a vetted third party at list price. You opt in per request.
from openai import OpenAI
client = OpenAI(base_url="https://api.llmcloud.ai/v1", api_key=LLMCLOUD_KEY)
client.chat.completions.create(
model="llmcloud/deepseek-v3.2-exp",
messages=[{"role": "user", "content": "Summarise this diff"}],
stream=True,
)Serverless FAQ
Are there cold starts on serverless?+
No. Serverless models stay warm on shared capacity at all times. Cold starts apply only to dedicated endpoints that scale to zero.
How is serverless billed?+
Per million input and output tokens, at the rate shown for each model. Cached input tokens cost 25% of the input rate. Or use an unlimited flat plan instead of per-token billing.
What if you don't host the model I need?+
Turn on gateway fallback, and the request goes to a vetted third-party provider at their list price with no markup. The response metadata shows where it ran.
Figures are illustrative until published measurement is available. See the full hosted fleet.