llmcloud.ai
Products · Serverless

Tokens on tap — capacity is our problem

Shared, always-warm endpoints for the most-used open models. You don't deploy anything or pay for idle time. Change one base URL and pay per token at a published rate.

Model$/1M in$/1M outtok/sTTFTStatus
llmcloud/deepseek-v3.2-exp$0.22$0.85240190msLive
llmcloud/llama-4-maverick$0.28$0.90210150msLive
llmcloud/qwen3-235b-a22b$0.20$0.70265130msLive
llmcloud/qwen3-coder-480b$0.35$1.20190210msLive
llmcloud/gpt-oss-120b$0.10$0.4032095msLive
llmcloud/mistral-large-3$0.50$1.50140220msBeta
llmcloud/kimi-k2-0905$0.45$1.80120280msBeta
llmcloud/qwen3-vl-235b$0.40$1.40110320msComing soon

What you get

speed

Speculative decoding by default

Draft models and custom kernels on H200 and B200 keep decode speed high, even with long prompts.

cache

Prompt caching

Repeated prefixes are billed at 25% of the input rate. Agent loops and RAG with shared system prompts get cheaper automatically.

limits

Rate limits that scale

Limits start at 600 requests/min and rise with usage. Need a hard guarantee? Move to dedicated capacity.

format

Structured outputs

JSON schema mode, function calling and grammar-constrained decoding work on every live text model.

integrity

Published precision

Each model lists the precision it's served at. We never swap in a lighter quantization under load.

overflow

Gateway fallback

If a model is at capacity, or we don't host it, the request can route to a vetted third party at list price. You opt in per request.

OpenAI SDK, one line changed
from openai import OpenAI
client = OpenAI(base_url="https://api.llmcloud.ai/v1", api_key=LLMCLOUD_KEY)

client.chat.completions.create(
    model="llmcloud/deepseek-v3.2-exp",
    messages=[{"role": "user", "content": "Summarise this diff"}],
    stream=True,
)

Serverless FAQ

Are there cold starts on serverless?+

No. Serverless models stay warm on shared capacity at all times. Cold starts apply only to dedicated endpoints that scale to zero.

How is serverless billed?+

Per million input and output tokens, at the rate shown for each model. Cached input tokens cost 25% of the input rate. Or use an unlimited flat plan instead of per-token billing.

What if you don't host the model I need?+

Turn on gateway fallback, and the request goes to a vetted third-party provider at their list price with no markup. The response metadata shows where it ran.

Figures are illustrative until published measurement is available. See the full hosted fleet.