llmcloud.ai
Platform · Analytics

Every request, priced and traced.

Spend, tokens, latency and errors — sliced by model, provider, key, project or user. No SDK: the router logs it.

MVP · join waitlist
Spend · 30d
$1,284
+12%
Tokens · 30d
42.7M
+8%
p50 TTFT
187ms
-14ms
Error rate
0.09%
-0.03pp
Top spend · 30d
by model
$1,284.03
anthropic/claude-4.5-sonnet
$539.30
openai/gpt-5
$308.16
google/gemini-3-pro
$192.60
fireworks/deepseek-v3.1
$141.24
auto:cost (routed)
$102.73

What you can slice

spend

Cost per anything

Group by model, provider, region, key, project, user, or any custom metadata tag.

perf

Latency histograms

p50/p95/p99 for TTFT and total, per model and route. Catch regressions early.

reliability

Error & retry rates

4xx / 5xx / timeout / retry counts per upstream. Every failover records the winner.

logs

Per-request logs

Full payloads, redacted on ingest. Filter by status, cost, latency, model or text.

tags

Custom metadata

Attach x-llmcloud-tag headers — user IDs, flags, agent names — and slice on them.

export

Export & webhooks

CSV, Parquet, S3 sync, and spend webhooks. Snowflake, BigQuery, Datadog on the roadmap.

How long are logs retained?+

30 days free, 90 days on Pro, configurable on Enterprise. Redaction runs on ingest, so flagged prompts never persist.

Can I get spend alerts?+

Yes — daily and monthly thresholds per key, delivered by email, Slack or webhook. Pair with per-key budgets for a hard stop.

Do you show cache hit rates?+

Yes. Every request carries cache_status (miss, exact, semantic) so you can measure caching ROI.