llmcloud.ai
Platform · Analytics

Every request, priced and traced.

Real-time spend, tokens, latency and error rates broken down by model, provider, key, project and end-user. No SDK to install — the router logs it all.

MVP · join waitlist
Spend · 30d
$1,284
+12%
Tokens · 30d
42.7M
+8%
p50 TTFT
187ms
-14ms
Error rate
0.09%
-0.03pp
Top spend · 30d
by model
$1,284.03
anthropic/claude-4.5-sonnet
$539.30
openai/gpt-5
$308.16
google/gemini-3-pro
$192.60
fireworks/deepseek-v3.1
$141.24
auto:cost (routed)
$102.73

What you can slice

spend

Cost per anything

Group by model, provider, region, key, project, end-user or a custom metadata tag you pass on the request.

perf

Latency histograms

p50, p95, p99 for TTFT and total, per model and per route. Spot regressions before your users do.

reliability

Error & retry rates

Track 4xx / 5xx / timeout / retry counts by upstream. Every failover is logged with the winning provider.

logs

Per-request logs

Full request/response payloads, redacted on ingest. Filter by status, cost, latency, model, or free-text search.

tags

Custom metadata

Attach x-llmcloud-tag headers or a metadata field — user IDs, feature flags, agent names — and slice by any of them.

export

Export & webhooks

CSV, Parquet, S3 sync, and a spend webhook per key. Snowflake, BigQuery and Datadog destinations on the roadmap.

How long are logs retained?+

30 days on the free tier, 90 days on Pro, configurable on Enterprise. Redaction rules apply on ingest so raw prompts never persist when guardrails flag them.

Can I get spend alerts?+

Yes — set daily and monthly thresholds per key. Alerts go to email, Slack or a webhook. Combine with per-key budgets to hard-stop at the cap.

Do you show cache hit rates?+

Yes. Every request is annotated with cache_status (miss, exact, semantic) so you can measure caching ROI per prompt shape.