Every request, priced and traced.
Real-time spend, tokens, latency and error rates broken down by model, provider, key, project and end-user. No SDK to install — the router logs it all.
What you can slice
Cost per anything
Group by model, provider, region, key, project, end-user or a custom metadata tag you pass on the request.
Latency histograms
p50, p95, p99 for TTFT and total, per model and per route. Spot regressions before your users do.
Error & retry rates
Track 4xx / 5xx / timeout / retry counts by upstream. Every failover is logged with the winning provider.
Per-request logs
Full request/response payloads, redacted on ingest. Filter by status, cost, latency, model, or free-text search.
Custom metadata
Attach x-llmcloud-tag headers or a metadata field — user IDs, feature flags, agent names — and slice by any of them.
Export & webhooks
CSV, Parquet, S3 sync, and a spend webhook per key. Snowflake, BigQuery and Datadog destinations on the roadmap.
How long are logs retained?+
30 days on the free tier, 90 days on Pro, configurable on Enterprise. Redaction rules apply on ingest so raw prompts never persist when guardrails flag them.
Can I get spend alerts?+
Yes — set daily and monthly thresholds per key. Alerts go to email, Slack or a webhook. Combine with per-key budgets to hard-stop at the cap.
Do you show cache hit rates?+
Yes. Every request is annotated with cache_status (miss, exact, semantic) so you can measure caching ROI per prompt shape.