Every request, priced and traced.
Spend, tokens, latency and errors — sliced by model, provider, key, project or user. No SDK: the router logs it.
What you can slice
Cost per anything
Group by model, provider, region, key, project, user, or any custom metadata tag.
Latency histograms
p50/p95/p99 for TTFT and total, per model and route. Catch regressions early.
Error & retry rates
4xx / 5xx / timeout / retry counts per upstream. Every failover records the winner.
Per-request logs
Full payloads, redacted on ingest. Filter by status, cost, latency, model or text.
Custom metadata
Attach x-llmcloud-tag headers — user IDs, flags, agent names — and slice on them.
Export & webhooks
CSV, Parquet, S3 sync, and spend webhooks. Snowflake, BigQuery, Datadog on the roadmap.
How long are logs retained?+
30 days free, 90 days on Pro, configurable on Enterprise. Redaction runs on ingest, so flagged prompts never persist.
Can I get spend alerts?+
Yes — daily and monthly thresholds per key, delivered by email, Slack or webhook. Pair with per-key budgets for a hard stop.
Do you show cache hit rates?+
Yes. Every request carries cache_status (miss, exact, semantic) so you can measure caching ROI.