Every provider is graded before it serves you.
Most gateways list whoever shows up. We publish the admission bar, run the probes ourselves, and date every score. A provider below the bar gets demoted out of default routing — not quietly kept in the mix.
The rubric
Five criteria, weighted. Serving integrity carries the most weight because it is the failure mode you cannot detect from your own logs.
Serving integrity
30%Are you served the weights you asked for?
- Advertised weights match served weights on a weekly canary probe
- Quantization disclosed per endpoint (bf16 / fp8 / int4)
- No silent model substitution during capacity crunches
- Deterministic seeds and logprobs behave as documented
- Context window honoured at the advertised limit, not truncated silently
Bar · Any undisclosed substitution or quantization change is an immediate suspension.
Performance under load
20%Does it hold up at 50x concurrency, not just at 1x?
- p50 / p95 TTFT measured at 1x, 10x, 50x concurrency
- Sustained output tokens/sec at each concurrency step
- Queue behaviour and backpressure semantics documented
- Streaming keeps flowing under load (no mid-stream stalls > 2s)
Bar · p95 TTFT at 10x may not exceed 3x the 1x baseline.
Reliability & capacity honesty
20%Do failures show up as errors, or as bad answers?
- 30 and 90-day uptime measured from our probes, not their status page
- Typed error taxonomy (rate limit vs overload vs upstream)
- Honest 429s instead of degraded quality during peaks
- Incident comms within 15 minutes, public postmortems
Bar · 99.5% 90-day availability measured at the gateway.
Data terms
20%What happens to the prompt after it leaves us?
- Zero-retention option available for gateway traffic
- No training on prompts or completions, contractually
- Sub-processors published and change-notified
- Region pinning enforceable, with a written residency commitment
Bar · No provider without a written no-training clause serves default traffic.
Commercial transparency
10%Can you predict next month's bill?
- Published per-token list price, no negotiated-only pricing
- Price-change notice period of 30 days or more
- Rate limits documented per tier and per model
- Billing granularity matches ours (per-token, not per-request buckets)
Bar · Undocumented rate limits or retroactive price changes fail the review.
Onboarding a provider
Terms, sub-processors, retention, training clauses, price-change notice and rate-limit docs. A provider that cannot show a written no-training clause never reaches step 2.
Two weeks of canary prompts against every endpoint, comparing served output against reference weights to detect quantization drift and silent substitution.
1x / 10x / 50x concurrency sweeps measuring TTFT, tok/s, stall rate and error taxonomy. Results are published, pass or fail.
30 days mirroring a slice of real production traffic with no user impact. Quality and cost deltas are compared against incumbents.
Graded, listed with a dated scorecard, and enrolled in monthly or quarterly re-evaluation. Scores below the bar demote the provider out of default routing.
Current scores
→ Provider directory| Provider | Grade | Score | Serving | Performance | Reliability | Data | Commercial | Evaluated |
|---|---|---|---|---|---|---|---|---|
| llmcloud OSS | A+ | 97 | 98 | 95 | 95 | 98 | 96 | 2026-07-28 |
| Meta (Llama) | A | 94 | 97 | 93 | 94 | 90 | 92 | 2026-07-07 |
| Anthropic | A | 94 | 98 | 86 | 96 | 94 | 90 | 2026-07-28 |
| Fireworks | A | 93 | 95 | 95 | 94 | 90 | 90 | 2026-07-07 |
| Mistral | A | 93 | 96 | 88 | 93 | 95 | 90 | 2026-07-14 |
| AWS Bedrock | A | 93 | 96 | 85 | 97 | 96 | 86 | 2026-07-14 |
| Vertex AI | A | 93 | 95 | 89 | 96 | 94 | 86 | 2026-07-14 |
| Together AI | A | 93 | 94 | 94 | 93 | 89 | 91 | 2026-07-07 |
| Azure OpenAI | A | 92 | 96 | 84 | 97 | 95 | 84 | 2026-07-14 |
| Cohere | A | 92 | 94 | 88 | 93 | 92 | 90 | 2026-06-30 |
| OpenAI | A | 92 | 97 | 84 | 95 | 88 | 92 | 2026-07-28 |
| Google (Gemini) | A | 92 | 95 | 92 | 94 | 86 | 88 | 2026-07-14 |
| SambaNova | A | 91 | 92 | 94 | 91 | 88 | 84 | 2026-06-30 |
| Groq | A | 91 | 93 | 97 | 88 | 86 | 84 | 2026-07-21 |
| Cerebras | B+ | 89 | 92 | 98 | 86 | 84 | 82 | 2026-07-21 |
| Perplexity | B+ | 85 | 88 | 86 | 89 | 78 | 82 | 2026-06-30 |
| Alibaba (Qwen) | B+ | 85 | 90 | 91 | 90 | 66 | 86 | 2026-07-21 |
| xAI | B | 84 | 90 | 82 | 88 | 74 | 78 | 2026-06-30 |
| DeepSeek | B | 82 | 88 | 90 | 86 | 62 | 84 | 2026-07-21 |
| Replicate | B | 82 | 86 | 74 | 85 | 80 | 82 | 2026-06-16 |
| RunPod | C+ | 78 | 80 | 82 | 78 | 72 | 80 | 2026-06-16 |
Scores are gateway-measured plus contract review. Re-evaluation is monthly for specialty and OSS hosts, quarterly for frontier labs and hyperscalers.