llmcloud.ai
Provider standards

Every provider is graded before it serves you.

Most gateways list whoever shows up. We publish the admission bar, run the probes ourselves, and date every score. A provider below the bar gets demoted out of default routing — not quietly kept in the mix.

The rubric

Five criteria, weighted. Serving integrity carries the most weight because it is the failure mode you cannot detect from your own logs.

Serving integrity

30%

Are you served the weights you asked for?

  • Advertised weights match served weights on a weekly canary probe
  • Quantization disclosed per endpoint (bf16 / fp8 / int4)
  • No silent model substitution during capacity crunches
  • Deterministic seeds and logprobs behave as documented
  • Context window honoured at the advertised limit, not truncated silently

Bar · Any undisclosed substitution or quantization change is an immediate suspension.

Performance under load

20%

Does it hold up at 50x concurrency, not just at 1x?

  • p50 / p95 TTFT measured at 1x, 10x, 50x concurrency
  • Sustained output tokens/sec at each concurrency step
  • Queue behaviour and backpressure semantics documented
  • Streaming keeps flowing under load (no mid-stream stalls > 2s)

Bar · p95 TTFT at 10x may not exceed 3x the 1x baseline.

Reliability & capacity honesty

20%

Do failures show up as errors, or as bad answers?

  • 30 and 90-day uptime measured from our probes, not their status page
  • Typed error taxonomy (rate limit vs overload vs upstream)
  • Honest 429s instead of degraded quality during peaks
  • Incident comms within 15 minutes, public postmortems

Bar · 99.5% 90-day availability measured at the gateway.

Data terms

20%

What happens to the prompt after it leaves us?

  • Zero-retention option available for gateway traffic
  • No training on prompts or completions, contractually
  • Sub-processors published and change-notified
  • Region pinning enforceable, with a written residency commitment

Bar · No provider without a written no-training clause serves default traffic.

Commercial transparency

10%

Can you predict next month's bill?

  • Published per-token list price, no negotiated-only pricing
  • Price-change notice period of 30 days or more
  • Rate limits documented per tier and per model
  • Billing granularity matches ours (per-token, not per-request buckets)

Bar · Undocumented rate limits or retroactive price changes fail the review.

Onboarding a provider

1 · Paper review

Terms, sub-processors, retention, training clauses, price-change notice and rate-limit docs. A provider that cannot show a written no-training clause never reaches step 2.

2 · Integrity probe

Two weeks of canary prompts against every endpoint, comparing served output against reference weights to detect quantization drift and silent substitution.

3 · Load harness

1x / 10x / 50x concurrency sweeps measuring TTFT, tok/s, stall rate and error taxonomy. Results are published, pass or fail.

4 · Shadow traffic

30 days mirroring a slice of real production traffic with no user impact. Quality and cost deltas are compared against incumbents.

5 · Admission

Graded, listed with a dated scorecard, and enrolled in monthly or quarterly re-evaluation. Scores below the bar demote the provider out of default routing.

ProviderGradeScoreServingPerformanceReliabilityDataCommercialEvaluated
llmcloud OSSA+
97
98959598962026-07-28
Meta (Llama)A
94
97939490922026-07-07
AnthropicA
94
98869694902026-07-28
FireworksA
93
95959490902026-07-07
MistralA
93
96889395902026-07-14
AWS BedrockA
93
96859796862026-07-14
Vertex AIA
93
95899694862026-07-14
Together AIA
93
94949389912026-07-07
Azure OpenAIA
92
96849795842026-07-14
CohereA
92
94889392902026-06-30
OpenAIA
92
97849588922026-07-28
Google (Gemini)A
92
95929486882026-07-14
SambaNovaA
91
92949188842026-06-30
GroqA
91
93978886842026-07-21
CerebrasB+
89
92988684822026-07-21
PerplexityB+
85
88868978822026-06-30
Alibaba (Qwen)B+
85
90919066862026-07-21
xAIB
84
90828874782026-06-30
DeepSeekB
82
88908662842026-07-21
ReplicateB
82
86748580822026-06-16
RunPodC+
78
80827872802026-06-16
A+95.0/100

Scores are gateway-measured plus contract review. Re-evaluation is monthly for specialty and OSS hosts, quarterly for frontier labs and hyperscalers.