Live leaderboard
Rankings that match your workload.
Public benchmarks (MMLU-Pro, GPQA, LiveBench, SWE-Bench) blended with our task-specific evals and real gateway traffic. Updated continuously.
Live·4,700 requests/sec across 15 tracked models
#
Model
Modality
01
claude-opus-4.5
Anthropic
textvision
94.2
82
320
$15.00↑
02
gpt-5.5
OpenAI
textvisionaudio
93.8
74
410
$12.50↑
03
gemini-3.1-pro
Google
textvisionaudiovideo
92.1
118
180
$7.00↑
04
grok-5
xAI
textvision
90.6
95
260
$6.00·
05
deepseek-v4
DeepSeek
text
89.9
140
210
$0.55↑
06
qwen3-max
Alibaba
textvision
88.4
128
240
$0.90↑
07
deepseek-coder-v3
DeepSeek
text
88.2
155
200
$0.60↑
08
llama-4-405b
Meta
textvision
87.6
165
150
$2.40·
09
mistral-large-3
Mistral
text
86.1
110
280
$3.00↓
10
gpt-5.5-mini
OpenAI
textvision
85.3
210
120
$0.40↑
11
gemini-3.6-flash
Google
textvisionaudio
84.7
240
90
$0.30↑
12
claude-haiku-4.5
Anthropic
textvision
83.9
195
130
$0.80↑
13
qwen3-vl-72b
Alibaba
textvision
82.4
88
340
$1.20·
14
llama-4-70b
Meta
text
81.1
220
110
$0.90·
15
grok-5-fast
xAI
text
79.8
280
80
$0.50↑
Preview data · Quality = weighted average of MMLU-Pro, GPQA-Diamond, LiveBench, SWE-Bench-Verified, and our proprietary router-eval.