AI Model Rankings

Ranked by an AI Value score out of 100 — your chosen benchmarks, at your chosen weights, normalised against every model we track. Pick a preset or set the weights yourself; the ranking recalculates as you go.

Data through .Every figure carries its source: 21% refresh automatically from a live API, the rest are published results we cite but have not re-measured ourselves. How each benchmark works

Metrics & Weights

6 metrics selected — total weight: 100%

Preset

Active metrics

SWE-bench Verified 30%HumanEval+ 15%Chatbot Arena ELO 15%Output Speed 15%Input Cost 13%Output Cost 13%
#Model
1
Gemini 3 Flash
Google
82.6
2
GPT-5.2
OpenAI
81.2
3
Gemini 3.1 Pro
Google
80.6
4
Claude Opus 4.6
Anthropic
80.2
5
GPT-5.1
OpenAI
79.1
6
Claude Opus 4.5
Anthropic
77.8
7
GPT-5.1 Codex
OpenAI
77.7
8
GPT-5
OpenAI
75.7
9
Claude Sonnet 4.6
Anthropic
74.9
10
Gemini 3 Pro
Google
73.7
11
Grok 4.1 Fast
xAI
73.7
12
Claude Sonnet 4.5
Anthropic
73.6
13
Grok 4
xAI
73.1
14
Gemini 2.5 Pro
Google
71.4
15
DeepSeek V3.2
DeepSeek
71.3
16
o4 Mini
OpenAI
71.3
17
Claude Sonnet 4
Anthropic
71
18
Qwen 3.5 397B
Qwen
70.8
19
o3
OpenAI
69.7
20
Claude Opus 4
Anthropic
68.7
21
o3 Pro
OpenAI
68.2
22
GPT-5.1 Codex Mini
OpenAI
67.6
23
Qwen 3 Coder
Qwen
67.5
24
Claude 3.7 Sonnet
Anthropic
65.5
25
GPT-5 Mini
OpenAI
64.2
26
DeepSeek R1 0528
DeepSeek
64
27
Grok 4 Fast
xAI
63.6
28
Gemini 2.5 Flash
Google
63.6
29
o3 Mini
OpenAI
62.5
30
GPT-4.1
OpenAI
62.4
31
Mistral Large 25.12
Mistral
60.3
32
DeepSeek R1
DeepSeek
60.2
33
o1 Mini
OpenAI
59.4
34
Qwen 3 Max
Qwen
59.4
35
DeepSeek V3.1
DeepSeek
59.3
36
Grok 3
xAI
59
37
o1
OpenAI
59
38
Qwen 3 235B
Qwen
56.8
39
Claude 3.5 Sonnet
Anthropic
56.4
40
GPT-4o
OpenAI
56
41
Codestral
Mistral
56
42
Claude Fable 5
Anthropic
55
43
GPT-4.1 Mini
OpenAI
54.9
44
Gemini 2.0 Flash
Google
53.4
45
Llama 4 Maverick
Meta
53.3
46
Claude Haiku 4.5
Anthropic
52.1
47
Grok 3 Mini
xAI
51.4
48
Qwen 3 32B
Qwen
50.9
49
GPT-5 Nano
OpenAI
50.8
50
Pixtral Large
Mistral
50.7
51
Gemini 2.5 Flash Lite
Google
49.7
52
GPT-4o Mini
OpenAI
49.6
53
Llama 4 Scout
Meta
49.3
54
Mistral Medium 3.1
Mistral
49
55
Command A
Cohere
48.5
56
Qwen 2.5 72B
Qwen
47.9
57
Grok 2
xAI
47.6
58
GPT-4.5
OpenAI
47.2
59
Claude 3.5 Haiku
Anthropic
45.8
60
Mistral Small 3.2
Mistral
43.7
61
Llama 3.3 70B
Meta
42.6
62
GPT-4.1 Nano
OpenAI
40.3
63
Claude 3 Opus
Anthropic
40.1
64
GPT-OSS 120B
OpenAI
40
65
GLM-5
Zhipu AI
39.9
66
Command R+
Cohere
39.7
67
MiniMax M2.5
MiniMax
39.6
68
Nova Pro
Amazon
38.9
69
GPT-OSS 20B
OpenAI
38.7
70
Gemma 3 27B
Google
35.6
71
Command R
Cohere
32.3
72
Ministral 3 8B
Mistral
32.2
73
Phi-4 Mini
Microsoft
31.1
74
Nova Lite
Amazon
31.1
75
Reka Flash 3
Reka AI
30.5
76
Phi-4
Microsoft
28.5
77
Gemma 3 12B
Google
28.3
78
GLM-4.6V
Zhipu AI
27.3
79
Kimi K2.5
Moonshot AI
26.1
80
Sonar Pro
Perplexity
25.9
81
Jamba 1.5 Mini
AI21 Labs
25.7
82
Gemma 3 4B
Google
25
83
Yi Lightning
01.AI
25
84
Phi-4 Reasoning Plus
Microsoft
25
85
GPT-5.6 Luna
OpenAI
25
86
GPT-5.4 nano
OpenAI
25
87
Gemini Embedding 001
Google
25
88
text-embedding-3-small
OpenAI
25
89
text-embedding-3-large
OpenAI
25
90
Lyria 3 Clip Preview
Google
25
91
Qwen3-ASR Flash
Alibaba
25
92
Qwen3.6 35B-A3B
Alibaba
25
93
Qwen Plus
Alibaba
25
94
Qwen3.6 Flash
Alibaba
25
95
Qwen3.8 Flash
Alibaba
25
96
Qwen3 8B
Alibaba
25
97
Gemini Embedding 2
Google
25
98
Lyria 3 Pro Preview
Google
25
99
DeepSeek V4 Flash Vision Exp
DeepSeek
25
100
DeepSeek V4 Flash
DeepSeek
25
101
DeepSeek V4.1 Flash
DeepSeek
25
102
Muse Spark 1.2 Contributor
Meta
25
103
Muse Spark 1.3 Contributor
Meta
25
104
Pixtral 12B
Mistral
25
105
Devstral Small
Mistral
25
106
Mistral Embed
Mistral
25
107
Devstral Small 2505
Mistral
25
108
Devstral Small 2
Mistral
25
109
Mistral 7B
Mistral
25
110
Voxtral Small (latest)
Mistral
25
111
Ministral 8B (latest)
Mistral
25
112
Mistral Nemo
Mistral
25
113
Mistral Small (latest)
Mistral
25
114
Open Mistral Nemo
Mistral
25
115
Mistral Small 4
Mistral
25
116
Ministral 3B (latest)
Mistral
25
117
Qwen2.5 72B Instruct
Alibaba
25
118
DeepSeek V4 Flash 0731
Alibaba
25
119
Qwen2.5-Omni 7B
Alibaba
25
120
Qwen-MT Turbo
Alibaba
25
121
Qwen3-Next 80B-A3B Instruct
Alibaba
25
122
Qwen3 Coder Flash
Alibaba
25
123
Qwen-VL Plus
Alibaba
25
124
Qwen3.5 27B
Alibaba
25
125
Qwen Flash
Alibaba
25
126
Qwen Turbo
Alibaba
25
127
Qwen3-VL 30B-A3B
Alibaba
25
128
Qwen3-VL Plus
Alibaba
25
129
Qwen2.5 7B Instruct
Alibaba
25
130
Qwen3-Coder 30B-A3B Instruct
Alibaba
25
131
Gemma 2 2b It
NVIDIA
25
132
Qwen3 14B
Alibaba
25
133
GLM-5.3-Flash
Zhipu AI
25
134
GLM-4.7-Flash
Zhipu AI
25
135
GLM-4.7-FlashX
Zhipu AI
25
136
GLM-4.5-Flash
Zhipu AI
25
137
Qwen-Omni Turbo
Alibaba
25
138
Command R7B Arabic
Cohere
25
139
North Mini Code
Cohere
25
140
Command R7B
Cohere
25
141
MiniMax-M2
MiniMax
25
142
GLM-4.5-Air
Zhipu AI
25
143
MiniMax-M3
MiniMax
25
144
MiniMax-M2.7
MiniMax
25
145
Step 3.5 Flash
StepFun
25
146
Step 3.5 Flash 2603
StepFun
25
147
Qwen Image
NVIDIA
25
148
Qwen Image Edit
NVIDIA
25
149
Laguna XS 2.1
NVIDIA
25
150
Mistral Large 3 675B Instruct 2512
NVIDIA
25
151
Ministral 3 14B Instruct 2512
NVIDIA
25
152
mistral-small-4-119b-2603
NVIDIA
25
153
Mistral: Mixtral 8x7B Instruct
NVIDIA
25
154
Mistral-7B-Instruct-v0.3
NVIDIA
25
155
Magistral Small 2506
NVIDIA
25
156
streampetr
NVIDIA
25
157
Llama 3.3 Nemotron Super 49B v1.5
NVIDIA
25
158
usdcode
NVIDIA
25
159
Nemotron 3 Nano Omni
NVIDIA
25
160
cosmos-transfer1-7b
NVIDIA
25
161
nemotron-voicechat
NVIDIA
25
162
studiovoice
NVIDIA
25
163
nemotron-3-content-safety
NVIDIA
25
164
cosmos-transfer2.5-2b
NVIDIA
25
165
bevformer
NVIDIA
25
166
Llama 3.1 Nemotron Nano VL 8B v1
NVIDIA
25
167
llama-nemotron-embed-vl-1b-v2
NVIDIA
25
168
synthetic-video-detector
NVIDIA
25
169
Nemotron 3 Super
NVIDIA
25
170
llama-nemotron-rerank-vl-1b-v2
NVIDIA
25
171
usdvalidate
NVIDIA
25
172
Active Speaker Detection
NVIDIA
25
173
Llama 3.1 Nemotron Ultra 253B
NVIDIA
25
174
llama-3_2-nemoretriever-300m-embed-v1
NVIDIA
25
175
nv-embedcode-7b-v1
NVIDIA
25
176
llama-3.1-nemotron-safety-guard-8b-v3
NVIDIA
25
177
nemotron-mini-4b-instruct
NVIDIA
25
178
cosmos-predict1-5b
NVIDIA
25
179
nemotron-content-safety-reasoning-4b
NVIDIA
25
180
riva-translate-4b-instruct-v1_1
NVIDIA
25
181
BGE M3
NVIDIA
25
182
sparsedrive
NVIDIA
25
183
gliner-pii
NVIDIA
25
184
Llama 3.1 Nemotron Nano 8B v1
NVIDIA
25
185
nv-embed-v1
NVIDIA
25
186
rerank-qa-mistral-4b
NVIDIA
25
187
Nemotron 3.5 Lightning 30B A3B
NVIDIA
25
188
nemotron-3-nano-30b-a3b
NVIDIA
25
189
paligemma
NVIDIA
25
190
Ministral 3 3B
Amazon Bedrock
25
191
nvidia-nemotron-nano-9b-v2
NVIDIA
25
192
Cosmos Reason2 8B
NVIDIA
25
193
Gemma 3 4B IT
NVIDIA
25
194
Gemma 3n E2b It
NVIDIA
25
195
Gemma 3n E4b It
NVIDIA
25
196
Llama 3.1 8B Instruct
NVIDIA
25
197
Llama Guard 4 12B
NVIDIA
25
198
Llama-3.2-90B-Vision-Instruct
NVIDIA
25
199
Llama 3.2 1b Instruct
NVIDIA
25
200
esmfold
NVIDIA
25
201
Llama 3.2 11b Vision Instruct
NVIDIA
25
202
Llama 3.1 70b Instruct
NVIDIA
25
203
Llama 3.3 70b Instruct
NVIDIA
25
204
ByteDance-Seed/Seed-OSS-36B-Instruct
NVIDIA
25
205
sarvam-m
NVIDIA
25
206
Phi 4 Multimodal
NVIDIA
25
207
Llama 3.1 Nemotron 70B Instruct
NVIDIA
25
208
dracarys-llama-3.1-70b-instruct
NVIDIA
25
209
Whisper Large v3
NVIDIA
25
210
solar-10.7b-instruct
NVIDIA
25
211
FLUX.1-schnell
NVIDIA
25
212
Gemma 3 12B IT
NVIDIA
25
213
Llama 3.2 3B Instruct
NVIDIA
25
214
Llama 4 Maverick 17b 128e Instruct
NVIDIA
25
215
esm2-650m
NVIDIA
25
216
FLUX.2 Klein 4B
NVIDIA
25
217
FLUX.1-dev
NVIDIA
25
218
FLUX.1-Kontext-dev
NVIDIA
25
219
Mercury 2.5
Inception
25
220
Mercury Edit 2
Inception
25
221
Mercury 2
Inception
25
222
solar-mini
Upstage
25
223
Solar Pro 4
Upstage
25
224
solar-pro3
Upstage
25
225
solar-pro2
Upstage
25
226
Qwen3 235B A22B Instruct
Vertex
25
227
Qwen3 Coder Next
Amazon Bedrock
25
228
Qwen3 235B-A22B Instruct 2507
Amazon Bedrock
25
229
Voxtral Mini 3B 2507
Amazon Bedrock
25
230
NVIDIA Nemotron Nano 12B v2 VL BF16
Amazon Bedrock
25
231
NVIDIA Nemotron Nano 3 30B
Amazon Bedrock
25
232
Gemma 4 E2B IT
Amazon Bedrock
25
233
Llama 4 Maverick 17B Instruct (US)
Amazon Bedrock
25
234
Ministral 14B 3.0
Amazon Bedrock
25
235
GPT OSS Safeguard 20B
Amazon Bedrock
25
236
Llama 4 Scout 17B Instruct (US)
Amazon Bedrock
25
237
NVIDIA Nemotron 3 Super 120B A12B
Amazon Bedrock
25
238
GPT OSS Safeguard 120B
Amazon Bedrock
25
239
Voxtral Small 24B 2507
Amazon Bedrock
25
240
Mistral Small 3.1
Azure
25
241
Phi-4-mini-reasoning
Azure
25
242
Llama 4 Maverick 17B 128E Instruct FP8
Azure
25
243
Embed v3 English
Azure
25
244
Embed v4
Azure
25
245
Gemma 3 27B IT
Amazon Bedrock
25
246
Llama 4 Scout 17B 16E Instruct
Azure
25
247
Model Router
Azure
25
248
Embed v3 Multilingual
Azure
25
249
Phi-4-reasoning
Azure
25
250
Codestral 25.01
Azure
25
251
text-embedding-ada-002
OpenAI
25
252
Qwen-Omni Turbo Realtime
Alibaba
25
253
MiniMax-M2.1
MiniMax
25
254
Step 3.7 Flash
StepFun
25
255
mistral-nemotron
NVIDIA
25
256
magpie-tts-zeroshot
NVIDIA
25
257
Nemotron Nano 12B v2 VL
NVIDIA
25
258
Llama 3.3 Nemotron Super 49B v1
NVIDIA
25
259
Nova Micro (US)
Amazon Bedrock
25
260
Sonar
Perplexity
24.9
261
Gemini 3.6 Flash
Google
24.9
262
Gemini 3.1 Flash Lite
Google
24.9
263
Nano Banana 2
Google
24.9
264
GPT-3.5-turbo
OpenAI
24.9
265
Nano Banana 2 Lite
Google
24.9
266
Qwen3-Omni Flash
Alibaba
24.9
267
Qwen2.5-VL 72B Instruct
Alibaba
24.9
268
Qwen3.5 122B-A10B
Alibaba
24.9
269
Qwen3-Omni Flash Realtime
Alibaba
24.9
270
Qwen2.5-VL 7B Instruct
Alibaba
24.9
271
Qwen2.5 14B Instruct
Alibaba
24.9
272
Qwen Plus Character (Japanese)
Alibaba
24.9
273
Qwen3 235B-A22B
Alibaba
24.9
274
Qwen3.6 27B
Alibaba
24.9
275
Gemini Flash-Lite Latest
Google
24.9
276
Gemini 3.8 Flash
Google
24.9
277
Gemini 3.7 Flash
Google
24.9
278
Gemini Flash Latest
Google
24.9
279
DeepSeek V4 Pro
DeepSeek
24.9
280
Magistral Small
Mistral
24.9
281
Devstral 2
Mistral
24.9
282
Mixtral 8x7B
Mistral
24.9
283
GLM-5.2
Mistral
24.9
284
Mistral Medium 3
Mistral
24.9
285
Mistral Large 3
Mistral
24.9
286
Grok Build 0.1
xAI
24.9
287
Qwen3 Coder Plus
Alibaba
24.9
288
Qwen-VL Max
Alibaba
24.9
289
Qwen3.6 Plus
Alibaba
24.9
290
Qwen3.5 35B-A3B
Alibaba
24.9
291
Qwen-VL OCR
Alibaba
24.9
292
QwQ Plus
Alibaba
24.9
293
Qwen3.5 397B-A17B
Alibaba
24.9
294
Qwen2.5 32B Instruct
Alibaba
24.9
295
Qwen3-VL 235B-A22B
Alibaba
24.9
296
Qwen3.5 Plus
Alibaba
24.9
297
GLM-5.1
Zhipu AI
24.9
298
GLM-4.5V
Zhipu AI
24.9
299
GLM-4.5
Zhipu AI
24.9
300
GLM-4.6
Zhipu AI
24.9
301
GLM-4.7
Zhipu AI
24.9
302
Kimi K2.7 Code
Moonshot AI
24.9
303
MiniMax-M2.7-highspeed
MiniMax
24.9
304
Qwen2.5 Coder 32b Instruct
NVIDIA
24.9
305
DeepSeek V4 Pro 0813
NVIDIA
24.9
306
MiniMax-M2.5-highspeed
MiniMax
24.9
307
Nemotron 3 Ultra 550B A55B
NVIDIA
24.9
308
Kimi K2 Thinking
Vertex
24.9
309
Morph v3 Large
Morph
24.9
310
Morph v3 Fast
Morph
24.9
311
Auto
Morph
24.9
312
Kimi K2 0905
NVIDIA
24.9
313
Nova 2 Lite (EU)
Amazon Bedrock
24.9
314
Qwen3 VL 235B A22B Instruct
Amazon Bedrock
24.9
315
Magistral Small 1.2
Amazon Bedrock
24.9
316
Devstral 2 123B
Amazon Bedrock
24.9
317
GPT-3.5 Turbo 1106
Azure
24.9
318
GPT-3.5 Turbo 0125
Azure
24.9
319
DeepSeek-V3.2-Speciale
Azure
24.9
320
Nano Banana
Google
24.9
321
Gemini 3.5 Flash Lite
Google
24.9
322
Devstral Medium
Mistral
24.9
323
Qwen3.7 Plus
Alibaba
24.9
324
Muse Glimmer 30B
NVIDIA
24.9
325
GPT-5.4 mini
OpenAI
24.8
326
Gemini 3.1 Flash Live Preview
Google
24.8
327
Qwen3.6 Max Preview
Alibaba
24.8
328
Muse Spark 1.1
Meta
24.8
329
Muse Spark 1.2
Meta
24.8
330
Grok 4.3
xAI
24.8
331
Grok 4.20 (Reasoning)
xAI
24.8
332
Grok 4.20 Multi-Agent
xAI
24.8
333
Qwen3.7 Max
Alibaba
24.8
334
Qwen3-Next 80B-A3B (Thinking)
Alibaba
24.8
335
QVQ Max
Alibaba
24.8
336
GLM-5.3
Zhipu AI
24.8
337
GLM-5V-Turbo
Zhipu AI
24.8
338
Kimi K2.6
Moonshot AI
24.8
339
Inkling
NVIDIA
24.8
340
Sonar Reasoning Pro
Perplexity
24.8
341
Gemini 2.5 Flash TTS
Vertex
24.8
342
Palmyra X5 (US)
Amazon Bedrock
24.8
343
Codex Mini
Azure
24.8
344
GPT-3.5 Turbo Instruct
Azure
24.8
345
Muse Spark 1.3
Meta
24.8
346
Jamba 1.5 Large
AI21 Labs
24.7
347
Gemini 2.5 Computer Use Preview 10-2025
Google
24.7
348
Gemini 3.5 Flash
Google
24.7
349
Qwen3.8 Max
Alibaba
24.7
350
Magistral Medium (latest)
Mistral
24.7
351
Mixtral 8x22B
Mistral
24.7
352
Mistral Medium (latest)
Mistral
24.7
353
Mistral Large (latest)
Mistral
24.7
354
Mistral Large 2.1
Mistral
24.7
355
Grok 4.5
xAI
24.7
356
Grok 4.6
xAI
24.7
357
Qwen Max
Alibaba
24.7
358
Qwen3-Coder 480B-A35B Instruct
Alibaba
24.7
359
Mistral: Mixtral 8x22B Instruct
NVIDIA
24.7
360
Perplexity Sonar Deep Research
Perplexity
24.7
361
GPT-5-Codex
Azure
24.7
362
GPT-5.1 Codex Max
Azure
24.7
363
Mistral Medium 3.5
Mistral
24.7
364
Kimi K2.7 Code HighSpeed
Moonshot AI
24.7
365
Claude Sonnet 5
Anthropic
24.6
366
GPT-5.6 Sol
OpenAI
24.6
367
GPT-5.3 Codex
OpenAI
24.6
368
Deep Research Max Preview (Apr-21-2026)
Google
24.6
369
Nano Banana Pro
Google
24.6
370
Deep Research Preview (Apr-21-2026)
Google
24.6
371
GPT-5.6 Terra
OpenAI
24.6
372
GPT-5.2 Chat
OpenAI
24.6
373
GPT-5.3 Chat (latest)
OpenAI
24.6
374
Qwen-MT Plus
Alibaba
24.6
375
Command A Plus
Cohere
24.6
376
Command A Reasoning
Cohere
24.6
377
Command A Vision
Cohere
24.6
378
Command A Translate
Cohere
24.6
379
Step 1 (32K)
StepFun
24.6
380
Palmyra X4 (US)
Amazon Bedrock
24.6
381
GPT-5.2 Codex
Azure
24.6
382
GPT-5.3 Codex Spark
OpenAI
24.6
383
GPT-5.4
OpenAI
24.5
384
Gemini 3.1 Flash TTS Preview
Google
24.5
385
Gemini Omni Flash Preview
Google
24.5
386
Kimi K3
Moonshot AI
24.5
387
Gemini 2.5 Pro TTS
Vertex
24.5
388
Nova Premier (US)
Amazon Bedrock
24.5
389
AU Anthropic Claude Sonnet 4.6
Amazon Bedrock
24.4
390
GPT-5.6
OpenAI
24.3
391
Gemini 3.5 Live Translate Preview
Google
24.3
392
GPT-Realtime-2.1
OpenAI
24.2
393
Step 2 (16K)
StepFun
24.2
394
Claude Opus 4.7
Anthropic
24.1
395
Claude Opus 4.8
Anthropic
24.1
396
Claude Opus 5
Anthropic
24.1
397
GPT-5.5
OpenAI
24
398
Qwen3-LiveTranslate Flash Realtime
Alibaba
24
399
GPT Chat Latest
Azure
24
400
gpt-image-2
OpenAI
24
401
GPT-4 Turbo
OpenAI
23.5
402
GPT-4 Turbo Vision
Azure
23.5
403
Claude Fable 5.1
Anthropic
23.1
404
GPT-6 Astra
OpenAI
23.1
405
Claude Mythos 5
Azure
23.1
406
Claude Opus 4.1
Vertex
22.2
407
AU Anthropic Claude Opus 4.6
Amazon Bedrock
21.9
408
GPT-5 Pro
OpenAI
21.3
409
GPT-4
OpenAI
21.3
410
GPT-5.2 Pro
OpenAI
19.8
411
GPT-5.5 Pro
OpenAI
18.8
412
GPT-5.4 Pro
OpenAI
18.8
413
Doubao Seed 2.0
ByteDance
14
414
Veo 3.1
Google
11.2
415
Sora 2
OpenAI
9.9
416
Nova 2.0 Lite
Amazon
9.5
417
Veo 3
Google
8.8
418
Seedream 4.5
ByteDance
7.9
419
MiMo V2 Flash
Xiaomi
6.4
420
Qwen 3 Next 80B
Qwen
5.6
421
Ministral 3 14B
Mistral
3
422
Ring Flash 2.0
InclusionAI
3
423
Nemotron 3 Nano
NVIDIA
2.7
424
Step 2.5 Flash
StepFun
2.3
425
K-EXAONE
LG AI Research
2
426
Qwen 3 Coder 480B
Qwen
1.9
427
Qwen 3 VL 235B
Qwen
1.3
428
Magistral Medium 1.2
Mistral
0.5
429
DALL-E 3
OpenAI
0
430
Kling 2.5 Turbo
Kuaishou
0
431
ERNIE 4.5
Baidu
0
432
ERNIE X1
Baidu
0
433
Hermes 4 70B
Nous Research
0
434
Imagen 4
Google
0
435
Luma Ray 3
Luma AI
0
436
Runway Gen-4
Runway
0
437
Midjourney v7
Midjourney
0
438
Midjourney v6.1
Midjourney
0
439
Stable Diffusion 3.5
Stability AI
0
440
Flux 1.1 Pro
Black Forest Labs
0
441
Flux 1.0 Dev
Black Forest Labs
0
442
Ideogram 3.0
Ideogram
0
443
Pika 2.5
Pika
0
444
o1-pro
OpenAI
0
445
chatgpt-image-latest
OpenAI
0
446
gpt-image-1
OpenAI
0
447
gpt-image-1-mini
OpenAI
0
448
gpt-image-1.5
OpenAI
0
449
Veo 3.1 fast
Google
0
450
Gemma 4 26B A4B IT
Google
0
451
Veo 3.1 lite
Google
0
452
Gemma 4 31B IT
Google
0
453
Voxtral Mini TTS (latest)
Mistral
0
454
Voxtral Mini (latest)
Mistral
0
455
Grok Imagine Video 1.5
xAI
0
456
Grok Imagine Image 2.0
xAI
0
457
Grok Imagine Image Quality
xAI
0
458
Grok Imagine Video
xAI
0
459
Aya Expanse 32B
Cohere
0
460
Aya Expanse 8B
Cohere
0
461
Aya Vision 8B
Cohere
0
462
Aya Vision 32B
Cohere
0
463
Step TTS 2
StepFun
0
464
StepAudio 2.5 ASR
StepFun
0
465
StepAudio 2.5 TTS
StepFun
0
466
Grok Imagine Image
xAI
0

Data sourced from Chatbot Arena, OpenRouter, and public benchmarks. Updated daily. Scores are dynamically computed based on your selected metrics.

How It Works

1

Choose a Persona or Build Your Own

Select a preset (Developer, Researcher, Business, etc.) that pre-selects relevant benchmarks and weights. Or customize everything from scratch.

2

See Real Data, Not Abstract Scores

Every metric shows the actual value — ELO 1410, $2.50/1M tokens, 65 tok/s. No normalization black box. The raw numbers are always visible.

3

Dynamic Ranking — Your Weights, Your Score

For each metric you select, we find the min/max across all models, normalize to 0-100, then compute a weighted composite. Missing data is excluded and weights renormalize automatically.

4

Share Your Rankings

Your exact configuration is encoded in the URL. Share it with your team or embed it — they'll see your exact ranking.

Data sourced from Chatbot Arena, OpenRouter, SWE-bench, and public research papers.

45 benchmarks across General, Coding, Math, Reasoning, Speed, Cost, and Context categories.

Frequently Asked Questions

What is the best AI model in 2026?

The best AI model depends on your use case. As of 2026, top contenders include Gemini 3.1 Pro, Claude Opus 4.6, and GPT-5.2 for general intelligence. For coding, Claude Opus 4.6 and GPT-5.2 lead on SWE-bench. For budget-conscious users, DeepSeek V3.2 and Gemini 2.5 Flash offer excellent performance per dollar. Use our rankings to sort models by YOUR priorities.

How do AI benchmarks work?

AI benchmarks are standardized tests that evaluate language models across specific capabilities. Common benchmarks include Chatbot Arena ELO (human preference voting), SWE-bench (real software engineering tasks), MMLU-Pro (knowledge and reasoning), GPQA Diamond (graduate-level science), and MATH (competition mathematics). Each benchmark tests a different aspect of model capability, and no single benchmark tells the whole story.

What is Chatbot Arena ELO?

Chatbot Arena ELO is a human preference ranking system where users compare AI model responses in blind head-to-head matchups. The ELO rating (borrowed from chess) reflects how often a model is preferred over others. Higher ELO means the model is more frequently preferred. It's considered one of the most reliable benchmarks because it uses real human judgment rather than automated scoring.

Which AI model is best for coding?

For coding tasks in 2026, the top models are Claude Opus 4.6 (80.8% SWE-bench), GPT-5.2 (80.0% SWE-bench), and Gemini 3.1 Pro (80.6% SWE-bench). For more affordable coding, GPT-5.1 Codex and Qwen 3 Coder offer strong performance at lower costs. The Developer persona pre-weights coding benchmarks to help you find the best fit.

Which AI model is cheapest?

The cheapest AI models by API pricing include GPT-5 Nano ($0.05/1M input), GPT-4.1 Nano ($0.10/1M input), Nova Lite ($0.06/1M input), and Gemini 2.5 Flash Lite ($0.10/1M input). For the best balance of quality and cost, DeepSeek V3.2 ($0.14/1M input) and Qwen 3 32B ($0.10/1M input) offer strong performance at budget prices.