Chatbot Arena ELO
Human preference ELO from blind head-to-head votes
What each metric we track measures and why it matters. We track 45 benchmarks across 8 categories.
Broad knowledge and reasoning benchmarks that test general intelligence across many domains.
Human preference ELO from blind head-to-head votes
Massive Multitask Language Understanding professional benchmark
Instruction following quality scored by GPT-4 as judge
Multi-turn conversation quality on 80 curated dialogues
Strict instruction following accuracy on verifiable constraints
Factual accuracy on short, verifiable questions
Resistance to generating false but plausible answers
The Intelligence Index is a single headline number that Artificial Analysis calculates by running a model through ten separate tests and blending the results. Those tests are grouped into four areas: agentic work (can the model carry out a multi-step task without a human steering it), coding, general knowledge and reasoning, and scientific reasoning.
Each underlying score is rescaled before blending so that one test with a wide range cannot dominate the total. Artificial Analysis runs these tests themselves rather than accepting numbers supplied by the model's maker, and they repeat several of the evaluations more than ten times per model, which lets them publish a confidence interval of under one percent.
What you can take from it: treat it as a general-purpose shortlist filter rather than a verdict. A model three or four points above another is meaningfully stronger across the board; a gap of under a point is noise. Because it averages very different skills, a specialist model can be held back by the categories you do not care about — if you only need code, read the Coding Index beside it, and if the work involves tools and multiple steps, read the Agentic Index.
Metrics that evaluate code generation, understanding, and real-world software engineering capabilities.
Code generation correctness with extended tests
Real-world software engineering task resolution
Live competitive programming benchmark
Multi-language code editing accuracy with real git repos
Comprehensive code generation across diverse programming tasks
Berkeley Function Calling Leaderboard — tool use accuracy
The Coding Index is the programming slice of Artificial Analysis's testing, reported on its own so it is not diluted by a model's general knowledge. It draws on tasks that ask a model to work inside a realistic codebase and on terminal-style benchmarks where the model has to run commands, read the output and react to it, rather than write a single function from a description.
What you can take from it: this is the number to weigh most heavily if the model will sit behind a coding assistant, a code-review step or an automated build fix. It tracks real-world usefulness better than a pure code-completion score does, because the underlying tasks penalise a model that writes plausible code but cannot recover when the code fails.
Read it next to the Agentic Index — strong coding with weak agentic scores usually means a model that writes well but needs a human to drive it.
Design Arena runs head-to-head votes on what models actually build: a website, a UI component, a data visualisation, a small game, an SVG, a slide deck. People see two results from the same brief and pick the better one, and the votes become an ELO the same way a chess rating works. This figure is a model's mean ELO across every category it has been scored in.
What you can take from it: this is the closest thing to a measure of whether a model produces work someone would accept, as opposed to code that merely runs. A model can score well on coding benchmarks — which check whether tests pass — and still produce an ugly, unusable interface, and this is where that shows up.
Weigh it heavily if the output is something a person will look at, and ignore it for backend or data work where nobody sees the result.
Head-to-head votes on what models produce when asked to build a web page: layout, styling and whether the result looks like something you would ship. Scored separately from the blended Design Arena figure because it is the category with the widest coverage and the one most people actually care about.
What you can take from it: coding benchmarks check whether code runs; this checks whether the result is usable. A model can pass tests and still produce a page nobody would accept, and the gap between a model's coding score and this one is usually where that shows up.
Mathematical problem-solving benchmarks ranging from competition-level to graduate-level difficulty.
Competition mathematics problem solving
Grade school math word problems
American Invitational Mathematics Examination — competition math
Benchmarks focused on logical reasoning, scientific understanding, and complex problem decomposition.
Graduate-level science Q&A by domain experts
Abstraction and Reasoning Corpus for general intelligence
Big-Bench Hard — 23 challenging multi-step reasoning tasks
Commonsense reasoning via pronoun resolution
The Agentic Index measures how well a model copes when a job cannot be finished in one reply: it has to plan, call tools, read what comes back, notice when something has gone wrong and change course, sometimes over dozens of steps. The underlying tests include office-style work with real documents and long-running automation tasks judged on whether the end result is correct, not on whether each individual step looked reasonable.
What you can take from it: this is the single most important score if you are building an assistant that uses connectors, fills in forms, files records or runs a workflow on its own. Models that score well on knowledge benchmarks frequently score poorly here, because staying coherent over many steps is a different skill from knowing facts.
A low agentic score is the usual explanation for an assistant that starts a task confidently and then stalls or loops.
The model plays an airline service agent across a multi-turn conversation: it has to read a policy document, call the right tools in the right order, refuse what the policy forbids, and get the customer to a correct outcome. A run counts only if the final state of the booking system is right — a plausible-sounding reply that left the database wrong scores zero.
What you can take from it: this is the closest public measure of whether a model can be trusted to act on someone's behalf rather than just advise them. It is the benchmark to weigh if you are building anything that calls tools, fills forms or changes records, and models that look strong on knowledge tests often fall over here, because following a policy over many turns is a different skill from knowing things.
OpenRouter runs it continuously across 123 models with 48 tasks per run.
Performance metrics measuring how quickly a model responds and generates output tokens.
Tokens generated per second
Latency before first token arrives
How long the model takes to return a finished image, measured from request to result at default settings. What you can take from it: this decides whether the model can sit in front of a user or has to run in the background. Under about five seconds a person will wait; past fifteen you need a queue, a progress state and somewhere to put the result.
It matters far more for an interactive tool than for a batch job, where throughput and cost dominate.
Pricing metrics showing the cost per million tokens for input and output.
Cost per 1M input tokens
Cost per 1M output tokens
Most providers will store the unchanging front part of a prompt — a system prompt, a long document, the earlier turns of a conversation — and charge a heavily reduced rate when the same text is sent again. This is the price of re-reading that stored text, and it is typically a tenth or less of the standard input price.
What you can take from it: on any workload that sends the same context repeatedly, this number decides the bill, not the headline input price. A long chat, a document the user keeps asking about, or an agent that carries its instructions into every step will hit the cached rate on the great majority of its input tokens.
Two models with identical list prices can differ several-fold in what they actually cost you once caching is taken into account, which is why the list price alone is a poor guide. A blank here means the provider does not publish a cached rate, so assume you pay full price on every token.
What a single generated image costs at the provider's default resolution and quality. Image models are billed per image rather than per token, so this is the number that maps to your bill: a hundred product shots at $0.04 is $4, whatever the prompt said. What you can take from it: compare it against how many attempts you actually need.
A cheaper model that takes six tries to get a usable result costs more than a dearer one that lands in two, so read this beside prompt adherence rather than on its own. Providers usually charge more for higher resolutions and for a quality or upscale setting, so treat this as a floor.
What one second of finished video costs. Video models bill by output duration, so a price that looks small multiplies fast: at $0.50 a second, a thirty-second cut is $15 per attempt, and you will not use the first attempt. What you can take from it: budget by finished minute and by how many takes you expect, not by the headline rate.
Read it beside maximum clip length — a model that only produces five seconds at a time needs several generations, and the joins between them, to reach the length you actually want.
Metrics related to the maximum amount of text a model can process in a single request.
Maximum context window size
Long-context understanding and retrieval accuracy at depth
Benchmarks evaluating visual understanding and image-text reasoning capabilities.
Massive Multi-discipline Multimodal Understanding
Visual mathematical reasoning across diagrams and charts
The largest image the model generates directly, before any separate upscaling step, expressed in megapixels so different aspect ratios can be compared. What you can take from it: for anything going to print or to a large screen, native resolution is the ceiling on quality — upscaling afterwards invents detail rather than recovering it.
For web use almost every current model is already past what you need, so this should not drive the decision.
How reliably the model draws legible, correctly spelled words when a prompt asks for them — a poster headline, a label on a product, a sign in a scene. It is scored separately because it is the single most common failure in image generation and the models differ enormously on it. What you can take from it: if your use is marketing assets, mockups, thumbnails or anything carrying a brand name, this matters more than overall image quality, because a beautiful image with mangled lettering is unusable.
If you only ever generate textless imagery, ignore it entirely.
Whether the image contains what you asked for — the right number of objects, the right colours on the right things, the spatial relationships you specified — rather than something merely in the right style. What you can take from it: this is what decides how many attempts a usable result takes, so it drives real cost far more than the per-image price does.
A model that scores well here can be directed; one that scores badly has to be coaxed, and you pay for every discarded attempt.
People are shown two generated images from the same prompt and pick the better one; the votes become an ELO rating. This is the image-generation category of Design Arena specifically, rather than a model's average across every category it competes in. What you can take from it: it is the only broad public measure of whether an image model produces results people actually prefer, which is not the same as resolution or speed.
Read it beside generation time and cost per image — a model that wins on preference but takes three minutes per image is a different proposition from one that is nearly as good in twenty seconds.
How much video the model produces in a single generation, in seconds. Almost every current model is capped well below the length of a finished piece. What you can take from it: this sets how much stitching your workflow needs. A model capped at five seconds cannot produce a twenty-second shot — you generate four clips and join them, and each join is a place where lighting, faces and motion can visibly jump.
A longer cap is worth real money on any narrative work, and almost nothing on short social cuts.
The tallest frame the model outputs natively — 720, 1080, or 2160 for 4K — before any upscaling. What you can take from it: 1080p is the practical floor for anything client-facing or going to a large screen; 720p is fine for social and for previewing. As with images, upscaling afterwards adds pixels rather than detail, so native resolution is the real ceiling.
Whether the model produces audio along with the picture — dialogue, effects and ambience timed to what is on screen — or returns silent video you have to score separately. Scored 100 for yes, 0 for no. What you can take from it: native audio removes a whole production step and, more importantly, removes the lip-sync problem, which is very hard to fix after the fact.
If your output is silent B-roll it is irrelevant; if anyone on screen speaks, it is close to decisive.
We track 45 benchmarks across 8 categories: General Intelligence, Coding, Mathematics, Reasoning, Speed & Latency, Cost & Pricing, Context Window, Multimodal. These cover everything from general knowledge and coding ability to mathematical reasoning, speed, and cost.
SWE-bench Verified is a curated subset of 500 real-world GitHub issues drawn from 12 popular open-source Python repositories, where each problem has been manually validated by software engineers. It gives an AI agent the full repository codebase plus the original issue text, then requires the agent to locate the bug, edit the correct files, and produce a patch that passes all tests.
MMLU-Pro is a significantly harder evolution of the original MMLU benchmark featuring over 12,000 rigorously curated multiple-choice questions across 14 academic domains. Unlike the original MMLU's 4-option format, MMLU-Pro expands each question to 10 answer choices, reducing random guessing from 25% to 10%.
Chatbot Arena is a crowdsourced platform where real users submit prompts and receive responses from two anonymous AI models side by side, then vote for the one they prefer. The platform uses the Bradley-Terry model to convert millions of pairwise votes into a ranked leaderboard, having collected over 6 million votes across 400+ models.
It depends on your use case. For general tasks, Chatbot Arena ELO and MMLU-Pro are key indicators. For software development, prioritize SWE-bench and HumanEval scores. For cost-sensitive applications, compare input and output pricing. For real-time applications, look at output speed and time to first token (TTFT). Use our rankings to weight metrics based on your priorities.
We track 45 benchmarks across 8 categories to give you a comprehensive view of AI model capabilities. Scores are sourced from official benchmark leaderboards, provider announcements, and independent evaluation platforms.
Use the AI Models to weight these benchmarks based on your priorities, or compare models side-by-side. Browse all model profiles or check pricing details.