xAI

Grok 2 — Benchmark Scores, Pricing & Performance Analysis

MIDxAI
Chatbot Arena ELO
1250
Output Speed
80 tok/s
Input Cost
$2.0/1M
Output Cost
$10.0/1M
Context Window
131K

Grok 2 is an AI model by xAI. View detailed benchmark scores, pricing data, and performance metrics on Serenities AI Models.

General Benchmarks

MMLU-Pro
68.0%46th of 66

Massive Multitask Language Understanding professional benchmark

HuggingFace Leaderboard
Chatbot Arena ELO
125054th of 74

Human preference ELO from blind head-to-head votes

LMSYS / HuggingFace
TruthfulQA
79.0%9th of 13

Resistance to generating false but plausible answers

Papers
MT-Bench
8.88th of 13

Multi-turn conversation quality on 80 curated dialogues

LMSYS
AlpacaEval 2.0
55.0%5th of 10

Instruction following quality scored by GPT-4 as judge

tatsu-lab
SimpleQA
36.4%5th of 14

Factual accuracy on short, verifiable questions

OpenAI

Coding Benchmarks

HumanEval+
78.0%48th of 66

Code generation correctness with extended tests

Papers
LiveCodeBench
32.0%52nd of 66

Live competitive programming benchmark

livecodebench.github.io
SWE-bench Verified
28.0%51st of 67

Real-world software engineering task resolution

swebench.com

Math Benchmarks

GSM8K
88.0%47th of 66

Grade school math word problems

Papers
MATH
70.0%50th of 66

Competition mathematics problem solving

Papers

Reasoning Benchmarks

ARC-AGI
15.0%48th of 66

Abstraction and Reasoning Corpus for general intelligence

arcprize.org
GPQA Diamond
45.0%119th of 140

Graduate-level science Q&A by domain experts

Papers
Winogrande
88.5%10th of 14

Commonsense reasoning via pronoun resolution

Papers

Speed Benchmarks

Time to First Token
350 ms31st of 92

Latency before first token arrives

Aggregated
Output Speed
80 tok/s45th of 94

Tokens generated per second

Aggregated

Cost Benchmarks

Output Cost
$10.0335th of 413

Cost per 1M output tokens

OpenRouter API
Input Cost
$2.0336th of 413

Cost per 1M input tokens

OpenRouter API

Context Benchmarks

Context Length
131K212th of 451

Maximum context window size

OpenRouter API

Multimodal Benchmarks

MathVista
60.5%13th of 14

Visual mathematical reasoning across diagrams and charts

Papers

Grok 2 — Benchmark Scores Overview

Scores normalized to percentage scale for visual comparison. ELO scores mapped to 0-100 range (1100-1500).

Grok 2 — Frequently Asked Questions

How intelligent is Grok 2?

Grok 2 scores 1250 on the Chatbot Arena ELO rating, making it a mid-tier AI model. This score is based on blind head-to-head human preference voting.

How much does Grok 2 cost?

Grok 2 costs $2.0 per 1M input tokens and $10.0 per 1M output tokens. This is mid-range pricing for its capability level.

How fast is Grok 2?

Grok 2 generates output at 80 tokens per second, which is moderate compared to other models. The time to first token is 350 ms.

How good is Grok 2 at coding?

Grok 2 achieves 28.0% on SWE-bench Verified, demonstrating basic real-world software engineering capability. This benchmark tests the model's ability to resolve actual GitHub issues.

How good is Grok 2 at math and reasoning?

Grok 2 scores 70.0% on the MATH benchmark (competition-level mathematics). It also achieves 45.0% on GPQA Diamond, a graduate-level science reasoning benchmark.

What is the context window of Grok 2?

Grok 2 has a context window of 131K tokens. This determines how much text, conversation history, and code the model can process in a single request.

Who created Grok 2?

Grok 2 was created by xAI. It is classified as a mid model in our catalogue.

Is Grok 2 open source?

No, Grok 2 is a proprietary model. It is available through xAI's API and compatible providers.