OpenAI

GPT-4.1 — Benchmark Scores, Pricing & Performance Analysis

Chatbot Arena ELO
1340
Output Speed
70 tok/s
Input Cost
$2.0/1M
Output Cost
$8.0/1M
Context Window
1.0M
Max Output
33K
Knowledge Cutoff
Jun 2024
Accepts
image, text, file

GPT-4.1 by OpenAI demonstrates strong general intelligence, solid coding performance. View detailed benchmark data including scores across coding, math, reasoning, speed, and cost metrics.

Specifications last checked against the provider on . Benchmark scores carry their own source, linked beside each number.

General Benchmarks

MMLU-Pro
78.0%29th of 66

Massive Multitask Language Understanding professional benchmark

HuggingFace Leaderboard
Chatbot Arena ELO
134033rd of 74

Human preference ELO from blind head-to-head votes

LMSYS / HuggingFace
IFEval
88.0%12th of 25

Strict instruction following accuracy on verifiable constraints

Google Research
AlpacaEval 2.0
60.2%3rd of 10

Instruction following quality scored by GPT-4 as judge

tatsu-lab
TruthfulQA
85.0%3rd of 13

Resistance to generating false but plausible answers

Papers
MT-Bench
9.42nd of 13

Multi-turn conversation quality on 80 curated dialogues

LMSYS

Coding Benchmarks

Website Arena ELO
105192nd of 108

Human preference ELO for building a working web page from a brief

Design Arena (via OpenRouter)
LiveCodeBench
55.0%23rd of 66

Live competitive programming benchmark

livecodebench.github.io
HumanEval+
89.0%16th of 66

Code generation correctness with extended tests

Papers
SWE-bench Verified
50.0%27th of 67

Real-world software engineering task resolution

swebench.com
Design Arena ELO
1034101st of 120

Human preference ELO for building things — websites, UI, games, charts, SVG

Design Arena (via OpenRouter)

Math Benchmarks

GSM8K
94.0%27th of 66

Grade school math word problems

Papers
MATH
83.0%31st of 66

Competition mathematics problem solving

Papers

Reasoning Benchmarks

τ²-Bench Airline
48.7%73rd of 93

Multi-turn service agent making tool calls under strict policy constraints

OpenRouter (measured)
ARC-AGI
40.0%22nd of 66

Abstraction and Reasoning Corpus for general intelligence

arcprize.org
GPQA Diamond
64.9%94th of 140

Graduate-level science Q&A by domain experts

Papers
Winogrande
93.0%2nd of 14

Commonsense reasoning via pronoun resolution

Papers

Speed Benchmarks

Time to First Token
450 ms52nd of 92

Latency before first token arrives

Aggregated
Output Speed
70 tok/s51st of 94

Tokens generated per second

Aggregated

Cost Benchmarks

Cached Input Cost
$0.50120th of 136

Cost per 1M cached input tokens — the price that actually applies to a long agent conversation

OpenRouter API
Output Cost
$8.0327th of 413

Cost per 1M output tokens

OpenRouter API
Input Cost
$2.0336th of 413

Cost per 1M input tokens

OpenRouter API

Context Benchmarks

Context Length
1.0M48th of 451

Maximum context window size

OpenRouter API
RULER
93.2%2nd of 13

Long-context understanding and retrieval accuracy at depth

Papers

Available from 27 providers

The same model costs different amounts depending on who serves it. You can bring your own key for 4 of these — connect it here.

ProviderInput / 1MOutput / 1MCached inContextYour key
Poe$1.8$7.2$0.451.0M—
302.AI$2$8—1.0M—
Abacus$2$8$0.51.0M—
Azure$2$8$0.51.0MSupported
Azure Cognitive Services$2$8$0.51.0M—
Cloudflare AI Gateway$2$8$0.51.0M—
DevPass (LLM Gateway)$2$8$0.51.0M—
Eden AI$2$8$0.51.0M—
FastRouter$2$8$0.51.0M—
Impossibl$2$8$0.51.0M—
Kilo Gateway$2$8$0.51.0M—
LLM Gateway$2$8$0.51.0M—
Merge Gateway$2$8$0.51.0M—
NanoGPT$2$8$0.51.0M—
NEAR AI Cloud$2$8$0.51.0M—
Ofox$2$8$0.51.0M—
OpenAI$2$8$0.51.0MSupported
OpenRouter$2$8$0.51.0MSupported
OrcaRouter$2$8$0.51.0M—
Pioneer$2$8$11.0M—
SAP AI Core$2$8$0.321.0M—
Vercel AI Gateway$2$8$0.51.0MSupported
Cortecs$2.192$8.769$0.5461.0M—
Requesty$2.2$8.8$0.551.0M—
AnyAPI———1.0M—
Model Oracle AI———1.0M—
Snowflake Cortex———1.0M—

GPT-4.1 — Benchmark Scores Overview

Scores normalized to percentage scale for visual comparison. ELO scores mapped to 0-100 range (1100-1500).

GPT-4.1 — Frequently Asked Questions

How intelligent is GPT-4.1?

GPT-4.1 scores 1340 on the Chatbot Arena ELO rating, making it a mid-tier AI model. This score is based on blind head-to-head human preference voting.

How much does GPT-4.1 cost?

GPT-4.1 costs $2.0 per 1M input tokens and $8.0 per 1M output tokens. This is mid-range pricing for its capability level.

How fast is GPT-4.1?

GPT-4.1 generates output at 70 tokens per second, which is slower, prioritizing quality over speed compared to other models. The time to first token is 450 ms.

How good is GPT-4.1 at coding?

GPT-4.1 achieves 50.0% on SWE-bench Verified, demonstrating strong real-world software engineering capability. This benchmark tests the model's ability to resolve actual GitHub issues.

How good is GPT-4.1 at math and reasoning?

GPT-4.1 scores 83.0% on the MATH benchmark (competition-level mathematics). It also achieves 64.9% on GPQA Diamond, a graduate-level science reasoning benchmark.

What is the context window of GPT-4.1?

GPT-4.1 has a context window of 1.0M tokens. This determines how much text, conversation history, and code the model can process in a single request.

Who created GPT-4.1?

GPT-4.1 was created by OpenAI. It is classified as a mid model in our catalogue.

Is GPT-4.1 open source?

No, GPT-4.1 is a proprietary model. It is available through OpenAI's API and compatible providers.