GPT-4o is an AI model by OpenAI. View detailed benchmark scores, pricing data, and performance metrics on Serenities AI Models.
Specifications last checked against the provider on . Benchmark scores carry their own source, linked beside each number.
General Benchmarks
Massive Multitask Language Understanding professional benchmark
HuggingFace LeaderboardHuman preference ELO from blind head-to-head votes
LMSYS / HuggingFaceCoding Benchmarks
Human preference ELO for building things — websites, UI, games, charts, SVG
Design Arena (via OpenRouter)Human preference ELO for building a working web page from a brief
Design Arena (via OpenRouter)Math Benchmarks
Reasoning Benchmarks
Multi-turn service agent making tool calls under strict policy constraints
OpenRouter (measured)Speed Benchmarks
Cost Benchmarks
Cost per 1M cached input tokens — the price that actually applies to a long agent conversation
OpenRouter APIContext Benchmarks
Available from 22 providers
The same model costs different amounts depending on who serves it. You can bring your own key for 4 of these — connect it here.
| Provider | Input / 1M | Output / 1M | Cached in | Context | Your key |
|---|---|---|---|---|---|
| Cloudflare AI Gateway | $1.25 | $5 | $0.625 | 128K | — |
| 302.AI | $2.5 | $10 | — | 128K | — |
| Abacus | $2.5 | $10 | — | 128K | — |
| Azure | $2.5 | $10 | $1.25 | 128K | Supported |
| Azure Cognitive Services | $2.5 | $10 | $1.25 | 128K | — |
| DevPass (LLM Gateway) | $2.5 | $10 | $1.25 | 128K | — |
| Eden AI | $2.5 | $10 | $1.25 | 128K | — |
| FrogBot | $2.5 | $10 | $1.25 | 128K | — |
| Impossibl | $2.5 | $10 | $1.25 | 128K | — |
| Kilo Gateway | $2.5 | $10 | $1.25 | 128K | — |
| LLM Gateway | $2.5 | $10 | $1.25 | 128K | — |
| Merge Gateway | $2.5 | $10 | $1.25 | 128K | — |
| NanoGPT | $2.5 | $10 | $1.25 | 128K | — |
| Ofox | $2.5 | $10 | $1.25 | 128K | — |
| OpenAI | $2.5 | $10 | $1.25 | 128K | Supported |
| OpenRouter | $2.5 | $10 | $1.25 | 128K | Supported |
| OrcaRouter | $2.5 | $10 | $1.25 | 128K | — |
| Pioneer | $2.5 | $10 | $1.25 | 128K | — |
| Vercel AI Gateway | $2.5 | $10 | $1.25 | 128K | Supported |
| Cortecs | $2.659 | $10.635 | $1.33 | 128K | — |
| Venice AI | $3.125 | $12.5 | — | 128K | — |
| Poe | — | — | — | 128K | — |
GPT-4o — Benchmark Scores Overview
Scores normalized to percentage scale for visual comparison. ELO scores mapped to 0-100 range (1100-1500).
Compare GPT-4o With
GPT-4o — Frequently Asked Questions
How intelligent is GPT-4o?
GPT-4o scores 1280 on the Chatbot Arena ELO rating, making it a mid-tier AI model. This score is based on blind head-to-head human preference voting.
How much does GPT-4o cost?
GPT-4o costs $2.5 per 1M input tokens and $10.0 per 1M output tokens. This is mid-range pricing for its capability level.
How fast is GPT-4o?
GPT-4o generates output at 143 tokens per second, which is moderate compared to other models. The time to first token is 450 ms.
How good is GPT-4o at coding?
GPT-4o achieves 30.7% on SWE-bench Verified, demonstrating moderate real-world software engineering capability. This benchmark tests the model's ability to resolve actual GitHub issues.
How good is GPT-4o at math and reasoning?
GPT-4o scores 74.6% on the MATH benchmark (competition-level mathematics). It also achieves 50.3% on GPQA Diamond, a graduate-level science reasoning benchmark.
What is the context window of GPT-4o?
GPT-4o has a context window of 128K tokens. This determines how much text, conversation history, and code the model can process in a single request.
Who created GPT-4o?
GPT-4o was created by OpenAI. It is classified as a mid model in our catalogue.
Is GPT-4o open source?
No, GPT-4o is a proprietary model. It is available through OpenAI's API and compatible providers.