Big claims.Measured results. See how AI models perform on the tests behind the headlines. Explore the scores, compare like for like, and follow the evidence.
Automatic source updates· Hourly checks while in use
267 models
59 benchmarks
2,842 results
Open data from Epoch AI
EXPLORE THE EVIDENCE
Choose a benchmark. Retrieved 2026-09-07 18:18 UTC
BenchmarkANLI APEX-Agents ARC AI2 ARC-AGI ARC-AGI-2 Aider polyglot BBH Balrog CL-bench CL-bench Life CadEval Chess Puzzles CritPt Cybench DTBench DeepResearch Bench DeepSWE EBR-bench ExploitBench Fiction.LiveBench FrontierCode FrontierMath-2025-02-28-Private FrontierMath-Tier-4-2025-07-01-Private FrontierMath-Tier-4-v2-Private FrontierMath-Tiers-1-3-v2-Private GDPval GPQA diamond GSM8K GSO-Bench GeoBench HLE HellaSwag LAMBADA LMCA Lech Mazur Writing MATH level 5 METR Time Horizons MMLU MirrorCode Mystery Game Puzzles OSWorld OSWorld 2.0 OTIS Mock AIME 2024-2025 OpenBookQA PIQA PostTrainBench ProofBench Remote Labor Index SWE-Bench verified ScienceQA SimpleBench SimpleQA Verified Surface Evolver Bench Terminal Bench The Agent Company TriviaQA VPCT WeirdML Winogrande Find a model Sort byHighest score Newest model release
GPQA diamond 1 matching results · higher is better Adjusted scores on a 0–100 scale, not raw accuracy percentages. Epoch rescales some tests for chance performance or known errors. Compare within the same benchmark. Model release dates are not test dates.
“Reported in” identifies the report containing the result; it can compare models from other companies. Effort is shown only when specified in the dataset.
Read the score with its context. This is Epoch AI’s curated collection of model benchmark results, including third-party evaluations. It does not cover every benchmark or every new model. A high score on one test does not establish overall superiority. Benchmarks with different versions remain separate.
We check for a newer source snapshot hourly when this page is used. If the source is unavailable, we keep the last successful data and show its retrieval time. The source does not provide a test timestamp for every result.
Data: Epoch AI, AI Benchmarking Hub , under CC BY 4.0 . This uses the processed ECI dataset: adjusted benchmark scores are scaled to 0–100 and rounded for display. These are not the composite ECI rating. See the score adjustments . External Aider Polyglot and Terminal-Bench data retain their Apache 2.0 licence . No benchmark questions are reproduced. Source & methodology