BenchmarksLiveAInews
NEW MODELGPT-6 AstraThe briefingGo Pro
THE MODEL SCOREBOARD

Big claims.
Measured results.

See how AI models perform on the tests behind the headlines. Explore the scores, compare like for like, and follow the evidence.

Automatic source updates· Hourly checks while in use
267models
59benchmarks
2,842results

Open data from Epoch AI

EXPLORE THE EVIDENCE

Choose a benchmark.

Retrieved 2026-09-07 18:18 UTC

GPQA diamond

1 matching results · higher is better

Adjusted scores on a 0–100 scale, not raw accuracy percentages. Epoch rescales some tests for chance performance or known errors. Compare within the same benchmark. Model release dates are not test dates.

“Reported in” identifies the report containing the result; it can compare models from other companies. Effort is shown only when specified in the dataset.

ModelAdjusted scoreModel releasedReported in
GPT-6 AstraEffort: Max94.4/100Epoch AI datasetOriginal report not specified