Browse

LLM benchmark scores by model, since 2019

Public
Explore tables

LLM benchmark scores by model, since 2019

Compare large language models on more than 80 benchmarks, from GPQA Diamond and SWE-bench Verified to FrontierMath and Humanity's Last Exam. The table gathers every result published in Epoch AI's Benchmarking Hub, with each score in its own unit, who ran the evaluation and a link to the source. Use it to rank models, follow how scores rose with each release, or set lab and paper numbers beside independent runs.

On 1 October 2026 the hub held about 6,800 results for about 1,200 model versions of 710 models from OpenAI, Anthropic, Google DeepMind, Meta, Alibaba and about 60 other developers. Each row is one model version on one benchmark under one evaluation setup.

Coverage

  • Window: models released from October 2019 to Epoch AI's latest update, most of them in 2025 and 2026. Dates are UTC calendar dates.
  • Grain: one row per published result, so a model can have several rows on one benchmark when the shots, harness, reasoning effort or source differ.
  • Cadence: the table is rebuilt every day from Epoch AI's download. Epoch AI changed 36 of its benchmark tables between 29 September and 1 October 2026.
  • New benchmarks: a benchmark Epoch AI adds to its Capabilities Index appears automatically, with one best score per model family, until its full table is added to the build.

Columns

  • result_id: Epoch AI's record id for the result, or a hash of the row where Epoch AI publishes none.
  • benchmark, benchmark_release_date: the benchmark as Epoch AI names it, and the date it was published.
  • model, model_version: the model family and the exact version evaluated, usually the developer's API name plus Epoch AI's reasoning-setting suffix.
  • organization, country, model_access, model_release_date: the developer, its country, how the model is released (API or open weights) and when.
  • training_compute_flop: Epoch AI's estimate of training compute, in floating-point operations.
  • score, score_unit, higher_is_better: the score as published, its unit and its direction. Units are fraction, percent, points out of 10, Elo rating, rating, speedup factor, Brier score, index points or US dollars.
  • score_fraction: the score as a fraction from 0 to 1, for benchmarks scored as a proportion.
  • stderr: the standard error the source reports, in the unit of score.
  • result_date: the run, grading, addition or update date, where the source gives one.
  • reported_by: Epoch AI for its own runs, Leaderboard, Model developer, Paper or report, or Not stated.
  • source, source_url: the cited paper, leaderboard or site and its link, or Epoch AI's evaluation log for its own runs.
  • evaluation_setup: the shots, harness, agent, reasoning effort or provider that tell several results for one model apart.
  • notes, epoch_file: Epoch AI's note on the result and the file in its download the row comes from.

Missing values

An empty cell means the source did not publish that value. Rows without a numeric score or a model name are left out, about 70 of the published rows on 1 October 2026. Results for older models often lack an organisation, access type or release date. Most results carry no standard error or result date. Epoch AI does not flag lab self-reports, so labs' own numbers appear under Paper or report and Model developer, next to academic papers.

Suitable for

  • Ranking models on one benchmark
  • Tracking how the best scores changed with model release dates
  • Comparing Epoch AI's independent runs with numbers from leaderboards, papers and labs
  • Relating benchmark scores to training compute and model access

Source and rights

Every row comes from the Epoch AI Benchmarking Hub at epoch.ai/benchmarks, read daily from its download. Epoch AI publishes its own evaluation runs under the Creative Commons Attribution 4.0 licence. Results Epoch AI compiled from other projects keep those projects' licences, so credit Epoch AI and the source linked in each row.

Tables

Table healthLivellm_benchmark_resultsTable detailsNamellm_benchmark_resultsColumns22
Table overviewPublished rows6,843Columns22 rows22 cols
Update detailsStatusLiveLast published2 Oct 2026

Sources

1 publisher
Epoch AIEpoch AI is a research institute investigating key trends and questions that will shape the trajectory and governance of Artificial Intelligence.2 endpoints
About
Epoch AI is a research institute investigating key trends and questions that will shape the trajectory and governance of Artificial Intelligence.
Usage rights
Licensed by contract.
Requests
85 requests across 2 endpoints
EndpointRequests / coverage
/data/benchmark_data.zipANLI scores that Epoch AI compiled from published leaderboards, papers and reports.84 requests
/data/processed_data_for_eci.csvBest score per model family on each Epoch Capabilities Index benchmark, scaled between the benchmark's random baseline and ceiling.1 request

Details

Contents
1 table · 6,843 rows (est.) · 22 columns
Updated
2 October 2026
Published
2 October 2026
Version
v1
License
Not stated
Visibility
Public
Publisher
Mostly Right
Topics
llm benchmarks · large language models · ai evaluation +4

Activity

Views254+254 in the last 30 days
Likes0+0 in the last 30 days
Uses10+10 in the last 30 days

Comments

0

No comments yet. Questions about coverage, licensing, or how a column is derived belong here, in the open, beside the data.

Sign in to join the conversation.

Use this data for free

A free workspace gives you an API key to query whole tables, download Parquet and connect your AI tools.

We are in open beta, so expect changes week to week.

Create a free workspace

Already have an account? Log in

DocsPricingTermsPrivacyStatus