AI Benchmarks

What each AI model scores, and what it costs per task, at every reasoning effort.

About AI Benchmarks

What each AI model scores, and what that score costs, at every reasoning effort. The Curves view draws each model’s effort levels as one line of score against cost. The Leaderboard ranks the same models by score, with cost as a budget you set.

Where the data comes from

Every score and every cost on this site comes from Artificial Analysis, which benchmarks AI models independently. We read it through their free Data API, never by scraping their pages. Their free tier asks for credit wherever the data appears, so every page names them in its footer.

The data refreshes every hour, on the hour. Each refresh reads the full model list, checks that it has the shape we expect, and joins each model’s effort settings into one line. If a refresh fails, or the new list has lost more than half of its data points, we keep showing the last good copy instead.

Whenever a refresh brings different numbers, the History view records every score, price and release that changed, with a short AI-written summary of what moved.

The data on this page was last updated .

Where the idea came from

In his video “Anthropic Actually Fixed Opus”, Theo (t3.gg) showed a chart that drew each model’s reasoning-effort levels as one curve of Intelligence Index against cost per task. It made the trade-off between paying more and getting a smarter answer easy to see. The Curves view is that chart, and the Leaderboard builds on it.

What the numbers mean

Intelligence Index
Artificial Analysis’s combined score across its set of evaluations (version 4.3 in the data shown here). Their methodology page lists what goes into it.
Coding Index and Agentic Index
Narrower scores from Artificial Analysis, each an equal-weighted average of the evaluations in its area. Their capability indices page defines both.
Cost per task
What Artificial Analysis measured it cost, in US dollars, to run one Intelligence Index task at that setting. We use the same cost beside Coding and Agentic scores, so a setting costs the same whichever score you rank by.
Effort level
Many models let you choose how much they reason before they answer, usually named low, medium, high, xhigh or max. Artificial Analysis tests each setting as its own entry. We join the settings of one release into one line, lowest effort first. More effort usually scores higher and costs more.
lowmedhighxhmaxover budget
Dots grow with effort. On the Leaderboard, a hollow ring is a setting over your budget.
Why some models have fewer levels
A model shows only the settings its maker offers and Artificial Analysis has tested. A model with one setting is a single dot.
Why new models can lack Coding or Agentic scores
The newest releases can have an Intelligence Index score before Artificial Analysis publishes their Coding or Agentic scores. Until it does, we list them as not yet scored, never as zero.

Who made this

AI Benchmarks was created by Daniel Miessler. It is independent, and not affiliated with Artificial Analysis, or with Theo or t3.gg.