Evidence before rankings

Methodology & sources

A ranking is useful when you can see what was measured, when it was measured, and how uncertain the result is. Here is how to read AI 114 and check the evidence yourself.

Reviewed Source-linked dataNo proprietary composite score
Download evidence · JSONExplore rankings →

Preference has a scope

Arena measures preferences within its evaluation setting. A high score is not a guarantee of accuracy or success on your task.

Uncertainty stays visible

Read scores together with confidence intervals, rank spreads, vote counts and any preliminary status.

Missing stays missing

Unverified data is excluded from public rankings. We do not invent values, substitute another model or manufacture a trust percentage.

01 · What the rankings measure

Arena is the primary source for the preference rankings shown here. Its standard leaderboards aggregate comparisons of model responses using a Bradley–Terry model. This estimates relative preference from pairwise outcomes; it is not an exam score or a measure of universal intelligence. Arena also publishes distinct evaluation types such as AutoEval, which require their own labels.

Text / Coding
Preferences on text responses, or the coding-related subset of those conversations. This does not directly establish repository-level software engineering performance.
Web development
Evaluation of generated web applications in Code Arena. Its task environment and category filters differ from general coding conversations and other agent benchmarks.
Other categories
Image, video, vision, search and other boards measure their respective tasks. Scores from different boards are not interchangeable.
Adjustment setting
The source category and raw / style-control setting must be identified. “Raw” means no style adjustment; style control estimates preference while adjusting for selected style features. Neither setting removes every possible bias.
Independent context: Artificial Analysis offers separate intelligence and coding-agent evaluations. Its Intelligence Index is primarily text-based and English-language. We link to these evaluations for context and do not merge them into an AI 114 overall score.

02 · Read beyond the rank

Arena score
A relative estimate in one evaluation. Higher is more preferred within that setting. Differences are not percentage gains in intelligence or correctness.
95% confidence interval
An interval describing uncertainty in the estimated score under the statistical method. It is not a 95% chance that a response is correct, nor a guarantee that the data represents every user or use case.
Rank / rank spread
Raw rank orders score estimates. Rank spread describes optimistic and pessimistic ranks calculated from score confidence intervals. Overlap is a reason to avoid claiming a decisive winner; it does not prove identical capability.
Votes
The source's comparison count, not a count of unique users or independent customers. More observations can reduce estimation uncertainty; they do not eliminate selection, prompt or judging bias.
Preliminary
A provisional status supplied by the source. We preserve it instead of relabeling an early estimate as a settled result.
AutoEval
An evaluation using proxy votes from a reward model trained on human preferences. This is distinguished from human-vote results. A missing official rank or vote count remains unavailable.
We preserve source uncertainty when available. If an interval or rank spread cannot be verified, it is shown as unavailable rather than reconstructed from guesses. We do not assign editorial “reliability percentages.”

03 · Compare the actual models

  • Keep model identity. Versions, reasoning settings and evaluation configurations matter. If the exact model is absent from a category, the result is unavailable. Another model from the same company cannot take its place.
  • Keep the source rank. Search, open-weight and company filters change the visible rows, not the model's original ranking. A company's best listed model does not measure that entire company's products.
  • Do not invent head-to-head outcomes. AI 114 does not turn score differences into advertised win probabilities, or count category wins as an overall verdict.
  • Use comparable evidence. Compare models within the same source, category, adjustment and data release. Consider uncertainty and task fit alongside point estimates.

04 · Separate dates. Check before publishing.

Source date is the date of the source's leaderboard release. Retrieved at is when AI 114 obtained it. Reviewed / deployed at describes our own work. Fetching an older leaderboard today does not make its observations current.

  1. Record the source URL, category, adjustment setting, source date and retrieval time with the data.
  2. Check that the response contains the expected board and valid model records. Preserve uncertainty, original ranks and evaluation status; do not convert missing values into zero.
  3. Publish only data that passes verification. A changed page, parse error or failed request must not silently replace the last checked dataset.
  4. When refresh or validation fails, use the last verified saved snapshot with its date and saved-data status. If no verified snapshot exists, show the category as unavailable.

Server caching can delay a refresh. “Fetched” or “cached” describes delivery, not the age of the source evaluation. The downloadable evidence record is provided so that displayed data can be traced to its source.

05 · Prices, users & release claims

Prices: A price needs a source, check date, currency, billing unit and relevant conditions. API token pricing and consumer subscription pricing are different. Input and output rates, caching, model settings, region, tax and media resolution can affect the amount paid. Unverified prices are withheld; a price–score chart is not a universal value ranking.

User counts: Monthly and weekly active users, registered accounts, visits, app-only audiences and regional counts are different measures. We do not combine them into a user ranking. A published figure needs the exact metric, scope, period and source. Unsupported figures are excluded.

Model availability: An OpenRouter listing date is not a provider's official launch date. Availability, planned releases and subscription details require current source evidence. Rumors and unsupported future dates are not presented as confirmed facts.

06 · Read the original evidence

These links explain the source data and its limits. Live source pages may change after this review; the evidence download records the data used for this release.

  1. Arena · Text leaderboard ↗Original text rankings, source release date, scores, rank spreads and votes.
  2. Arena · Code / WebDev leaderboard ↗Web-development task scope, category settings and evaluation status.
  3. Arena · Ranking method ↗Raw ranks and the calculation and interpretation of rank spreads.
  4. Arena · Open-source methodology ↗Bradley–Terry implementation, confidence intervals and style adjustment.
  5. Arena · AutoEval ↗Reward-model evaluations and how they differ from human-vote evidence.
  6. Arena · Leaderboard changelog ↗Method and data changes that can affect comparability over time.
  7. Artificial Analysis · Intelligence methodology ↗Supplementary source: evaluation composition, settings and scope. Not combined with Arena scores.
  8. Artificial Analysis · Coding Agent methodology ↗Supplementary source: repository tasks, terminal workflows and agent evaluation.
  9. OpenRouter · Model catalog API ↗Catalog properties and pricing fields; listing metadata is not a provider launch announcement.

Policy review: September 24, 2026. A successful data refresh verifies retrieval and structure; it does not independently reproduce the source's evaluation or certify the model's answers.