01 · What the rankings measure
Arena is the primary source for the preference rankings shown here. Its standard leaderboards aggregate comparisons of model responses using a Bradley–Terry model. This estimates relative preference from pairwise outcomes; it is not an exam score or a measure of universal intelligence. Arena also publishes distinct evaluation types such as AutoEval, which require their own labels.
- Text / Coding
- Preferences on text responses, or the coding-related subset of those conversations. This does not directly establish repository-level software engineering performance.
- Web development
- Evaluation of generated web applications in Code Arena. Its task environment and category filters differ from general coding conversations and other agent benchmarks.
- Other categories
- Image, video, vision, search and other boards measure their respective tasks. Scores from different boards are not interchangeable.
- Adjustment setting
- The source category and raw / style-control setting must be identified. “Raw” means no style adjustment; style control estimates preference while adjusting for selected style features. Neither setting removes every possible bias.
Independent context: Artificial Analysis offers separate intelligence and coding-agent evaluations. Its Intelligence Index is primarily text-based and English-language. We link to these evaluations for context and do not merge them into an AI 114 overall score.
02 · Read beyond the rank
- Arena score
- A relative estimate in one evaluation. Higher is more preferred within that setting. Differences are not percentage gains in intelligence or correctness.
- 95% confidence interval
- An interval describing uncertainty in the estimated score under the statistical method. It is not a 95% chance that a response is correct, nor a guarantee that the data represents every user or use case.
- Rank / rank spread
- Raw rank orders score estimates. Rank spread describes optimistic and pessimistic ranks calculated from score confidence intervals. Overlap is a reason to avoid claiming a decisive winner; it does not prove identical capability.
- Votes
- The source's comparison count, not a count of unique users or independent customers. More observations can reduce estimation uncertainty; they do not eliminate selection, prompt or judging bias.
- Preliminary
- A provisional status supplied by the source. We preserve it instead of relabeling an early estimate as a settled result.
- AutoEval
- An evaluation using proxy votes from a reward model trained on human preferences. This is distinguished from human-vote results. A missing official rank or vote count remains unavailable.
We preserve source uncertainty when available. If an interval or rank spread cannot be verified, it is shown as unavailable rather than reconstructed from guesses. We do not assign editorial “reliability percentages.”
03 · Compare the actual models
- Keep model identity. Versions, reasoning settings and evaluation configurations matter. If the exact model is absent from a category, the result is unavailable. Another model from the same company cannot take its place.
- Keep the source rank. Search, open-weight and company filters change the visible rows, not the model's original ranking. A company's best listed model does not measure that entire company's products.
- Do not invent head-to-head outcomes. AI 114 does not turn score differences into advertised win probabilities, or count category wins as an overall verdict.
- Use comparable evidence. Compare models within the same source, category, adjustment and data release. Consider uncertainty and task fit alongside point estimates.
04 · Separate dates. Check before publishing.
Source date is the date of the source's leaderboard release. Retrieved at is when AI 114 obtained it. Reviewed / deployed at describes our own work. Fetching an older leaderboard today does not make its observations current.
- Record the source URL, category, adjustment setting, source date and retrieval time with the data.
- Check that the response contains the expected board and valid model records. Preserve uncertainty, original ranks and evaluation status; do not convert missing values into zero.
- Publish only data that passes verification. A changed page, parse error or failed request must not silently replace the last checked dataset.
- When refresh or validation fails, use the last verified saved snapshot with its date and saved-data status. If no verified snapshot exists, show the category as unavailable.
Server caching can delay a refresh. “Fetched” or “cached” describes delivery, not the age of the source evaluation. The downloadable evidence record is provided so that displayed data can be traced to its source.
05 · Prices, users & release claims
Prices: A price needs a source, check date, currency, billing unit and relevant conditions. API token pricing and consumer subscription pricing are different. Input and output rates, caching, model settings, region, tax and media resolution can affect the amount paid. Unverified prices are withheld; a price–score chart is not a universal value ranking.
User counts: Monthly and weekly active users, registered accounts, visits, app-only audiences and regional counts are different measures. We do not combine them into a user ranking. A published figure needs the exact metric, scope, period and source. Unsupported figures are excluded.
Model availability: An OpenRouter listing date is not a provider's official launch date. Availability, planned releases and subscription details require current source evidence. Rumors and unsupported future dates are not presented as confirmed facts.
Policy review: September 24, 2026. A successful data refresh verifies retrieval and structure; it does not independently reproduce the source's evaluation or certify the model's answers.