Most preferred AI model in blind human testing
The highest style-controlled text rating on the Arena human-preference leaderboard, held by the model people choose most often in blind head-to-head comparisons.
View chart data
| As of | Arena style-controlled Elo rating |
|---|---|
| Jul 21, 23:59 UTC | 1,507 Elo |
About this metric
Benchmarks measure what models can do; preference ratings measure what people actually choose. The Arena leaderboard collects blind head-to-head votes: a person asks a question, two anonymous models answer, and the person picks the better response. This chart tracks the top style-controlled text rating on that board, whichever model holds it.
Ratings are Elo scores computed from the full pairwise vote record, republished here from the official style-controlled text leaderboard, which corrects for formatting effects such as response length and markdown density. Elo is relative, so the number is only meaningful against the rest of the field, and a lead of a few points within the published confidence interval is a statistical tie.
Human preference is the metric assistants are ultimately tuned for, and it often diverges from objective benchmarks: a model can lead on reasoning scores while losing blind votes on helpfulness and tone. When the same model tops both this line and the objective-benchmark line, the field has a clear frontier; when they split, the market is choosing between capability and likability.