Leaderboards are only worth reading if you know who benefits from the ranking. Most sites in this space publish nothing on this. Here is everything that could bias what you see here.
| Source | Who operates it | Bias to be aware of |
|---|---|---|
| Arena formerly LMArena | Arena (commercial; sells evaluation data to AI labs) | Scores are human preference votes, so they reward answers people like, which is not the same as answers that are correct. Published research has also raised concerns about preferential pre-release testing for large labs. |
| OpenRouter | Commercial inference marketplace | Prices are what OpenRouter lists, which can differ from going direct to a provider. Usage rankings reflect OpenRouter's own customer mix, which skews toward indie and agent workloads. |
| SWE-bench | Academic (Princeton/Stanford) | Since late 2025 the Verified board restricts submissions to academic teams and established labs, so commercial agents are under-represented. Some entries are self-reported rather than maintainer-checked; we label which is which. |
| GitHub | Microsoft | Stars measure attention, not quality, and are trivially influenced by launch publicity. |