Conflicts of interest

Leaderboards are only worth reading if you know who benefits from the ranking. Most sites in this space publish nothing on this. Here is everything that could bias what you see here.

What we have a stake in

gcrab โ€” an open-source Go agent runtime built by the operators of this site, listed on the agent frameworks page. It is labelled "our project" everywhere it appears, it is ranked by the same public GitHub metrics as everything else, and it receives no scoring adjustment. If its numbers are ever missing, that means the public source had none, not that we hid them.
We are not paid by any model provider, framework, or benchmark operator. There is no sponsored placement, no paid inclusion, and no affiliate or referral link anywhere on this site. If that ever changes it will be disclosed on this page and labelled inline, not buried in a footer.

Where our data comes from, and who runs it

SourceWho operates itBias to be aware of
Arena
formerly LMArena
Arena (commercial; sells evaluation data to AI labs)Scores are human preference votes, so they reward answers people like, which is not the same as answers that are correct. Published research has also raised concerns about preferential pre-release testing for large labs.
OpenRouterCommercial inference marketplacePrices are what OpenRouter lists, which can differ from going direct to a provider. Usage rankings reflect OpenRouter's own customer mix, which skews toward indie and agent workloads.
SWE-benchAcademic (Princeton/Stanford)Since late 2025 the Verified board restricts submissions to academic teams and established labs, so commercial agents are under-represented. Some entries are self-reported rather than maintainer-checked; we label which is which.
GitHubMicrosoftStars measure attention, not quality, and are trivially influenced by launch publicity.

Things we deliberately don't do

  • No invented composite score. We never average incompatible benchmarks into a single "intelligence index" that can't be audited or reproduced.
  • No estimated numbers. A model with no published score on a board is left out of that board, not given a guess.
  • No ranking ourselves first. gcrab appears where the public metrics put it.
  • No stale data behind a fresh date. Every table carries the retrieval date of the data in it, not the date the page was rebuilt.

Corrections

Found something wrong or something we should disclose? Email [email protected]. Corrections ship in the next daily build, and material ones get noted here.