Every AI software pitch arrives with benchmark claims attached. "Scores highest on reasoning." "Leads the coding leaderboard." What the pitch does not include is which benchmarks were left out, or how the same model performs on the tests where it does not come first.
BenchmarkList at benchmarklist.com is a free registry tracking over 2,500 AI evaluations and more than 10,000 models. It is not affiliated with any AI company. It groups results by the capability being tested, coding, reasoning, agentic tasks, safety, speed, and cost, so you can check a specific model across multiple categories in one place.
What the Numbers on BenchmarkList Show
FrontierCode is a coding evaluation published by Cognition AI that tests whether models can complete difficult tasks against production-codebase standards. The August 2026 results show Claude Opus 5 at 53.4%, Claude Fable 5 at 46.3%, and Claude Sonnet 5 at 38.8%. The top-ranked model passes just over half of the hardest coding tasks. None of them pass two-thirds. That context does not appear in vendor materials.
ARC-AGI-2, a reasoning benchmark designed to resist memorisation, shows GPT-5.5 leading at roughly 85%, above the human average of around 66%. On TAU3-Bench, which tests banking and customer-service agent work, DeepSeek V4 Flash and DeepSeek V4 Pro hold the top two positions. The model that leads on reasoning does not lead on customer-service tasks. These are separate capability categories with separate leaderboards.
Why the Benchmarks in Most AI Marketing Now Mean Little
MMLU, HumanEval, and GSM8K, the evaluations most commonly cited in AI marketing materials, are saturated. Every frontier model scores above 90% on all three. A vendor citing a 91% MMLU result in 2026 is citing a number that no longer separates any model from its closest competitors.
BenchmarkList carries both the saturated legacy benchmarks and the newer evaluations that still differentiate models. It also hosts the Artificial Analysis Intelligence Index, a composite score weighting agents, coding, scientific reasoning, and general capability, last updated July 2026. That combination makes it straightforward to see which claim a vendor is resting their case on: an active benchmark with real spread between models, or one where every frontier model already maxes out.
What a South African Business Should Do With This
Before committing to an AI platform for a specific business function, look up the underlying model and check the benchmark category that matches your use case. Customer-facing work maps to agent task benchmarks. Document processing maps to long-context and instruction-following tests. Coding tools belong on FrontierCode or SWE-bench, not MMLU.
If a supplier cannot name the specific model powering their AI feature, that is worth noting. BenchmarkList's model directory covers over 10,000 models. If a model is real and has been publicly evaluated, it will be in there.
For a Gauteng business spending R10,000 or more per year on an AI platform, this check takes fifteen minutes and answers one specific question before you sign: is the capability claim this vendor is making supported by the evaluations that actually test it, or by the ones where every frontier model already scores above 90%?