Skip to content
App Builder Index

News

Why We Stopped Trusting Vendor Benchmarks, and What We Test Instead

Every builder in this index ships a benchmark page claiming category leadership. We used to cite them. Here is why we built our own six briefs instead, and what changed once we did.

Tom Brackett ·

Open the marketing site of any tool in this index and you will find a benchmark. A completion rate, a speed comparison, a leaderboard position, usually against a hand-picked set of rivals on a task the vendor itself designed. We used to link to these in early drafts of our reviews as supporting evidence. We stopped, and the reason is worth explaining because it is the same reason our own numbers are worth checking.

The problem with a benchmark the vendor controls

A vendor choosing its own test, its own comparison set, and its own success criteria is not lying when it publishes a favourable number. It is doing exactly what any of us would do with an assignment we get to grade ourselves. The task gets selected for a strength, the comparison set gets selected for a weakness, and the result is a number that is technically true and practically useless for deciding whether the tool will hold up on your actual project.

We saw this most clearly on agent performance claims. Three different tools in this index have, at different points, claimed the top spot on a coding-agent leaderboard. All three cannot be leading the same leaderboard at the same time unless the leaderboards are measuring different things, which they are: different task sets, different scoring, different disclosure about how many attempts were allowed before the best one was reported.

What we test instead

Six fixed briefs, described in full on how we review, run identically on every tool: a CRUD application, a marketing site, a live third-party integration, a scope change forty prompts in, a deliberately broken migration, and an export-and-run-elsewhere exit test. Every tool gets the same six briefs, the same two testers of different technical backgrounds, three runs each. Nobody chooses their own test.

The difference this makes shows up most in brief four, the scope change. It is trivial to demonstrate an agent that writes a CRUD app well. It is much harder to demonstrate one that can hold a multi-tenancy decision in its head forty prompts later without contradicting itself, and vendor benchmarks essentially never test for it, because it does not produce a clean number for a landing page.

The trade-off we are making

Our approach is slower and covers fewer scenarios than a vendor's internal test suite, which can run thousands of tasks in parallel. We accept that trade because the six briefs are public, repeatable, and run the same way on every single row in this index. A reader who disagrees with our scoring can read exactly what we did and argue with the method, which is not possible with a benchmark whose task list was never published. That is the standard we are holding ourselves to as well, which is why every axis score traces back to a documented brief rather than a number we simply believed.