Skip to content
App Builder Index

News

The Index Now Covers Thirty-Eight Builders, and We Recomputed Every Score

What started as eleven tools at launch is thirty-eight today. Widening the index surfaced a bug in how we rolled axis scores into the overall rating, and we have corrected it across every row.

Owen Pryce · · Updated

The index launched with eleven builders and a promise: score everything on the same ten axes, publish the weights, and let the arithmetic decide the order. Thirty-eight builders in, that promise needed a stress test, and it failed one.

What went wrong

The overall star rating shown next to each tool is supposed to be the weighted mean of its ten axis scores, using the weights published on how we review: reliability and integrations at sixteen percent each, SEO and GEO at thirteen, design at twelve, agent performance at ten, speed and value at nine each, scalability at seven, API and MCP access at five, and code ownership at three.

As the index grew from eleven tools to thirty-eight, the overall figure stopped being derived from that formula and started drifting from it, tool by tool, as new rows were added by hand. The table and the methodology page were describing two different scoring systems while claiming to describe one. That is the kind of inconsistency that makes a review site's arithmetic unverifiable, which is the opposite of the point of publishing weights at all.

What we did about it

Every one of the thirty-eight rows now has its overall rating recomputed directly from its ten stored axis scores and the published weights, rounded to one decimal place. Nothing about the axis scores themselves changed as part of this pass, with one exception noted below. The order on rankings is now a direct, checkable consequence of the axis table beneath it: take any row, take the weights, do the multiplication, and you land on the number displayed.

The one exception is our own listing. Totalum's axis scores had not been revisited since the eleven-tool launch, while the tool itself had shipped a full quarter of improvements to reliability, scalability, agent performance and API access. We re-tested it against the same six briefs we run on every other tool and updated its axis scores to match what the current product does, not what it did at launch. Its speed score was not touched, and it remains behind Base44, Bolt.new and v0 on that axis specifically, because first-preview latency is a real, unresolved weakness we are not going to paper over with a score.

Why we are telling you this instead of quietly fixing it

A correction you do not disclose is not a correction, it is a cover-up with better arithmetic. The whole value of a published methodology is that a reader can catch us being wrong. Someone did the multiplication on our own numbers and it did not add up, which is exactly the failure mode a public weights table is supposed to make visible. We would rather publish this note than have the inconsistency sit there unexplained.

If you spot another row where the displayed rating does not match the weighted mean of its axis scores, the fastest way to reach us is the correction form on contact. We re-run the multiplication in front of you and fix it the same day.