Skip to content
App Builder Index

Guides

The true cost of credit-based pricing

We tracked every euro across six briefs on eleven platforms. The advertised price predicted the real bill in exactly three cases.

Owen Pryce · · Updated

Nine of the eleven builders in our index meter usage. The units differ, credits, tokens, messages, agent minutes, but the economics are identical: you buy a monthly allowance, and both successful and failed work draws it down.

We recorded actual spend on every test build. Here is what the data says.

Failed attempts are the entire story

On a good run, a medium-complexity feature cost between 4 and 11 EUR of usage across the platforms we tested. On a bad run, the same feature cost between 9 and 38 EUR. The spread is not about the feature. It is about how many times the agent tried and missed.

That produces a counter-intuitive result: the platforms with the cheapest headline price were not the cheapest in practice. A tool that costs less per attempt but needs three attempts is more expensive than one that costs more per attempt and lands first time.

Debugging is the most expensive activity

The pattern was consistent. Building new features consumed credits predictably. Fixing a subtle bug consumed them unpredictably, because the agent cannot tell in advance whether its next hypothesis is correct, and each hypothesis costs the same as a working change.

One of our reviewers spent 40 credits on a single row-level security policy. A developer would have fixed it in under ten minutes. That is not a criticism of the agent's intelligence, it is an observation about what happens when you meter a process that involves guessing.

The two flat-priced exceptions

Two platforms in our index charge a flat monthly fee with generous limits rather than metering agent work. Our testers behaved differently on those platforms, and the difference was visible in the logs: they tried more things, abandoned more experiments, and did not ask us whether a retry was allowed.

We are wary of drawing a large conclusion from a small sample, but the effect matches what every reviewer reported qualitatively. A meter changes how you build, and not in the direction of better software.

How to budget honestly

Take the advertised plan price. Now assume you will spend the same amount again on top-ups in any month where you are actively building. That approximation matched our real spend within about twenty percent on seven of the nine metered platforms.

If that number is unacceptable to your finance team, restrict your shortlist to the flat-priced options, or accept that you are buying an unbounded service and put a hard cap on the account.

What we want vendors to publish

Two things would fix most of this. First, cost per completed feature on a published reference brief, not credits per message, because nobody can convert one to the other. Second, a visible running total in the interface with a configurable ceiling.

Three of the eleven tools show a running total. None publish cost per completed feature. Until they do, our value axis is doing that work on the buyer's behalf, and the methodology behind it is on our how we review page.