Blog
Explainer

What is usage-based billing, and what it does to your invoices

Usage-based billing charges for what a customer consumed rather than for access. The definition is easy. Defending the number on the invoice is the actual product.

Dharmendra Jagodana6 min read

In short

Usage-based billing charges a customer for what they consumed in a period rather than a flat fee for access: API calls, tokens, renders, gigabytes. It moves the hard problem from pricing to counting, because every line on the invoice is a claim about something that already happened.

Usage-based billing charges a customer for what they consumed in a period rather than a flat fee for having access. The unit is whatever you can count: API calls, tokens, renders, gigabytes held overnight.

That definition takes ten seconds and it is where most explanations of this stop. The interesting part is what the model does to the rest of your company, because an invoice that varies is an invoice a customer reads. Nobody audits a $99 line that has been $99 for a year. Everybody audits the month it says $412.

Why the model exists at all

Flat pricing asks every customer to pay the average. That works while your customers look alike and stops working the moment one of them is thirty times heavier than another, which is the normal shape of an API product and the guaranteed shape of anything doing inference.

So the model exists to solve a fairness problem and a margin problem at once. The heavy customer pays for being heavy, the light one is not quietly subsidising them, and your cost of goods tracks your revenue instead of diverging from it. That is the whole pitch, and it is a good one.

What the pitch leaves out is that you have just made your revenue a function of something you have to measure correctly, in production, on the first try.

How it actually works

Three pieces, and only the first is interesting.

Recording. Something happens, and you write it down against the customer and a unit. This has to be the same call that does the work:

const usage = await bb.usage.record(workspaceId, {
  quotaSlug: 'api-calls',
  quantity: 1,
  idempotencyKey: req.headers.get('x-request-id') ?? undefined,
});

Two details in there carry most of the weight. Recording and checking are one round trip, so there is no window where fifty concurrent requests each read "plenty left" before any of them is counted. On a hard-capped plan record() itself rejects at the cap, so there is no guard to write after it; server-side quota enforcement covers what to do with the return value. And the idempotency key means a retry does not bill twice, which stops being optional the first time you put a queue in front of this.

Aggregating. Sum the records for the period. Easy, until someone asks what happens to an event recorded at 23:59:58 on the last day, and you discover your period boundary and your customer's timezone disagree.

Invoicing. Turn the sum into money. Also easy, and also where the arguments land, because this is the first time the customer sees any of it.

The part the generic explainers skip

An invoice under this model is a claim about the past. You are telling a customer that 41,200 things happened last month and that they owe you for them. They cannot check it. They have to take your word, or take their own logs and find that the two do not match, which they will not, because your logs and their logs were never counting the same thing.

So the real requirement is not accurate metering. It is defensible metering, which is a stronger property with three parts:

  • Attributable. Every recorded unit traces to a request the customer can recognise. A count with no drill-down is an assertion.
  • Idempotent. A retried or replayed event does not double. Without this, your own reliability engineering inflates the bill.
  • Visible before the invoice. The customer sees the running number during the period, not after it. A surprise at the end of the month is a dispute even when the number is correct.

Miss the third one and you will be arguing about the first two in an email thread with someone who has already decided you are wrong.

Key takeaway

Usage-based billing moves the hard problem from pricing to counting. The invoice is a claim about events that already happened, so the system's real job is to make that claim attributable, idempotent, and visible before it becomes a bill.

Picking the unit

The unit is the decision that outlives everything else, because changing it later means re-pricing every customer.

A good unit has three properties. The customer can predict it before acting - "per image" passes, "per compute-second" mostly does not. It survives your own architecture changing, which is why "per API call" ages better than "per database query". And the cost to you moves with it, otherwise you have built variable pricing on a fixed cost and rediscovered flat pricing with extra steps.

Then there is the question everyone defers: does a failed request bill? There is no correct answer. There is only the version you wrote down before launch and the version you decided during a support ticket, and the first is much cheaper.

When it is the wrong choice

Usage pricing is a bad fit when the buyer needs a number for a purchase order. A procurement process cannot approve "roughly $600 a month, probably". Some of the largest deals you will ever sign require a fixed figure, and a pure usage model either loses them or forces you to invent a committed-spend tier.

It is also wrong when the customer cannot control their own usage. Charging by seats is charging for a decision your customer makes. Charging by webhook deliveries is charging for a decision their traffic makes, and a customer who gets a bill they could not have prevented does not conclude that the pricing is fair.

And it is wrong early. A product with eleven customers does not have a usage distribution, it has eleven anecdotes. Flat pricing for the first year is not a failure of nerve, it is declining to optimise something you cannot see yet.

We price ourselves accordingly, and it is worth saying out loud since we sell the metering: BuildBase plans are flat. Launch is $49 a month with 25,000 MAU included, Grow is $99, Scale is $199, and the only line that meters is storage, at $0.10 per GB past the plan allowance. Everything else is a hard allowance that stops rather than bills over. We build usage-based billing for products where the unit is obvious; ours is not one of them, and pretending otherwise would make our own pricing page worse.

What it is not

Usage-based billing is not the same as a quota, and it is not the same as credits, although it is usually implemented with one or both. A quota is an allowance that resets; a credit is a balance that carries. Which you need is a separate decision with its own consequences, laid out in quotas are not credits.

It is also not entitlements. Whether a workspace may call an endpoint at all is a different question from how many times it did, even though both end up as a condition around the same component in the UI - see entitlements are not feature flags.

If you are weighing the model against building the meter yourself, the cost of building usage metering yourself runs that arithmetic. And the route handler that puts all of this into one Next.js endpoint is charging per API call.

billing
metering
pricing

Frequently Asked Questions

What is usage-based billing?

Usage-based billing charges a customer for the quantity they consumed during a billing period rather than a fixed fee for access. The unit can be anything you can count reliably - API calls, tokens processed, images rendered, gigabytes stored - and the invoice is the sum of recorded events for that period.

What is the difference between usage-based and subscription billing?

A subscription charges the same amount every period regardless of consumption. Usage-based billing charges a variable amount derived from recorded events. Many products run both at once: a recurring fee that includes an allowance, then a usage charge for whatever goes past it.

What counts as a billable unit?

A good unit has three properties: the customer can predict it before acting, it survives your own architecture changing, and your cost moves with it. Then decide whether a failed request bills, and write that down before launch rather than during a support ticket.

Does usage-based billing need metering to be real-time?

Recording does. Invoicing does not. The record has to happen in the same call that does the work, otherwise a burst of requests all pass the check before any of them is counted. Showing the customer their current usage can lag by minutes without anyone minding, as long as the lag is not a surprise.

Put this into practice

The modules in this post are one console away. 7-day free trial, no credit card.

7-day free trialNo credit card requiredCancel anytime