Guides
Build guide

Charge per credit in Next.js, and why the debit goes last

Charge per credit in Next.js: check the balance before the work, debit after it, return 402 when the balance is empty, and pass an idempotencyKey so a retry is not charged twice.

Dharmendra Jagodana8 min read

In short

A credit debit cannot be undone from application code, so in Next.js the consume call goes after the work succeeds, not before it. Check the balance first to avoid doing free work, catch INSUFFICIENT_CREDITS and return 402 rather than 429, and pass an idempotencyKey so a retry is not charged twice.

The SDK's bb.credits surface has eight methods on it. Not one of them undoes a spend. You can read a balance, consume, purchase, list your packages, list the public ones, list transactions, list buckets, read what is expiring. There is no refund, no reverse, no way to put a consumed credit back from your own application code.

That single absence decides the shape of every per-credit route you will write. If the debit cannot be taken back, it has to happen after the work it pays for. This guide is about that debit. Where the credits come from in the first place - packages, grants, top-ups and what expires when - is the other half.

This is where charging per credit stops resembling charging per API call. A metered call records one unit of something that already happened. A credit prices work that has not happened yet, at an amount you chose, and the work can fail after the money is gone. If you are still deciding which of the two you are selling, quotas are not credits is the argument for picking one.

Before you start

A BuildBase project with the credits module and at least one credit package defined, so a workspace has a balance to spend. Credit pricing is yours to set per org, so every amount in this guide is illustrative.

Step one: two calls, with two different jobs

The route needs a cheap check before the work and an authoritative debit after it. Those are not redundant. The check stops you burning compute nobody can pay for, and the debit is the one that actually moves the balance.

// app/api/generate/route.ts
import BuildBase from '@buildbase/sdk';

const bb = BuildBase({ serverUrl, orgId, getSessionId });

const COST = 10;

export async function POST(request: Request) {
  const { workspaceId, prompt } = await request.json();

  // Advisory. Cheap, and it stops you doing work nobody can pay for.
  const balance = await bb.credits.getBalance(workspaceId);

  if ((balance?.available ?? 0) < COST) {
    return Response.json(
      { error: 'Insufficient credits', available: balance?.available ?? 0 },
      { status: 402 }
    );
  }

  const output = await generate(prompt);

  // Authoritative. The balance moves here, after the work succeeded.
  const debit = await bb.credits.consume(workspaceId, {
    amount: COST,
    description: 'AI generation',
  });

  return Response.json({ output, remaining: debit.balanceAfter });
}

What you should see: call it once and the balance drops by ten. Make generate throw and the balance does not move, because the line that moves it never runs.

Report the balance from balanceAfter on the consume response, not from the advisory read minus the cost. The first is what the ledger recorded. The second is arithmetic on a number that was already out of date when you read it.

So what about the gap between the check and the debit? It is not a hole you need to close. On the server the debit lands as one conditional update that requires the balance to still cover the amount, so two requests racing for the last credit cannot both win. The check is a courtesy to your own infrastructure. The debit is the gate.

Step two: the empty balance is a 402, and the 429 means something else

consume throws rather than returning a falsy result, and the error carries what you need to say to the user.

try {
  await bb.credits.consume(workspaceId, {
    amount: COST,
    description: 'AI generation',
  });
} catch (err) {
  if (err.code === 'INSUFFICIENT_CREDITS') {
    return Response.json(
      {
        error: err.message,
        available: err.available,
        requested: err.requested,
      },
      { status: 402 }
    );
  }
  throw err;
}

err.available and err.requested are both on the error, so the response can say "you need ten and have three" rather than "something went wrong". That is a support ticket you never receive.

Use 402 here, not 429. Payment required sends the user to a screen with a buy button; too many requests sends them to a screen that tells them to wait, and waiting will not help. The distinction matters more than it sounds, because this route can genuinely return both: the consume endpoint is rate limited at 30 requests per 60 seconds per IP, and that limiter answers with a 429 of its own. A 429 from this route is never an empty balance.

Check where that limit lands before you size anything around it. Called from a Next.js route handler, the IP the limiter sees is your own server's, not the end user's, so all your tenants share one bucket of thirty.

Key takeaway

402 means buy more credits. 429 from the same endpoint means you are calling it too fast. Treating them as one error sends a paying customer to the wrong screen.

Step three: make the retry safe with a key, not a try/catch

A client that retries is the normal case, not the edge case. Without a key, the second attempt is just new work. It charges again.

await bb.credits.consume(workspaceId, {
  amount: COST,
  description: 'AI generation',
  idempotencyKey: request.headers.get('x-request-id') ?? undefined,
});

The ledger carries a unique index on the workspace and the key together. A repeat of a key it has already seen returns the original outcome and leaves the balance alone.

Two caveats before you rely on it. The replayed response echoes the amount you asked for rather than what the original call deducted, so treat a replay as confirmation that the debit happened, not as a fresh reading of the balance. And a debit large enough to span several credit buckets stamps the key on the first bucket's ledger row only. The guard still holds, because the guard is the index on that one row. The ledger is just less uniform than you might assume when you go reading it later.

The key has to be stable for the logical job. A fresh uuid per attempt is not an idempotency key, it is a new charge with extra steps.

Step four: surface the balance before it runs out

Server-side enforcement is the part that counts, and it is also the part the user never sees until it stops them. The React side exists to make sure zero is not a surprise.

import {
  useCreditBalance,
  useExpiringCredits,
  WhenCreditsExhausted,
  WhenCreditsLow,
} from '@buildbase/sdk/react';

function GenerateButton({ workspaceId }) {
  const { balance } = useCreditBalance(workspaceId);
  const { expiringCredits } = useExpiringCredits(workspaceId, 7);

  return (
    <>
      <p>Credits: {balance?.available}</p>
      {expiringCredits > 0 && <p>{expiringCredits} credits expire soon</p>}

      <WhenCreditsLow threshold={10}>
        <p>Running low. Buy more credits.</p>
      </WhenCreditsLow>

      <WhenCreditsExhausted>
        <p>No credits left.</p>
      </WhenCreditsExhausted>

      <button>Generate (10 credits)</button>
    </>
  );
}

threshold on WhenCreditsLow is required, which is the right call: there is no sensible default for a number you price yourself. useExpiringCredits handles the other failure, a balance that lapses rather than empties. The seven-day window above is the difference between an expiry and an accusation.

One thing to get right: the number and the guards refresh by different routes. useCreditBalance is a plain fetch hook, so after a debit your route made it holds a stale figure until you call its own refetch. invalidateCreditBalance() does not touch it. That function refreshes the credit-balance context, which is what WhenCreditsLow and WhenCreditsExhausted read from. So one call updates the gate and the other updates the number, and a UI that only does one of them will cheerfully show the old balance to someone who just spent. Import both from @buildbase/sdk/react. The same reasoning applies to gating the control itself, which has its own guide.

When you cannot debit after the work

Everything above assumes you can await the work inside the request. Queue it instead, or hand it to a job that runs for four minutes, and the choice disappears: you have to take the debit up front, because the request is over before anyone knows whether the work succeeded.

There is no clean answer to this in the SDK today, and I would rather say so than invent one. Nothing reverses a consumed credit from application code. Two paths put credits back deliberately, and neither is a reversal. The admin grant from the console needs a credit package, so it returns that package's amount and not the amount you debited. The workflow grant_credits action has a worse problem. It passes no idempotency key and runs up to three attempts, so a failure handler can grant twice.

Both gaps look closable to me, and naming them beats pretending the pattern is clean: either consume grows a compensating reversal, or the grant path gets exposed with an arbitrary amount and an idempotency key. Until one of those lands, the workable shape for queued work is to debit up front, emit your own failure event, and reconcile from the console.

The failure modes worth knowing

Not every workspace member can spend. The consume endpoint requires a usage-record permission on top of workspace membership. Admins and members hold it by default. The read-only viewer role does not, which is what stops a read-only seat draining the balance. Purchasing credits needs billing manage and reading the balance needs billing view. If credits map to real money in your product, it is still worth putting the authorization you want in your own route, above the SDK call.

The low-balance event is not a threshold you set. A credit.low_balance event fires automatically, at twenty percent of the amount you just debited. So a ten-credit debit alerts at two or under. A five-hundred-credit debit alerts at a hundred or under. Not at zero, though: the event needs a balance above it, so the debit that empties the account warns nobody. Per call, not per workspace, and there is no setting for any of it. The threshold prop in the React guard is a separate client-side thing that happens to share the word.

You cannot email on every consume. credit.consumed is a system event and a webhook event, but it is deliberately absent from the notifiable set. Purchases, expiries, grants, revocations and low balances can notify a user. Every single spend cannot, which is correct for anything billed by the thousand.

Trusting the client with the amount. amount comes from your server, priced by the work you just did. Take it from the request body and your price list is a suggestion.

Doing the work twice to avoid charging twice. The instinct when a debit fails is to retry the whole route. Read the error first. A timeout or a dropped connection is worth retrying, and retrying the consume alone with the same key is the cheap fix, because the work is already done and only the debit is missing. An INSUFFICIENT_CREDITS error is not worth retrying at all: the balance will not have changed, so the second attempt fails the same way. That one is a 402 and a buy button, not a retry.

Whether any of this is cheaper than building the ledger, the buckets and the idempotency guard yourself is a separate question, and the cost of DIY usage metering does that arithmetic.

Install

npm i @buildbase/sdk
billing
credits
nextjs

Frequently Asked Questions

Should I deduct credits before or after the operation?

After, whenever you can await the work inside the request. There is no call in the SDK that reverses a consumed credit, so a debit taken before work that then fails has to be corrected by hand from the console. Check the balance before the work, debit after it.

What should I return when a workspace is out of credits?

A 402. The consume call throws with code INSUFFICIENT_CREDITS and carries available and requested, which is enough to tell the user the exact shortfall. A 429 from the same route means something else: the consume endpoint is rate limited to 30 requests per 60 seconds per IP.

Does an idempotencyKey stop a retried request being charged twice?

Yes. The ledger has a unique index on the workspace and key together, so a repeat of the same key returns without touching the balance, though the replayed response echoes the requested amount, not a fresh balance. Pass a stable id for the logical job, not a fresh one per attempt, or every retry looks like new work.

How do I warn someone before their balance hits zero?

WhenCreditsLow takes a required threshold and renders when the balance is at or below it. useExpiringCredits(workspaceId, 7) covers the other case, a balance that is about to lapse rather than run out. Warning at a threshold sells a top-up; blocking at zero loses the session.

Ship it

Create a project and run this guide against your own workspace.

7-day free trialNo credit card requiredCancel anytime