How to enforce quotas server-side in Next.js
The record call is the lock, not the if-statement after it. How quota enforcement actually behaves on hard-cap and overage plans.
In short
Record usage with bb.usage.record() in the route handler. The record call is the enforcement: it increments atomically and rejects at the cap on hard-capped plans. The available <= 0 check from the docs snippet is not what stops the work, and it gives the wrong answer in all three situations it can meet.
The quota snippet in our own docs ends with a guard that never fires on a hard-capped plan, and fires wrongly on a plan with overage enabled:
if (usage.available <= 0) {
return Response.json({ error: 'Quota exceeded' }, { status: 429 });
}It looks like the enforcement. It is not. The enforcement already happened one
line earlier, inside record(), and what you do with the return value depends
on a plan setting that snippet does not mention.
This guide is the server half of enforcing quotas in React. That post gates what a user sees. This one is the part a user with devtools cannot skip.
Before you start
A quota item defined on the plan version, and a workspace on an active or
trialing subscription. The examples use the slug api-calls, matching
charging per API call.
Step one: record the usage
One call, in the route handler, before you do the expensive thing:
import BuildBase from '@buildbase/sdk';
const bb = BuildBase({
serverUrl: 'https://api.yourapp.com',
orgId: 'your-org-id',
getSessionId: async () => {
const cookies = await getCookies();
return cookies.get('bb-session')?.value ?? null;
},
});
export async function POST(req: Request) {
const usage = await bb.usage.record(workspaceId, {
quotaSlug: 'api-calls',
quantity: 1,
idempotencyKey: req.headers.get('x-request-id'),
});
return Response.json({ remaining: usage.available });
}What you should see: consumed goes up by one, and available comes back as
whatever is left. Fire the same request with the same x-request-id twice and
consumed moves once.
The increment is a single conditional update, not a read followed by a write. That matters more than it sounds. Two requests arriving together against the last remaining unit cannot both be told there is room, because there is no window between the check and the increment for the second one to land in.
Step two: handle the shape your plan actually has
Here is the part the snippet skips. Enforcement branches on whether the plan hard-caps, and a plan hard-caps when the subscription is trialing or when overage is disabled on that quota.
On a hard-capped plan, a rejection throws. It does not come back as
available: 0. The conditional update refuses to match, the service raises, and
the usage route returns HTTP 400 with the reason in the message:
try {
const usage = await bb.usage.record(workspaceId, {
quotaSlug: 'api-calls',
quantity: 1,
idempotencyKey: req.headers.get('x-request-id'),
});
return Response.json({ remaining: usage.available });
} catch (err) {
// Quota rejections arrive here, as a 400 with the reason in the message.
return Response.json({ error: 'Quota exceeded' }, { status: 429 });
}So the if never runs on these plans. The catch is the branch that matters,
and if you want your own callers to see a 429 you have to map it yourself,
because the platform does not send one.
On a plan with overage enabled, nothing throws at all. Usage past the
included amount is accepted and billed, and available is floored at zero. It
reports Math.max(0, included - consumed), so it sits at 0 for every request
after the cap. Keep the docs' guard here and you reject every overage call you
had already agreed to bill.
Key takeaway
available is computed after the increment. A call that consumes the last
unit succeeds and returns available: 0. Treating 0 as "this request was
refused" rejects the one request that was fine.
There is a third case worth knowing: quota.limit_exceeded fires only on the
first crossing into overage, not on every call past the cap. On a hard-capped
plan it never fires at all, because usage there can never exceed the included
amount. If you are waiting on that event to email someone, a trialing workspace
will wait forever.
Step three: make the retry safe
idempotencyKey is the difference between a retry and a double charge. A replay
is matched on workspace, quota slug and key together, and a match returns the
current status with used: 0 rather than incrementing a second time.
Because the slug is part of the match, one request id can safely record against several quotas:
await bb.usage.record(workspaceId, {
quotaSlug: 'api-calls',
quantity: 1,
idempotencyKey: requestId,
});
await bb.usage.record(workspaceId, {
quotaSlug: 'storage-mb',
quantity: 5,
idempotencyKey: requestId, // same key, different quota, both land
});Leave the key off and every retry your queue makes is a unit your customer pays for. Use something that is stable across retries of the same logical operation, not a fresh uuid per attempt.
At volume, stop sending one call per unit
The usage endpoint is rate limited to 60 requests per minute. Record one unit per API call synchronously and that limit, not your quota, becomes the thing that breaks first.
recordBatch takes up to 100 items per request and returns a result per item,
so partial failures are visible rather than silent:
const result = await bb.usage.recordBatch(workspaceId, {
items: [
{ quotaSlug: 'api-calls', quantity: 1, idempotencyKey: 'req-1' },
{ quotaSlug: 'storage-mb', quantity: 5, idempotencyKey: 'upload-1' },
],
});This is the shape any per-run product ends up needing. PlugNode bills per flow run rather than per seat, and a single run that fans out across steps produces far more than one billable event, so the batch endpoint is the only version of this that holds.
What breaks, and how you will know
A 400 you read as a bug. Quota rejection and "quota slug not on this plan version" both return 400 with different messages. Read the message before assuming the limit was hit. Credits get a machine-readable code. Quota rejections do not, so the message string is what you are left matching on.
Overage that bills late. Stripe sync is asynchronous. If the enqueue fails, the record stays unbilled until the sweep picks it up, and that sweep runs every 10 minutes - so a usage row that has not reached Stripe yet is late, not lost.
Counting the work you did not do. Recording before the operation means a failure after the record still consumed the unit. For cheap idempotent work that is the right trade. For an expensive job, record after it succeeds and accept that a caller can start one more than the plan allows.
A quota that resets when you did not expect. Allowances reset monthly whatever the billing interval, so a yearly subscription still gets a monthly reset rather than one annual pool. That is the opposite of what most people assume from an annual plan.
If you are deciding whether this should be a quota at all, the distinction from a balance you buy is in quotas are not credits.
Install
npm i @buildbase/sdk