A 429 when the quota should not be exhausted
A 429 on the usage endpoint is almost never the quota. Quota rejection returns 400. Here is how to reconcile the counter you read with the counter being enforced.
In short
The counter your dashboard reads and the counter enforcement increments are produced by two different code paths. A 429 on the usage endpoint comes from the rate limiter at 60 requests per minute per IP, not from the plan allowance. Quota rejection returns 400 and credits return 402.
The counter you are reading is not the counter being enforced. That sounds like a bug report, but it is the design, and once you see it the 429 stops being mysterious.
Two different code paths produce the two numbers. GET /usage/status and
GET /usage/all compute what the dashboard shows. POST /usage increments
what the plan enforces. They read the same field on the same workspace
document, and they still disagree, for about five reasons that are all
legitimate.
Start with the one that resolves most cases: a quota rejection does not
return 429. It returns 400. The usage route wraps the whole handler in a
try/catch and answers every failure with res.status(400), which is the
behaviour already documented in
enforcing quotas server-side.
So if you are holding a 429, something other than your plan allowance said no.
Before you start
A workspace on a plan version that defines the quota slug, and the ability to
read the raw HTTP response rather than only the SDK error. The SDK throws a
plain Error carrying the server's message string, with no status code
attached, so the status has to come from a fetch.
The status code names the ceiling
Three ceilings sit on the same request, and each one answers with a different code:
| Code | Who said no | Where it lives |
|---|---|---|
| 429 | the rate limiter | middleWares/rateLimit/index.ts |
| 400 | the plan quota | the usage route's catch block |
| 402 | the credit balance | InsufficientCreditsError |
The 402 is a separate system entirely. If you are debiting a wallet rather than metering an allowance, the distinction is the whole subject of quotas are not credits.
Because the SDK surfaces the message and drops the status, the message string is your real signal, and the two candidates do not look alike:
- Rate limiter:
Too many usage recording requests. Please slow down. - Quota:
Quota limit reached for 'api-calls'. Included: 250, used: 250, requested: 1, available: 0.
The rate limiter is also built with standardHeaders: true, so a 429 from it
carries RateLimit-Limit, RateLimit-Remaining and RateLimit-Reset. A
response with those headers was not a quota decision. This is the fastest check
you can run, and I have not run the snippet below as written, so treat it as the
shape rather than a tested script:
const res = await fetch(
`${serverUrl}/api/v1/public/workspaces/${workspaceId}/subscription/usage`,
{
method: 'POST',
headers: {
'content-type': 'application/json',
'x-session-id': sessionId,
},
body: JSON.stringify({ quotaSlug: 'api-calls', quantity: 1 }),
}
);
console.log(res.status, res.headers.get('ratelimit-remaining'));
console.log(await res.json());Key takeaway
A 429 with RateLimit-Remaining: 0 is the limiter. A 400 whose message begins
"Quota limit reached for" is the plan. They are different ceilings with
different units, different windows and different keys.
The rate limit counts requests; the quota counts units
This is where most of the confusion actually starts. A quota is a plan allowance: units, per workspace, reset monthly. A rate limit is throughput: requests, per minute, keyed on the caller's IP address.
The usage endpoint is limited to 60 requests per minute. Credit consume gets 30. Above both sits a global limiter at 500 requests per 3 seconds, and a separate per-key limiter at 1000 per minute that only engages when the request carries an org API key.
Per IP is the part that bites. Every workspace your backend records usage for shares one bucket, because they all leave through the same address. A hundred customers on a plan with room to spare will still collide at 60 requests a minute if your server records one unit per call.
We have been bitten by exactly this shape internally. The token exchange endpoint used to be capped at 10 requests per minute per IP, which sounds strict until you notice it is a server-to-server endpoint, so the per-IP ceiling was really a ceiling on one integration's entire throughput. Our own provisioning makes up to three exchanges per signup. Ten a minute meant about three signups a minute before the platform started refusing its own work. It is 120 now, tunable by env var, and the code comment explaining why is longer than the change.
One request is not one unit
POST /usage/batch takes up to 100 items and shares the same 60-per-minute
bucket as the single-record route. That is the intended escape hatch: one
request against the limiter, up to 100 units against the quota.
It also changes where failures appear. The batch handler runs items through
Promise.allSettled and returns HTTP 200 with a per-item result array, so a
quota rejection inside a batch is not a 400 at all. It is a success: false
object sitting in the results array of a 200 response, next to a failed
count. If you are only checking res.ok, you will record 100 items, get a 200,
and never see that 40 of them were refused.
And an item's quantity has a minimum of 1 but no maximum on the public route,
so a single item can consume as much of the allowance as it likes. Your
dashboard number can move by 5,000 on one request that cost you one unit of rate
limit.
A reset that has not landed, or will never land
Allowances reset monthly whatever the billing interval, which the enforcement guide already states. What it does not say is when.
The reset is anchored on usageAnchorDay, the day of the month the subscription
was created, not the 1st and not the invoice date. The repo's own worked example:
a subscription created on Jan 31 resets on Jan 31, Feb 28, Mar 31, Apr 30. The
short month clamps, and the month after goes back to the anchor rather than
cascading down. A sweep cron runs every 5 minutes, offset by one, at :01, :06,
:11 and so on, so a reset lands within about five minutes of its boundary and
not at the stroke of it.
Then there is the case that produces a counter stuck at the cap for days. Reset
eligibility is limited to active and trialing. Entitlement, the check that
decides whether usage may be recorded at all, allows active, trialing and
past_due. Read those two sets side by side and the gap is obvious: a past_due
subscription can keep recording usage, and will never be reset. It fills to the
included amount once, and then every call is refused until the payment recovers.
Nothing is broken. The counter is just correct and permanently full.
The replay that did not deduplicate
An idempotency key is matched on three fields together: workspace, quota slug and key. Change any one of them and it is a different operation, and the replay increments again.
That is deliberate for the slug, because it lets one request id record against several quotas in the same operation. It is a trap for the workspace, because a retry that resolves the workspace differently, say from a cached session rather than the path parameter, is a new match key and a second unit.
There is a narrower path worth knowing about. The increment and the log row are not one transaction. The increment happens first, and the log row that stores the idempotency key is written inside a try/catch that only logs on failure. If that write fails, the unit is spent and no record of the key exists, so the retry finds nothing to match and spends another. Rare, and it shows up as a counter two units ahead of the work you actually did.
Keys are capped at 255 characters, and the usual real cause is much simpler than any of the above: a fresh uuid per attempt instead of one stable id per logical operation.
You may be reading a different counter than you think
Usage lives at quotas.<slug> on the workspace document. Not on the org, not on
the user. If your app resolves a workspace from a session on the read and from a
path parameter on the write, you are looking at two documents.
Two more writers touch that same field. The admin path, PATCH /workspaces/:id/quota/usage, sets values absolutely rather than incrementing,
and when a Stripe record fails it falls back to writing the value locally. The
internal Stripe route takes workspaceId from the request body, while the
public route takes it from the URL. Same field, four ways in.
The read and the write also do not resolve your subscription the same way.
getQuotaStatus calls getWorkspaceSubscription, which returns any record that
is not canceled, including paused, incomplete and unpaid. Recording calls
getEntitledSubscription, which also rejects a dunning-suspended record
and a trialing record whose trialEnd has passed. So a workspace with an
expired trial that still says trialing will show you a full allowance on the
dashboard and throw No active subscription found on the record. A 400, and not
one word about quotas in it.
Triage, in the order that settles fastest
Read the status code. 429 is the limiter, 400 is the quota or the subscription, 402 is the credit balance. Then read the message, because 400 covers quota rejection, an unknown slug and a missing subscription with three different strings. Then check which workspace id the failing call actually sent, which is where the dashboard-versus-enforcement gap usually turns out to live.
If the answer is the limiter, batching is the fix rather than a higher ceiling, and the batch shape is covered in the server-side enforcement guide. If the answer is 402, you are on the credits path and the balance is not one number, which the credits wallet guide explains.
Install
npm i @buildbase/sdk