Usage metering middleware in Express: record before the handler or after the response (Hono and Fastify too)
Usage metering middleware for Express, Hono and Fastify: when to record before the handler, when to batch after the response, and the rate limit that decides it.
In short
Metering middleware in Express can record before the handler or after the response. Recording first enforces the quota but counts requests that fail, one call each against a 60-a-minute limit. Recording after a 2xx and flushing with recordBatch bills only successes at volume, but cannot stop a request at the cap.
Write the obvious metering middleware for an Express API, one
bb.usage.record() per request, and you have set the throughput of every
metered route you own to 60 requests a minute.
That is not your quota. It is the rate limit on BuildBase's usage endpoints, counted per IP, and every usage call comes from your server's one address. All your users share it. It breaks first.
So the decision is not where the middleware goes. It is which of two jobs the call does on this route:
- Mode one: record before the handler. The record call is also the quota check, so it can refuse the request. At one call per request, it suits expensive routes that stay well under 60 a minute, like report generation.
- Mode two: record after the response, in batches. Only successful responses are billed, and 100 of them go in one call. Nothing is refused per request, because by the time the batch runs, the responses are gone.
You cannot have both on one call. Most APIs want both, on different routes.
Before you start
An Express, Hono or Fastify app with sign-in working through its setup guide
(Express, Hono,
Fastify), and the workspace resolved per
request as in multi-tenant workspaces in
Express, which puts it on
req.workspace. In the console, a plan with the quotas you meter, here
reports and api-calls.
What the middleware knows, and when
Before the handler, middleware cannot know whether the work will succeed. After the response, it knows the status code and can no longer do anything about the request. Each mode gives up one half.
Our per-API-call guide says metering in
middleware is "wrong for anything expensive" because the middleware "does not
know whether the work succeeded". For mode one that is exactly right. A request
that fails after the record has still been counted, and bb.usage has no way
to take it back. quantity has a minimum of 1, so there is no negative record
either.
Mode one: the record is the gate
// middleware/meter.js
const crypto = require('node:crypto');
const buildbase = require('../config/buildbase');
/** Refuses the request when the quota is spent. One usage call per request. */
exports.meterFirst =
({ quota, quantity = 1 }) =>
async (req, res, next) => {
const key = req.get('Idempotency-Key') ?? crypto.randomUUID();
if (key.length > 128) {
return res.status(400).json({ message: 'Idempotency-Key is too long.' });
}
try {
await buildbase.forRequest(req).usage.record(req.workspace.id, {
quotaSlug: quota,
quantity,
idempotencyKey: `${quota}:${key}`,
source: 'api',
});
return next();
} catch (err) {
if (err.status === 400 && err.message.startsWith('Quota limit reached')) {
return res
.status(402)
.json({ message: `This workspace has used its ${quota} allowance.` });
}
if (err.status === 429) {
// BuildBase's rate limit, not the customer's quota.
res.set('Retry-After', '60');
return res
.status(503)
.json({ message: 'Usage metering is busy. Try again shortly.' });
}
return next(err);
}
};// routes/api.js
router.post('/reports', meterFirst({ quota: 'reports' }), reports.create);On a plan that is trialing, or a quota with overage switched off, record is
an atomic check-and-increment, so two requests racing for the last unit cannot
both get it. Server-side quota enforcement
walks through that, and why a separate "is there any left?" read is the wrong
check. With overage on, record never refuses. It bills the overage and the
request goes through.
Three choices in there are ours:
402 for a spent quota, not 429. A 429 tells a well-behaved client to back off and retry. Retrying a spent allowance fails until the period resets, so we answer with the status that says "this needs paying for".
503 when BuildBase rate-limits us. It fails closed. Doing the work unmetered would give away exactly the expensive requests this mode exists for. And the status is ours to pick, because the upstream 429 is about our server's call volume, not the customer.
The client's Idempotency-Key, when it sends one. A retried request with
the same key is counted once: the server matches the key on workspace, quota
and key, and a replay returns used: 0. Without a header we mint a UUID,
which makes the record safe to retry inside this request but treats a client
retry as a new request. A server-rendered form can hand out its key when the
page renders instead, as usage-based billing in React
Router does from the loader. The cap of 128 characters keeps the prefixed key under
the 255 the API accepts.
To check it, set the reports quota to 2 with overage off and send three
requests. The third answers 402.
Mode two: record after the response, in batches
The middleware only notes what happened. A buffer sends it later.
// middleware/meter.js (continued)
const usageBuffer = require('../lib/usage-buffer');
/** Bills 2xx responses only. Never refuses a request. */
exports.meterAfter =
({ quota, quantity = 1 }) =>
(req, res, next) => {
const clientKey = req.get('Idempotency-Key');
// A long client key is not truncated: two truncated keys could collide.
const key =
clientKey && clientKey.length <= 128 ? clientKey : crypto.randomUUID();
const workspaceId = req.workspace.id;
// Only a session that may record usage is worth keeping for the flush.
const sessionId = req.workspace.permissions.includes(
'workspace:usage:record'
)
? req.session.buildbaseSessionId
: null;
res.on('finish', () => {
if (res.statusCode < 200 || res.statusCode >= 300) return;
usageBuffer.add(workspaceId, sessionId, {
quotaSlug: quota,
quantity,
idempotencyKey: `${quota}:${key}`,
source: 'api',
});
});
next();
};Why only 2xx? A 4xx is the client's mistake and a 5xx is ours. Neither
belongs on an invoice. Node's docs are precise about what finish
means: the response "has been sent", which "does not imply that the client has
received anything yet". For billing, sent is the right line.
// lib/usage-buffer.js
// The one BuildBase call, set once at startup: the Express app passes the
// SDK, a Fastify app its gateway. Nothing else in this file knows which.
/** @typedef {(sessionId: string, workspaceId: string, items: Array<{ quotaSlug: string, quantity: number, idempotencyKey: string, source?: string }>) => Promise<unknown>} Sender */
/** @type {Sender | null} */
let send = null;
/** @param {Sender} fn */
exports.setSender = (fn) => {
send = fn;
};
const FLUSH_MS = 60_000;
const BATCH_LIMIT = 100; // the most recordBatch takes in one call
/** workspaceId -> items waiting to be sent */
const pending = new Map();
/** workspaceId -> the latest session from a member who may record */
const sessions = new Map();
/** Workspaces with a batch in flight, so two flushes never overlap. */
const flushing = new Set();
exports.add = (workspaceId, sessionId, item) => {
if (sessionId) sessions.set(workspaceId, sessionId);
const items = pending.get(workspaceId) ?? [];
items.push(item);
pending.set(workspaceId, items);
if (items.length >= BATCH_LIMIT) flush(workspaceId);
};
async function flush(workspaceId) {
const sessionId = sessions.get(workspaceId);
const queued = pending.get(workspaceId);
if (!send || !sessionId || !queued?.length) return; // nothing to do yet
if (flushing.has(workspaceId)) return; // the next tick picks the rest up
flushing.add(workspaceId);
const items = queued.slice(0, BATCH_LIMIT);
pending.set(workspaceId, queued.slice(BATCH_LIMIT));
try {
await send(sessionId, workspaceId, items);
} catch (err) {
if (
err.status === undefined &&
err.message === 'Failed to record batch usage'
) {
// The server refused some items (usually a hard cap) and recorded the
// rest.
// A resend would be refused again, so drop the batch and say so.
console.warn('usage partly refused', workspaceId, items.length);
} else {
// A 429, a 5xx or a network error: it may or may not have landed.
// Put the items back; their keys make the resend safe either way.
pending.set(workspaceId, [...items, ...(pending.get(workspaceId) ?? [])]);
console.error('usage flush failed', workspaceId, err.message);
}
} finally {
flushing.delete(workspaceId);
}
}
exports.flushAll = () => Promise.all([...pending.keys()].map(flush));
/** For shutdown: keep flushing until nothing is left, or the deadline passes. */
exports.drain = async (deadlineMs = 10_000) => {
const stop = Date.now() + deadlineMs;
while (Date.now() < stop) {
const left = [...pending.entries()].some(
([id, items]) => items.length && sessions.has(id)
);
if (!left && flushing.size === 0) return;
await exports.flushAll();
await new Promise((r) => setTimeout(r, 100));
}
};
setInterval(exports.flushAll, FLUSH_MS).unref();In the Express app, the sender is the SDK. forSession is one line added to
config/buildbase.js from the setup guide, next to forRequest:
// config/buildbase.js (next to forRequest)
exports.forSession = (sessionId) => buildbase().withSession(sessionId);
// app.js, once at startup
const buildbase = require('./config/buildbase');
const usageBuffer = require('./lib/usage-buffer');
usageBuffer.setSender((sessionId, workspaceId, items) =>
buildbase.forSession(sessionId).usage.recordBatch(workspaceId, { items })
);// routes/api.js
router.get('/search', meterAfter({ quota: 'api-calls' }), search.run);Send a few searches and nothing reaches BuildBase for up to a minute. Then one
recordBatch lands with all of them, and api-calls in the console moves by
the number of 2xx responses, not the number of requests.
Which session sends the batch? The requests are over, so there is no current
user, and the server SDK takes no API key: every call it makes runs as some
user's session. The buffer keeps, per workspace, the most recent session
from a member whose permissions include workspace:usage:record, which the
workspace middleware already
resolved. A viewer's session would only earn a 403, so it is never kept.
This is the weakest part of the design, and we would rather say so. A kept session can be signed out, or its member demoted, before the flush. That shows up as a rejected batch while the items wait for the next request from a member who may record. Usage from a workspace whose only active users are viewers waits until one comes along. An API token can be exchanged for a session, but that session's user still has to be a member of the workspace who may record, so it does not remove the problem for a multi-tenant API.
The arithmetic of the flush interval
A batch carries one workspace and counts once against the 60-a-minute limit. Flush every N seconds, with W workspaces active, and you make W × 60 / N calls a minute. Keep that at or under 60 and the rule falls out:
The flush interval, in seconds, has to be at least the number of workspaces active during it.
At the 60-second interval above, that holds for up to 60 busy workspaces at a
time. And the budget is shared: every mode-one record spends from the same
60, and so does every other process behind the same outgoing IP. Two API
servers behind one NAT halve it. Past that, lengthen the interval and accept
that usage shows up later.
A workspace that fills 100 items early flushes early and spends an extra call, so the busiest customers cost the most of the budget.
What the batch gives up
The cap. Every item goes through the same check as a single record, so on
a hard-capped quota the items past the limit are refused, after the customer
already has the response. That is overshoot, bounded by one flush interval of
traffic. If a route has to stop dead at the cap, it belongs in mode one.
Which items were refused. Over HTTP, the batch answers 200 with per-item
success: false entries in results (one request is not one
unit walks through them). Through
@buildbase/sdk 0.0.73 you never see them: a batch with any refused item
throws Failed to record batch usage, with no status and no results. That
is why the buffer tells that error apart from the rest. The accepted items are
already recorded and the refused ones will be refused again, so re-queueing the
batch would resend the same 100 items at every flush while newer usage waits
behind them.
Anything in memory at a crash. The buffer lives in the process. A deploy
that kills it without draining loses up to a minute of usage, so wire drain()
to shutdown. It waits for a batch already in flight and keeps flushing past the
first 100 items, up to a deadline:
process.on('SIGTERM', async () => {
await usageBuffer.drain(10_000);
process.exit(0);
});A crash still loses what was buffered. Notice which way that fails: toward
billing too little, never twice. A resend after a failed flush reuses every
item's key, so an item the server did count the first time comes back as a
replay rather than a second charge. (The server checks and increments in
separate steps, so two flushes of the same item at the same instant could both
count. The flushing set keeps one batch per workspace in flight, so one process
never does that.)
On Hono, the real status arrives after await next()
Code after await next() runs on the way back out, with the response already
built. Hono's docs say next() never throws: an error goes to app.onError(), or
becomes a 500, before it comes back up, so c.res.status is the real status
either way.
The Hono setup guide keeps forSession private inside its gateway. Expose the
usage calls with one more entry on the object buildbaseFor() returns:
// src/buildbase.ts, inside buildbaseFor()'s return
usage: (sessionId: string) => forSession(sessionId).usage,// src/meter.ts
import { getSignedCookie } from 'hono/cookie';
import { createMiddleware } from 'hono/factory';
import { buildbaseFor } from './buildbase';
import { config } from './config'; // the setup guide's env() helper, moved to its own file
export const meterAfter = (quotaSlug: string) =>
// Typed, so c.var.workspace is checked here. Mount it after requireWorkspace.
createMiddleware<{ Variables: { workspace: { id: string } } }>(
async (c, next) => {
await next();
if (c.res.status < 200 || c.res.status >= 300) return;
const workspace = c.get('workspace');
if (!workspace) return; // mounted without requireWorkspace in front
const sessionId = await getSignedCookie(
c,
config(c).COOKIE_SECRET!,
'bb_session'
);
if (!sessionId) return;
const recorded = buildbaseFor(config(c))
.usage(sessionId)
.record(workspace.id, {
quotaSlug,
quantity: 1,
idempotencyKey: `${quotaSlug}:${crypto.randomUUID()}`,
})
.catch((err) => console.error('usage not recorded', err.message));
try {
c.executionCtx.waitUntil(recorded); // Workers: keep the isolate alive for it
} catch {
// Node and Bun have no ExecutionContext; the promise runs on its own.
}
}
);The try is not decoration. On Node and Bun, reading c.executionCtx throws
This context has no ExecutionContext, and with the response already gone, the
only sign would be an error in your logs on every metered request.
That is still one call per request, so the 60-a-minute arithmetic applies. On
Node and Bun, a Hono app can use the Express buffer module unchanged, with its sender set, because
it is plain module state and a timer. Workers should not: Cloudflare says not to rely on global state across
requests, since an isolate can be evicted at any time, and there is no timer
that outlives a request to flush it. Batch there through
something that persists, such as a Cloudflare Queue whose consumer calls
recordBatch.
On Fastify, onResponse is mode two
The onResponse hook runs when "a response has been sent, so you will
not be able to send more data to the client", which is exactly mode two.
Mode one goes in a preHandler hook, the same shape as the Express
meterFirst, answering with reply.code(402).
The Fastify setup guide's BuildBaseGateway has no usage method, so add one,
and add it to the fake your tests pass in:
// src/plugins/app/buildbase.ts
export interface BuildBaseGateway {
// ...the six methods from the setup guide
recordUsageBatch(
sessionId: string,
workspaceId: string,
items: Array<{
quotaSlug: string;
quantity: number;
idempotencyKey: string;
}>
): Promise<void>;
}The real implementation awaits
client.withSession(sessionId).usage.recordBatch(workspaceId, { items }) and
returns nothing.
The test fake pushes the items into an array, which gives you a test that a
404 is never billed.
// src/plugins/app/meter.ts
import { randomUUID } from 'node:crypto';
import fp from 'fastify-plugin';
import * as usageBuffer from '../../lib/usage-buffer';
export default fp(async (fastify) => {
usageBuffer.setSender((sessionId, workspaceId, items) =>
fastify.buildbase.recordUsageBatch(sessionId, workspaceId, items)
);
fastify.addHook('onResponse', async (request, reply) => {
const quotaSlug = request.routeOptions.config.meter;
if (!quotaSlug) return;
if (reply.statusCode < 200 || reply.statusCode >= 300) return;
const workspace = request.workspace;
if (!workspace) return; // not a workspace route
const mayRecord = workspace.permissions.includes('workspace:usage:record');
usageBuffer.add(
workspace.id,
mayRecord ? request.session.buildbaseSessionId : null,
{ quotaSlug, quantity: 1, idempotencyKey: `${quotaSlug}:${randomUUID()}` }
);
});
});A route opts in with config: { meter: 'api-calls' } (declare meter on
Fastify's FastifyContextConfig so the read type-checks). The buffer is the
same module as the Express one, with the gateway as its sender, so the tests'
fake gateway sees every batch. It is plain JavaScript: in a TypeScript or ESM
Fastify project, port it to lib/usage-buffer.ts, which is a line-for-line
change plus types.
Do not use request.id as the key. Fastify's default is "value of
'request-id' header if provided or monotonically increasing integers" (and v5
ignores that header unless requestIdHeader is set), so two processes, or one process after a restart, hand out the same IDs. A reused key
is a replay, and a replay is recorded as used: 0: real usage, silently not
billed.
What breaks, and how you can tell
| You see | Cause, and the fix |
|---|---|
Too many usage recording requests. Please slow down. in your logs, or your own 503s from mode one | The metered routes together make over 60 usage calls a minute. Move the busiest route to mode two, or lengthen the flush interval. The 429 that is not your quota explains the limiter. |
| Console usage lags traffic by a minute | Mode two working. The flush interval is the lag. |
| Usage drops after every deploy | The buffer died with the process. Check that SIGTERM reaches Node (a shell wrapper in the container's start command swallows it) and that drain() runs. |
| Usage stops moving for one workspace, and "usage partly refused" is in the log | Its quota is hard-capped and full, or the workspace has no active subscription, or the quota slug is not on its plan. The batch was refused in part; the accepted items were recorded and the rest dropped. Raise the cap or turn overage on. |
| A flush fails with 403 over and over | The kept member was signed out or demoted since. The next request from a member who may record replaces the session. A workspace that only ever sees viewers never lands its usage, so keep metered routes off roles that cannot be billed. |
| Customers billed for requests that failed | The route is on mode one, which records before the work. That is the trade. If failures are common there, move it to mode two. |
| Metered usage far below traffic, and nothing errors | Keys are repeating. Look for request.id, a counter or a hard-coded key. |
Install
npm i @buildbase/sdk