Guides
Build guide

Usage metering middleware in Express: record before the handler or after the response (Hono and Fastify too)

Usage metering middleware for Express, Hono and Fastify: when to record before the handler, when to batch after the response, and the rate limit that decides it.

Dharmendra Jagodana10 min read

In short

Metering middleware in Express can record before the handler or after the response. Recording first enforces the quota but counts requests that fail, one call each against a 60-a-minute limit. Recording after a 2xx and flushing with recordBatch bills only successes at volume, but cannot stop a request at the cap.

Write the obvious metering middleware for an Express API, one bb.usage.record() per request, and you have set the throughput of every metered route you own to 60 requests a minute.

That is not your quota. It is the rate limit on BuildBase's usage endpoints, counted per IP, and every usage call comes from your server's one address. All your users share it. It breaks first.

So the decision is not where the middleware goes. It is which of two jobs the call does on this route:

  • Mode one: record before the handler. The record call is also the quota check, so it can refuse the request. At one call per request, it suits expensive routes that stay well under 60 a minute, like report generation.
  • Mode two: record after the response, in batches. Only successful responses are billed, and 100 of them go in one call. Nothing is refused per request, because by the time the batch runs, the responses are gone.

You cannot have both on one call. Most APIs want both, on different routes.

Before you start

An Express, Hono or Fastify app with sign-in working through its setup guide (Express, Hono, Fastify), and the workspace resolved per request as in multi-tenant workspaces in Express, which puts it on req.workspace. In the console, a plan with the quotas you meter, here reports and api-calls.

What the middleware knows, and when

Before the handler, middleware cannot know whether the work will succeed. After the response, it knows the status code and can no longer do anything about the request. Each mode gives up one half.

Our per-API-call guide says metering in middleware is "wrong for anything expensive" because the middleware "does not know whether the work succeeded". For mode one that is exactly right. A request that fails after the record has still been counted, and bb.usage has no way to take it back. quantity has a minimum of 1, so there is no negative record either.

Mode one: the record is the gate

// middleware/meter.js
const crypto = require('node:crypto');
const buildbase = require('../config/buildbase');

/** Refuses the request when the quota is spent. One usage call per request. */
exports.meterFirst =
  ({ quota, quantity = 1 }) =>
  async (req, res, next) => {
    const key = req.get('Idempotency-Key') ?? crypto.randomUUID();
    if (key.length > 128) {
      return res.status(400).json({ message: 'Idempotency-Key is too long.' });
    }
    try {
      await buildbase.forRequest(req).usage.record(req.workspace.id, {
        quotaSlug: quota,
        quantity,
        idempotencyKey: `${quota}:${key}`,
        source: 'api',
      });
      return next();
    } catch (err) {
      if (err.status === 400 && err.message.startsWith('Quota limit reached')) {
        return res
          .status(402)
          .json({ message: `This workspace has used its ${quota} allowance.` });
      }
      if (err.status === 429) {
        // BuildBase's rate limit, not the customer's quota.
        res.set('Retry-After', '60');
        return res
          .status(503)
          .json({ message: 'Usage metering is busy. Try again shortly.' });
      }
      return next(err);
    }
  };
// routes/api.js
router.post('/reports', meterFirst({ quota: 'reports' }), reports.create);

On a plan that is trialing, or a quota with overage switched off, record is an atomic check-and-increment, so two requests racing for the last unit cannot both get it. Server-side quota enforcement walks through that, and why a separate "is there any left?" read is the wrong check. With overage on, record never refuses. It bills the overage and the request goes through.

Three choices in there are ours:

402 for a spent quota, not 429. A 429 tells a well-behaved client to back off and retry. Retrying a spent allowance fails until the period resets, so we answer with the status that says "this needs paying for".

503 when BuildBase rate-limits us. It fails closed. Doing the work unmetered would give away exactly the expensive requests this mode exists for. And the status is ours to pick, because the upstream 429 is about our server's call volume, not the customer.

The client's Idempotency-Key, when it sends one. A retried request with the same key is counted once: the server matches the key on workspace, quota and key, and a replay returns used: 0. Without a header we mint a UUID, which makes the record safe to retry inside this request but treats a client retry as a new request. A server-rendered form can hand out its key when the page renders instead, as usage-based billing in React Router does from the loader. The cap of 128 characters keeps the prefixed key under the 255 the API accepts.

To check it, set the reports quota to 2 with overage off and send three requests. The third answers 402.

Mode two: record after the response, in batches

The middleware only notes what happened. A buffer sends it later.

// middleware/meter.js (continued)
const usageBuffer = require('../lib/usage-buffer');

/** Bills 2xx responses only. Never refuses a request. */
exports.meterAfter =
  ({ quota, quantity = 1 }) =>
  (req, res, next) => {
    const clientKey = req.get('Idempotency-Key');
    // A long client key is not truncated: two truncated keys could collide.
    const key =
      clientKey && clientKey.length <= 128 ? clientKey : crypto.randomUUID();
    const workspaceId = req.workspace.id;
    // Only a session that may record usage is worth keeping for the flush.
    const sessionId = req.workspace.permissions.includes(
      'workspace:usage:record'
    )
      ? req.session.buildbaseSessionId
      : null;
    res.on('finish', () => {
      if (res.statusCode < 200 || res.statusCode >= 300) return;
      usageBuffer.add(workspaceId, sessionId, {
        quotaSlug: quota,
        quantity,
        idempotencyKey: `${quota}:${key}`,
        source: 'api',
      });
    });
    next();
  };

Why only 2xx? A 4xx is the client's mistake and a 5xx is ours. Neither belongs on an invoice. Node's docs are precise about what finish means: the response "has been sent", which "does not imply that the client has received anything yet". For billing, sent is the right line.

// lib/usage-buffer.js
// The one BuildBase call, set once at startup: the Express app passes the
// SDK, a Fastify app its gateway. Nothing else in this file knows which.
/** @typedef {(sessionId: string, workspaceId: string, items: Array<{ quotaSlug: string, quantity: number, idempotencyKey: string, source?: string }>) => Promise<unknown>} Sender */
/** @type {Sender | null} */
let send = null;
/** @param {Sender} fn */
exports.setSender = (fn) => {
  send = fn;
};

const FLUSH_MS = 60_000;
const BATCH_LIMIT = 100; // the most recordBatch takes in one call

/** workspaceId -> items waiting to be sent */
const pending = new Map();
/** workspaceId -> the latest session from a member who may record */
const sessions = new Map();
/** Workspaces with a batch in flight, so two flushes never overlap. */
const flushing = new Set();

exports.add = (workspaceId, sessionId, item) => {
  if (sessionId) sessions.set(workspaceId, sessionId);
  const items = pending.get(workspaceId) ?? [];
  items.push(item);
  pending.set(workspaceId, items);
  if (items.length >= BATCH_LIMIT) flush(workspaceId);
};

async function flush(workspaceId) {
  const sessionId = sessions.get(workspaceId);
  const queued = pending.get(workspaceId);
  if (!send || !sessionId || !queued?.length) return; // nothing to do yet
  if (flushing.has(workspaceId)) return; // the next tick picks the rest up
  flushing.add(workspaceId);
  const items = queued.slice(0, BATCH_LIMIT);
  pending.set(workspaceId, queued.slice(BATCH_LIMIT));

  try {
    await send(sessionId, workspaceId, items);
  } catch (err) {
    if (
      err.status === undefined &&
      err.message === 'Failed to record batch usage'
    ) {
      // The server refused some items (usually a hard cap) and recorded the
      // rest.
      // A resend would be refused again, so drop the batch and say so.
      console.warn('usage partly refused', workspaceId, items.length);
    } else {
      // A 429, a 5xx or a network error: it may or may not have landed.
      // Put the items back; their keys make the resend safe either way.
      pending.set(workspaceId, [...items, ...(pending.get(workspaceId) ?? [])]);
      console.error('usage flush failed', workspaceId, err.message);
    }
  } finally {
    flushing.delete(workspaceId);
  }
}

exports.flushAll = () => Promise.all([...pending.keys()].map(flush));

/** For shutdown: keep flushing until nothing is left, or the deadline passes. */
exports.drain = async (deadlineMs = 10_000) => {
  const stop = Date.now() + deadlineMs;
  while (Date.now() < stop) {
    const left = [...pending.entries()].some(
      ([id, items]) => items.length && sessions.has(id)
    );
    if (!left && flushing.size === 0) return;
    await exports.flushAll();
    await new Promise((r) => setTimeout(r, 100));
  }
};

setInterval(exports.flushAll, FLUSH_MS).unref();

In the Express app, the sender is the SDK. forSession is one line added to config/buildbase.js from the setup guide, next to forRequest:

// config/buildbase.js (next to forRequest)
exports.forSession = (sessionId) => buildbase().withSession(sessionId);

// app.js, once at startup
const buildbase = require('./config/buildbase');
const usageBuffer = require('./lib/usage-buffer');

usageBuffer.setSender((sessionId, workspaceId, items) =>
  buildbase.forSession(sessionId).usage.recordBatch(workspaceId, { items })
);
// routes/api.js
router.get('/search', meterAfter({ quota: 'api-calls' }), search.run);

Send a few searches and nothing reaches BuildBase for up to a minute. Then one recordBatch lands with all of them, and api-calls in the console moves by the number of 2xx responses, not the number of requests.

Which session sends the batch? The requests are over, so there is no current user, and the server SDK takes no API key: every call it makes runs as some user's session. The buffer keeps, per workspace, the most recent session from a member whose permissions include workspace:usage:record, which the workspace middleware already resolved. A viewer's session would only earn a 403, so it is never kept.

This is the weakest part of the design, and we would rather say so. A kept session can be signed out, or its member demoted, before the flush. That shows up as a rejected batch while the items wait for the next request from a member who may record. Usage from a workspace whose only active users are viewers waits until one comes along. An API token can be exchanged for a session, but that session's user still has to be a member of the workspace who may record, so it does not remove the problem for a multi-tenant API.

The arithmetic of the flush interval

A batch carries one workspace and counts once against the 60-a-minute limit. Flush every N seconds, with W workspaces active, and you make W × 60 / N calls a minute. Keep that at or under 60 and the rule falls out:

The flush interval, in seconds, has to be at least the number of workspaces active during it.

At the 60-second interval above, that holds for up to 60 busy workspaces at a time. And the budget is shared: every mode-one record spends from the same 60, and so does every other process behind the same outgoing IP. Two API servers behind one NAT halve it. Past that, lengthen the interval and accept that usage shows up later.

A workspace that fills 100 items early flushes early and spends an extra call, so the busiest customers cost the most of the budget.

What the batch gives up

The cap. Every item goes through the same check as a single record, so on a hard-capped quota the items past the limit are refused, after the customer already has the response. That is overshoot, bounded by one flush interval of traffic. If a route has to stop dead at the cap, it belongs in mode one.

Which items were refused. Over HTTP, the batch answers 200 with per-item success: false entries in results (one request is not one unit walks through them). Through @buildbase/sdk 0.0.73 you never see them: a batch with any refused item throws Failed to record batch usage, with no status and no results. That is why the buffer tells that error apart from the rest. The accepted items are already recorded and the refused ones will be refused again, so re-queueing the batch would resend the same 100 items at every flush while newer usage waits behind them.

Anything in memory at a crash. The buffer lives in the process. A deploy that kills it without draining loses up to a minute of usage, so wire drain() to shutdown. It waits for a batch already in flight and keeps flushing past the first 100 items, up to a deadline:

process.on('SIGTERM', async () => {
  await usageBuffer.drain(10_000);
  process.exit(0);
});

A crash still loses what was buffered. Notice which way that fails: toward billing too little, never twice. A resend after a failed flush reuses every item's key, so an item the server did count the first time comes back as a replay rather than a second charge. (The server checks and increments in separate steps, so two flushes of the same item at the same instant could both count. The flushing set keeps one batch per workspace in flight, so one process never does that.)

On Hono, the real status arrives after await next()

Code after await next() runs on the way back out, with the response already built. Hono's docs say next() never throws: an error goes to app.onError(), or becomes a 500, before it comes back up, so c.res.status is the real status either way.

The Hono setup guide keeps forSession private inside its gateway. Expose the usage calls with one more entry on the object buildbaseFor() returns:

// src/buildbase.ts, inside buildbaseFor()'s return
usage: (sessionId: string) => forSession(sessionId).usage,
// src/meter.ts
import { getSignedCookie } from 'hono/cookie';
import { createMiddleware } from 'hono/factory';

import { buildbaseFor } from './buildbase';
import { config } from './config'; // the setup guide's env() helper, moved to its own file

export const meterAfter = (quotaSlug: string) =>
  // Typed, so c.var.workspace is checked here. Mount it after requireWorkspace.
  createMiddleware<{ Variables: { workspace: { id: string } } }>(
    async (c, next) => {
      await next();
      if (c.res.status < 200 || c.res.status >= 300) return;
      const workspace = c.get('workspace');
      if (!workspace) return; // mounted without requireWorkspace in front

      const sessionId = await getSignedCookie(
        c,
        config(c).COOKIE_SECRET!,
        'bb_session'
      );
      if (!sessionId) return;

      const recorded = buildbaseFor(config(c))
        .usage(sessionId)
        .record(workspace.id, {
          quotaSlug,
          quantity: 1,
          idempotencyKey: `${quotaSlug}:${crypto.randomUUID()}`,
        })
        .catch((err) => console.error('usage not recorded', err.message));

      try {
        c.executionCtx.waitUntil(recorded); // Workers: keep the isolate alive for it
      } catch {
        // Node and Bun have no ExecutionContext; the promise runs on its own.
      }
    }
  );

The try is not decoration. On Node and Bun, reading c.executionCtx throws This context has no ExecutionContext, and with the response already gone, the only sign would be an error in your logs on every metered request.

That is still one call per request, so the 60-a-minute arithmetic applies. On Node and Bun, a Hono app can use the Express buffer module unchanged, with its sender set, because it is plain module state and a timer. Workers should not: Cloudflare says not to rely on global state across requests, since an isolate can be evicted at any time, and there is no timer that outlives a request to flush it. Batch there through something that persists, such as a Cloudflare Queue whose consumer calls recordBatch.

On Fastify, onResponse is mode two

The onResponse hook runs when "a response has been sent, so you will not be able to send more data to the client", which is exactly mode two. Mode one goes in a preHandler hook, the same shape as the Express meterFirst, answering with reply.code(402).

The Fastify setup guide's BuildBaseGateway has no usage method, so add one, and add it to the fake your tests pass in:

// src/plugins/app/buildbase.ts
export interface BuildBaseGateway {
  // ...the six methods from the setup guide
  recordUsageBatch(
    sessionId: string,
    workspaceId: string,
    items: Array<{
      quotaSlug: string;
      quantity: number;
      idempotencyKey: string;
    }>
  ): Promise<void>;
}

The real implementation awaits client.withSession(sessionId).usage.recordBatch(workspaceId, { items }) and returns nothing. The test fake pushes the items into an array, which gives you a test that a 404 is never billed.

// src/plugins/app/meter.ts
import { randomUUID } from 'node:crypto';
import fp from 'fastify-plugin';

import * as usageBuffer from '../../lib/usage-buffer';

export default fp(async (fastify) => {
  usageBuffer.setSender((sessionId, workspaceId, items) =>
    fastify.buildbase.recordUsageBatch(sessionId, workspaceId, items)
  );

  fastify.addHook('onResponse', async (request, reply) => {
    const quotaSlug = request.routeOptions.config.meter;
    if (!quotaSlug) return;
    if (reply.statusCode < 200 || reply.statusCode >= 300) return;
    const workspace = request.workspace;
    if (!workspace) return; // not a workspace route
    const mayRecord = workspace.permissions.includes('workspace:usage:record');
    usageBuffer.add(
      workspace.id,
      mayRecord ? request.session.buildbaseSessionId : null,
      { quotaSlug, quantity: 1, idempotencyKey: `${quotaSlug}:${randomUUID()}` }
    );
  });
});

A route opts in with config: { meter: 'api-calls' } (declare meter on Fastify's FastifyContextConfig so the read type-checks). The buffer is the same module as the Express one, with the gateway as its sender, so the tests' fake gateway sees every batch. It is plain JavaScript: in a TypeScript or ESM Fastify project, port it to lib/usage-buffer.ts, which is a line-for-line change plus types.

Do not use request.id as the key. Fastify's default is "value of 'request-id' header if provided or monotonically increasing integers" (and v5 ignores that header unless requestIdHeader is set), so two processes, or one process after a restart, hand out the same IDs. A reused key is a replay, and a replay is recorded as used: 0: real usage, silently not billed.

What breaks, and how you can tell

You seeCause, and the fix
Too many usage recording requests. Please slow down. in your logs, or your own 503s from mode oneThe metered routes together make over 60 usage calls a minute. Move the busiest route to mode two, or lengthen the flush interval. The 429 that is not your quota explains the limiter.
Console usage lags traffic by a minuteMode two working. The flush interval is the lag.
Usage drops after every deployThe buffer died with the process. Check that SIGTERM reaches Node (a shell wrapper in the container's start command swallows it) and that drain() runs.
Usage stops moving for one workspace, and "usage partly refused" is in the logIts quota is hard-capped and full, or the workspace has no active subscription, or the quota slug is not on its plan. The batch was refused in part; the accepted items were recorded and the rest dropped. Raise the cap or turn overage on.
A flush fails with 403 over and overThe kept member was signed out or demoted since. The next request from a member who may record replaces the session. A workspace that only ever sees viewers never lands its usage, so keep metered routes off roles that cannot be billed.
Customers billed for requests that failedThe route is on mode one, which records before the work. That is the trade. If failures are common there, move it to mode two.
Metered usage far below traffic, and nothing errorsKeys are repeating. Look for request.id, a counter or a hard-coded key.

Install

npm i @buildbase/sdk
express
hono
fastify
billing

Frequently Asked Questions

Should usage be recorded before or after the response?

Before, when the work is expensive and the cap has to hold: the record call is what refuses the request. After, when the route is cheap and busy, and billing only successful responses matters more than stopping exactly at the cap.

Why not record usage once per request in middleware?

The usage endpoints allow 60 requests a minute per IP, and every call comes from your API server. One record per request makes that limit the throughput ceiling of every metered route combined.

When should I use recordBatch instead of record?

When a route is busy enough that one call per request would pass 60 a minute. A batch takes up to 100 items for one workspace and counts once against the limit. The cost is that nothing is refused per request, because the batch runs after the responses have gone.

What happens if the metering call fails after the response is sent?

The customer has the result and the usage is not counted yet. Keep the items and send them again with the same idempotency keys, so an item that did land the first time is not counted twice.

Ship it

Create a project and run this guide against your own workspace.

7-day free trialNo credit card requiredCancel anytime