Token expired mid-request on a long operation
A 401 halfway through an upload, while the user is still signed in for another 29 days. Two clocks, not one, and the fix is a refresh before the call rather than a retry after it.
In short
Two different clocks fail this request. SESSION_DEFAULTS in packages/shared holds tokenExpiry of 2-minute and sessionTTL of 30 days. The short one belongs to a handoff token crossing a server boundary, not to your API calls, so a user who is definitively still logged in can still see a 401 mid-flight.
The request that failed was running for four minutes. The user who started it is signed in for another 29 days. Both are true at once. That contradiction is the entire bug report.
What you pasted into Google was probably one of these:
Request failed (401: Unauthorized - Please check your session)
Token verification failed: jwt expired
BuildBase: Not authenticated (getSessionId returned null)None of them means the user logged out. They mean something in the path ran out of time while the work was still going, and the thing that ran out was not the session.
Two clocks, and the one you are watching is the long one
SESSION_DEFAULTS in packages/shared/src/constants/platform-stats.ts holds
both numbers, three lines apart:
export const SESSION_DEFAULTS = {
sessionTTL: '30 days',
sessionTTLShort: '30-day',
permissionRefresh: 'hourly',
tokenExpiry: '2-minute',
webhookMaxAge: '5 minutes',
batchLimit: '100',
} as const;sessionTTL: '30 days' is the session. It is a server-side record, and
SESSION_TTL in server/src/routes/v1/auth/constants.ts is the same value in
seconds. On a device the user asked you to remember it stretches further, to
whatever rememberMeDays the org set, defaulting to 90, under a hard ceiling of 180. That clock is generous on purpose.
tokenExpiry: '2-minute' is a different thing entirely, and the gap between the
two is where your request died.
Key takeaway
A 401 mid-request usually means a short-lived credential somewhere in the path expired, not that the session did. The session record can have 29 days left while a two-minute handoff token, minted when the flow started, ran out before the step that redeems it.
The two-minute clock is not sitting on your API calls
Here is the part that catches people. The two-minute expiry does not belong to every authenticated request. It belongs to tokens that cross a boundary between two servers and are meant to be redeemed the instant they arrive.
Look at signOrgSwitchToken in central-server/src/utils/es256-jwt.ts. It
signs with ES256, stamps a jti, locks aud to one specific tenant server URL,
and sets expiresIn: '2m'. The comment above it says what the number is for:
just enough for the client-to-tenant handshake. The token goes central, then
client, then tenant, and never back. The social login path does the same thing a
different way - LOGIN_HANDOFF in the tenant's Redis TTL table is 120 seconds,
described as the "token pickup window after a redirect," and the code behind it
is consumed with an atomic get-and-delete so it can only be read once.
So the short clock is deliberate, and raising it is not the fix you want to file a ticket for. Two minutes is not a guess about how fast your API is. It is the window in which a bearer credential is allowed to sit in a browser before it stops being useful to anyone who steals it. Stretch it to an hour to survive your long upload and you have not fixed the upload; you have widened the window on the one credential in the system that travels through a URL bar.
The honest gap in our own constants: tokenExpiry: '2-minute' does not say
which token it describes, and I had to read two server files to find out.
What actually straddles the boundary
Four shapes of work outlive a short credential, and they are the four that generate this error:
- A file upload where the browser is still streaming bytes when the call that authorized it has already aged out.
- A long report that queries, aggregates and renders before it answers.
- A batch import that walks thousands of rows under one request.
- An agent call that thinks for a while, where the model, not your code, decides how long the request takes.
Before you blame the token, check which clock you actually hit, because the
SDK has one of its own. timeout on the client config defaults to 30000ms, and
maxRetries defaults to 0 with the note that it "only retries on network errors
and 5xx responses, never on 4xx."
That gives you a clean split. If the work died at roughly 30 seconds with an abort, you hit the SDK timeout and no token was involved. If it died at some other point with a 401, the credential expired. And whichever it was, nothing retried it for you, because the SDK refuses to retry a 4xx and it is right to.
A third cause looks identical from outside. It is not a clock at all.
SessionService.getSession re-reads the user from the database roughly
hourly, which is the permissionRefresh: 'hourly' line in the same constant. If
that read comes back with the user blocked or deleted, the session is removed
everywhere and the next call fails. Your onSessionExpired callback receives
'expired' for that too: the type allows 'invalid', but 0.0.70 never sends
it. Refreshing will not help, and a retry loop that keeps trying is just noise
in your logs.
Where the refresh is supposed to fit
Before the long call. Not in a catch block after it.
That sounds obvious until you look at how the credential is resolved. I read the
shipped bundle of @buildbase/sdk 0.0.68 to check this rather than guessing
from the types: the server client's action modules resolve the session by
awaiting your getSessionId callback on every call. So a fresh cookie is
picked up by the next call without you doing anything. withSession(id) is the
opposite by design - it resolves once and pins that id for the life of the
scoped client, which is exactly what you want for a background job and exactly
what you do not want for a user-facing client running for hours.
Two things worth knowing before you lean on this:
bb.auth() does not phone the server. It calls your getSessionId, returns
{ sessionId } if it got a string, and returns null if it did not. That
proves the cookie is present. It proves nothing about whether the session behind
it is alive. To learn that, make a cheap authenticated call and let it fail
fast, while the request is young and a failure costs you nothing.
And BuildBaseSession is { sessionId: string }. There is no expiry on it, and
the SDK exposes no explicit refresh method - I grepped for one. getSessionId
is the refresh seam, and it is the only one.
The shape of the fix
Resolve the credential first, then hand the long work a session that belongs to the job rather than to somebody's browser tab. The canonical server setup is the one from our own snippets:
import BuildBase from '@buildbase/sdk';
const bb = BuildBase({
serverUrl: 'https://api.yourapp.com',
orgId: 'your-org-id',
getSessionId: async () => {
const cookies = await getCookies();
return cookies.get('bb-session')?.value ?? null;
},
});Then the import route. Every chunk carries a key derived from the batch and the row index, so a replay lands on the same keys and cannot count anything twice:
import BuildBase from '@buildbase/sdk';
export async function POST(req: Request) {
const batchId = req.headers.get('x-request-id') ?? crypto.randomUUID();
const { workspaceId, rows } = await parseUpload(req);
// 1. Resolve the credential while the request is young. A cheap call here
// fails in milliseconds; the same failure after the import costs minutes.
const session = await bb.auth();
if (!session) {
return Response.json({ error: 'Unauthorized' }, { status: 401 });
}
await bb.workspace.get(workspaceId);
// 2. Run the long part under a session the job owns, not the user's cookie.
const api = bb.withSession(process.env.BB_SERVICE_SESSION_ID!);
// 3. One request per 100 rows, each item keyed so a replay is a no-op.
for (let i = 0; i < rows.length; i += 100) {
await api.usage.recordBatch(workspaceId, {
items: rows.slice(i, i + 100).map((_row, n) => ({
quotaSlug: 'rows-imported',
quantity: 1,
idempotencyKey: `${batchId}:${i + n}`,
})),
});
}
return Response.json({ imported: rows.length });
}I have not run this against a live stack; parseUpload and the service session
id are yours to supply. The recordBatch ceiling of 100 items per request is
real and matches batchLimit: '100' in the same constant, and the batching
argument is worked through properly in
enforcing quotas server-side.
Splitting the work into chunks of 100 buys you something beyond the ceiling: a credential that expires at chunk 14 loses one chunk, not the import. Replay from 14 with the same keys and the first thirteen stay exactly as they were.
Warning
A retry without an idempotency key is how a refresh-and-replay turns one import into two. The key is matched on workspace, quota slug and key together, so a replay returns the current status instead of incrementing again. Use something stable across attempts of the same logical operation, never a fresh uuid per attempt.
Reading the 401 you actually got
Four failures, four different reads. Work through them in this order:
The message contains jwt expired. That is a handoff or org-switch token
being verified after its two minutes. Something is holding a token that was
meant to be redeemed immediately - a retry queue, a saved link, a tab that sat
open through a redirect. Mint a new one rather than extending the old one.
The body is bare text, not JSON. The tenant's unAuthorized() helper sends
a plain 401 status with no JSON body, so the SDK's parse fails and it falls back
to its own status map. That is where Unauthorized - Please check your session
comes from. It is the SDK's generic string for a 401, not a description of what
happened, so it tells you nothing about which clock fired.
getSessionId returned null. No credential was found at all. The cookie is
missing or your callback threw and swallowed it. This is not an expiry.
onSessionExpired('expired') with a session that should be alive. Check
whether the user was blocked or removed. The hourly re-read deletes the session
the moment it sees either, and that is a permissions event wearing a 401.
Building this plumbing yourself is a week you will spend twice, which is part of the arithmetic in the cost of building auth yourself. The session and workspace side of it is in add multi-tenant workspaces in an afternoon.
Install
npm i @buildbase/sdk