Connection pool exhausted as org count grows
MongoNotConnectedError under load is usually LRU eviction, not a dead database. The arithmetic behind the 30-connection tenant window and where your own ceiling lands.
In short
MongoNotConnectedError under load is usually eviction, not a dead database. A tenant server holds at most 30 open tenant connections regardless of org count, so sockets stay bounded near 67 per API process while evictions grow with the number of orgs active in any five-minute window.
Two lines show up in the tenant server log, usually a few minutes apart, and it is the second one that wakes someone up:
[MongoDB] Pool full (30), evicting LRU tenant <orgId>
MongoNotConnectedError: Client must be connected before running operations
The second error almost never names the org that caused it. It lands on some unrelated request that was mid-query on a connection the pool had just decided to close, which is why the first instinct is to go and check whether the database is up. It is up. You ran out of slots in a fixed-size window, and something else paid for it.
So the useful question is not "is the pool exhausted", it is "at what org count does it become exhausted, and how far am I from that". That is arithmetic you can run today on your own numbers.
Before you start
These figures are the defaults in packages/server-core/src/config.ts for a
BuildBase tenant server running per-org MongoDB databases. The shape of the
arithmetic holds for any database-per-tenant system with a bounded connection
cache. The constants will be yours.
The multiplication almost everyone runs first
Orgs times connections per org. You have 500 orgs, MONGO_ORG_POOL_SIZE is 3,
so the server wants 1,500 sockets and your database tier will not give you
1,500 sockets. Time to shard.
That calculation is wrong, and it is wrong in the direction that makes people solve the wrong problem. Sockets are not what grows.
The multiplication that actually runs
The cap is on tenant connections, not on orgs. From config.ts, on API
defaults:
| Variable | Default | What it bounds |
|---|---|---|
MONGO_CENTRAL_POOL_SIZE | 5 | Sockets to the central database |
MONGO_ORG_POOL_SIZE | 3 | Base sockets per tenant connection |
MONGO_ORG_POOL_SIZE_MIN | 2 | Floor the base scales down to |
MONGO_MAX_TENANT_CONNECTIONS | 30 | Tenant connections held at once |
MONGO_TENANT_IDLE_MS | 300000 | Idle time before one is reaped |
Workers get smaller numbers: central 2, base 2, floor 1. Same shape.
The base does not stay the base. Every tenant connection is created with a pool size computed from how many tenant connections are already live, and the scaler is logarithmic:
// Mirrors getTenantPoolSize() in
// packages/server-core/src/database/mongo.ts. Pure arithmetic - paste it into
// a scratch file with your own constants rather than trusting mine.
const BASE = 3; // MONGO_ORG_POOL_SIZE
const MIN = 2; // MONGO_ORG_POOL_SIZE_MIN
function tenantPoolSize(liveTenantConnections: number): number {
if (liveTenantConnections <= 1) return BASE;
return Math.max(MIN, Math.floor(BASE / Math.log2(liveTenantConnections + 1)));
}Run it and the curve collapses almost immediately. At two live connections,
floor(3 / log2(3)) is 1, which clamps to the floor of 2. Every value above
that clamps too. So connections one and two get 3 sockets each and connections
three through thirty get 2.
That gives one API process a hard ceiling of:
5 central
+ 6 first two tenant connections at 3
+ 56 remaining 28 tenant connections at 2
= 67 sockets
A worker process comes out at 34 by the same route. One API replica plus one worker is 101 sockets against your database, whether the deployment serves 30 orgs or 900.
Key takeaway
Adding orgs does not add sockets. MONGO_MAX_TENANT_CONNECTIONS caps the
tenant connection cache at 30 per process, so socket count is bounded at 67 on
API defaults no matter how many orgs exist. What grows with org count is the
eviction rate inside that fixed window.
The number that does grow
Call it A: the count of distinct orgs one process serves inside the idle window, which is five minutes by default. A is the number to measure, and it is not your total org count. It is your total org count divided across replicas, multiplied by how likely each org is to be active in any given five minutes.
Three regimes, and you can tell which one you are in from the logs alone:
- A at or under 30. Every org in the window is resident. No
Pool fulllines at all. The eviction sweep runs every 60 seconds and reports idle reaps, nothing more. - A just over 30. Occasional
Pool fulllines. The orgs that get evicted are the quiet ones, and when they come back they pay a cold connect. Users notice this as a slow first page after switching orgs. - A well over 30. Thrashing. The connection you evict to make room is one
that another in-flight request is still holding, and that request dies with
MongoNotConnectedErrorrather than anything that names an org or a pool.
We hit the third regime on a laptop, not in production, and the numbers are in
notes/VERIFICATION_RUNBOOK.md: 45,825 Pool full evictions in a single run
at 163 orgs, with MongoNotConnectedError surfacing on requests that had
nothing to do with the cause, including a failed database seed.
163 orgs against a 30-connection window should not produce forty-five thousand evictions. What produced them was a link redirector that did not know which org owned a link, so it asked all of them: one tenant query per org per click, and a full sweep on every miss. O(orgs) per redirect against a bounded pool. Any single code path that touches every org turns a comfortable A into a hostile one, and that path is usually something small that nobody thinks of as a tenant-heavy route.
Where the ceiling lands
Roughly 1,000 orgs per deployment is the documented connection-pool ceiling, and it is the number we plan against.
But it is not the number that will bind you first. A deployment with 300 orgs, one API replica and a cron that walks every tenant is in trouble; a deployment with 900 orgs across enough replicas that each one holds fewer than 30 active at a time is not. Total orgs is the headline. A is the operational limit.
The first symptom is latency, not an error
Long before anything throws, org switching gets slow. Here is what a cold
tenant connection costs, in order, from mongo.ts:
- An org-existence check, served from Redis on a 5 minute TTL, or a central database lookup on a miss.
mongoose.createConnectionagainst the tenant URL.- A wait for the connection to open, capped at 10 seconds, which fails as
Connection timeout for <orgId>. - An admin
pingto validate it.
All of that sits in front of the first query of the first request after an eviction. One user, switching orgs, sees a page that used to be fast take a second or two. Nobody files a bug for that.
Then it gets worse in a way that does not look like a pool problem. The connection factory retries five times with a 5 second gap between attempts, and each of the six attempts waits out a 5 second server-selection timeout first, so a request against a tenant database that cannot be reached sits there for close to a minute before anything is raised. Meanwhile every other request on that process is fine. You get a p99 that climbs while p50 sits flat, which reads as a network blip, an overloaded replica, a bad deploy - anything except a fixed-size cache doing exactly what it was configured to do.
Grep for these before you reach for a profiler:
grep 'Pool full' server.log | wc -l # evictions under pressure
grep 'Evicted .* tenant connections' server.log # the 60s idle sweep
grep 'Creating tenant connection' server.log # cold connectsThe third line is the one worth watching over time. A healthy process creates tenant connections at startup and then goes quiet. A thrashing one creates them forever.
The number you set may not be the number you are running
This one costs an afternoon if you do not know it.
On the cloud deploy path, the heredoc in .github/workflows/docker-server.yml
writes the .env that compose forwards into the container, and that heredoc
is the allowlist. A variable with a line there is live. A variable the code
reads but the heredoc omits is not using its configured value - it is unset,
and the code falls back to its default.
MONGO_ORG_POOL_SIZE has no line in that heredoc. Neither do
MONGO_ORG_POOL_SIZE_MIN, MONGO_MAX_TENANT_CONNECTIONS or
MONGO_TENANT_IDLE_MS. That is deliberate, and notes/ENVIRONMENT.md says so:
their defaults are right for cloud. But it means setting one in the environment
and redeploying changes nothing at all, silently, and you will spend the next
hour wondering why the eviction rate did not move.
Do not trust the environment. The server prints what it is running, at startup:
[MongoDB] Central DB: Standard MongoDB (pool: 5)
[MongoDB] Tenant DB pool: 3 per org
And every cold connect prints the scaled value in force at that moment:
[MongoDB] Creating tenant connection for org <orgId> (pool: 12 active, poolSize: 2)
poolSize: 2 there is the floor doing its job. If you expected 3 and see 2,
nothing is broken.
What actually moves the number
Lowering MONGO_ORG_POOL_SIZE barely helps. It only applies to the first
two tenant connections, so dropping it from 3 to 2 saves 2 sockets out of 67.
The variable that governs 28 of the 30 is MONGO_ORG_POOL_SIZE_MIN. Take that
to 1 and the API ceiling falls to 39. That is a real cut, and it costs you
concurrency inside each tenant: one socket per org means two simultaneous
queries for the same org queue behind each other.
Shorter idle reaping frees slots, not sockets. MONGO_TENANT_IDLE_MS at
300000 with a sweep every 60 seconds means a quiet org holds a slot for up to
five minutes after its last query. Cutting it to 60000 returns slots to the
window faster, which lowers Pool full evictions of connections that are still
in use. It does not lower peak socket count, and it buys more cold connects. On
a workload with a long tail of rarely-active orgs this is the cheapest win
available. On one where every org is active every minute it does nothing.
Raising MONGO_MAX_TENANT_CONNECTIONS trades sockets for evictions,
linearly. At 100 the ceiling becomes 207 sockets per API process. If your
database tier has the headroom and A sits somewhere between 30 and 100, this is
the correct fix and it is one line. Check what you have before you spend it:
mongosh "$MONGO_CONNECTION_URL" --eval \
'JSON.stringify(db.serverStatus().connections)'available against your replica count is the budget. Spend it deliberately.
Past the ceiling, config stops being the answer. Every knob above moves a constant. None of them changes the fact that one process can only cache so many tenant connections, and that a route touching every org will exhaust any window you give it. Two things actually change the slope:
Remove the O(orgs) paths. The redirector above got fixed by giving links an
owner index in Redis, link-owner:<linkId> -> orgId, written at creation and
backfilled on the rare sweep that still resolves one. Forty-five thousand
evictions became a single key lookup. No pool setting was touched. Go find your
equivalent: anything that iterates orgs to answer a request is a pool problem
wearing a different hat.
Then partition. Route each org consistently to the same replica so that a
process's working set is a stable subset rather than a random sample of
everything. The code already assumes this is where it goes - the comment above
the scaler in mongo.ts reads "with org-sticky routing, each instance handles
a subset of orgs, so fewer connections per org are needed". Sticky routing is
what makes A a number you control instead of a number you observe.
The trade you actually bought
A database per org is not free and it was never sold as free. What it buys is isolation you cannot get wrong: a query that forgets its tenant filter has nothing else to find, because there is nothing else in the database. What it costs is a connection per active tenant, and connections are the one resource that does not scale by adding application memory.
That is the trade. It is a good trade for a B2B product with hundreds of orgs and real isolation requirements, and a bad one for a consumer app with a million accounts. We picked it deliberately and we live with a connection ceiling because of it. The full cost breakdown is in what database-per-tenant actually costs, and the decision itself, including when a shared schema is the better answer, is in database-per-tenant vs shared schema.
If you are earlier than any of this and just wiring up tenancy, start with multi-tenant workspaces and come back to this page when the eviction lines appear.
Install
npm i @buildbase/sdk