Decide which axis is growing before you design for it
Scale is not one thing. A product can face thousands of small tenants, a handful of very large ones, a bursty write rate, or a document pipeline that dominates everything else. Each has a different answer, and designing for all of them at once produces a system that is expensive in every direction.
Write down the expected shape first: how many organisations, how concentrated, how much data per organisation, and which operations are genuinely the hot path. Revisit it as usage arrives, but be explicit about what you are assuming now.
The common error is scaling too early in the wrong place. Optimising the request path while one tenant holds most of your data, or caching hard before anything has been measured, adds complexity and buys nothing.
expected_shape:
organisations: thousands
concentration: low
hot_path: [create_record, search]
bulk_path: [import, export, reindex]
growth_lever: tenant_count
Choose a tenancy model on purpose
Three models are common. A shared schema with a tenant identifier is cheapest to operate and easiest to migrate. Schema per tenant isolates more and complicates migrations at volume. Database per tenant isolates almost completely and turns operations, reporting and cost into a serious engineering burden.
Choose based on customer expectations, data sensitivity and the operational capacity you actually have. Regulated buyers with contractual isolation requirements, and very large tenants with unusual performance profiles, both push away from the shared model. If you are unsure, start shared and keep every query tenant-scoped so a later move remains possible.
Record the decision and the conditions that would change it. Tenancy choices get revisited under pressure, which is the worst moment to rediscover the reasoning.
- Shared schema: simplest to run, isolation enforced in the data layer
- Schema per tenant: stronger isolation, migrations need orchestration
- Database per tenant: maximum isolation, highest operational cost
- Reversible data access patterns, so the model can be changed later
Make cross-tenant access a test failure, not a review question
The most expensive defect in a multi-tenant product is a query that quietly returns another customer’s data. It is rarely caught in review and frequently caught by a customer.
Enforce tenancy in the data layer rather than in each query, index with the tenant identifier leading so the database can use it, and write a test asserting that a record belonging to another tenant is not readable. That test should exist for every model holding customer data.
Background jobs and administrative tooling are where the discipline leaks, because they run outside the request context and quietly carry no scope at all.
- Tenant scope applied by default, with any opt-out explicit and reviewed
- Composite indexes led by the tenant identifier
- Cross-tenant access tests for every tenant-scoped model
- Queue workers carrying tenant context explicitly
- Administrative tooling reviewed separately from request paths
index: [organisation_id, created_at]
scope: organisation_id = :tenant
deny_by_default: true
Fan-out is the real scale problem
Systems rarely fail from one slow request. They fail from work that multiplies: a notification per record, a scheduled job per organisation, a reindex that visits every tenant inside a single process. Each is fine at small scale and ruinous at large scale.
Batch by tenant rather than by record, keep concurrency bounded per tenant so one large customer cannot occupy the whole worker pool, and move long operations onto a queue where they are resumable instead of blocking.
Watch queue depth and per-tenant lag. When the number of organisations grows faster than the workers, the failure mode is a growing backlog rather than a crash, and it stays invisible for a long time.
- Batch work per tenant rather than per row
- Per-tenant concurrency limits on the worker pool
- Resumable jobs with checkpoints rather than one long transaction
- Queue depth and oldest job age on the operational dashboard
Entitlements and limits belong in data, not in code
Plan limits and feature access get hard-coded, then fork. A customer on a grandfathered plan ends up on a code path that exists for exactly one account, and nobody can ever remove it.
Model entitlements as rows: a plan, a feature flag, a limit and an effective date. The product then reads the same definition for gating, for display and for enforcement, so the interface cannot promise something the backend will refuse.
Meter usage as it happens rather than counting at period end. A customer approaching a limit should see it approaching. Whatever you count, decide whether the count is advisory or blocking, and make that explicit per feature.
- Plan, feature and limit as versioned rows with effective dates
- One entitlement source used for gating, display and enforcement
- Usage metered continuously with visible progress against the limit
- Advisory and blocking limits clearly distinguished
- Grandfathering represented as data rather than a special code branch
Design for the noisy neighbour before it arrives
On a shared platform some tenants are heavier than others, and the damage is rarely uniform. One organisation running a bulk import can consume the queue, the database connections or the rate limit that everybody else depends on.
Give every tenant a share rather than first-come-first-served: per-tenant concurrency, its own queue priority, and quotas on the expensive operations. When a tenant exceeds its share, degrade it rather than the platform, and tell the customer what happened.
Scope caching to the tenant and evict per tenant. A shared cache key that forgets the tenant belongs to the same bug class as an unscoped query.
- Per-tenant concurrency and priority on shared resources
- Quotas on expensive operations with a documented response
- Cache keys and eviction scoped to the tenant
- Connection pools sized against the worst realistic tenant
Write the noisy-neighbour test before you need it. Simulating one heavy tenant is cheap; surviving one is not.
Takeaways
- Name the growth axis you are designing for, and revisit it with evidence
- Choose tenancy deliberately and record the conditions for changing it
- Enforce tenant scope in the data layer and test for cross-tenant reads
- Batch by tenant and bound concurrency per tenant
- Keep entitlements in data so grandfathering stays visible