Instances & Trust
The kernel, metadata, and package model describes what an org is and how it renders. It says nothing about where an org physically runs. At scale that becomes its own architecture: many orgs cannot share one database forever, the platform must stay redundant and detect trouble before customers do, and all of that must be reported to customers honestly. This is the operations plane — the fleet of database instances orgs are distributed across, the health system that watches them, and the customer-facing trust site that publishes their status.
Instances — the unit of hosting
Section titled “Instances — the unit of hosting”An instance is a self-contained platform stack — a database and its compute — that hosts many orgs. Multi-tenancy (metadata ≠ data) is exactly what lets many orgs share one instance: the shapes differ per org, the machinery is shared. The platform is therefore not one giant database but a fleet of instances, each hosting a slice of the customer base.
Every org has a home instance. The same domain-resolves-the-org step that routes a request also resolves the org to its instance, so directing traffic to the right instance is already in the request path — an org’s placement is a property the router reads, not a lookup bolted on later. This mirrors the well-established instance/pod model of large multi-tenant platforms, where a customer lives on a named instance rather than on an undifferentiated cloud.
Scaling — when an instance is full
Section titled “Scaling — when an instance is full”An instance has finite headroom along several axes: storage, connections and compute, IOPS, and tenant count. As it approaches those limits the response is horizontal, not vertical:
- New orgs land on a fresh instance. Provisioning places a new org on an instance with headroom rather than overloading a full one.
- Existing orgs can migrate between instances to rebalance load or free a crowded instance.
- Regions are first-class. Instances live in regions; an org is placed on an instance in the region it requires — for data residency (EU, US, …) or for latency — and sustained growth in a region spins up additional instances there.
The concrete sizing thresholds — the point at which a single database instance is “too large” — are a decision backed by infrastructure research, listed under open decisions below.
Sharding — two distinct meanings
Section titled “Sharding — two distinct meanings”The word covers two different operations, and conflating them causes confusion:
- Instance-level distribution (the primary axis). Spreading orgs across many independent instances is how the platform scales to a large customer base. This is “sharding” at tenant granularity, and it is the common case.
- Intra-org sharding (later, for whale tenants). A single very large org whose data outgrows one instance may eventually need its own data partitioned across shards. This is rarer and later than instance-level distribution, and is called out so the two are never confused.
Redundancy and recovery
Section titled “Redundancy and recovery”Each instance needs replicas (read replicas to spread load, a standby for failover), backups, and point-in-time recovery. A managed database platform provides much of this out of the box; operating at platform scale extends it toward cross-zone and, where warranted, cross-region redundancy. How much is delegated to the managed provider versus built above it is an open decision.
Degradation detection
Section titled “Degradation detection”Every instance is continuously monitored — request latency, error rates, connection saturation, replication lag, storage headroom, and synthetic checks against a known-good path. Automated detection opens an incident the moment an instance degrades, alerts operators, and publishes the state to the trust site. The goal is twofold: catch degradation before a customer reports it, and report it honestly when it happens rather than letting it be discovered.
The trust site
Section titled “The trust site”A separately-hosted domain — its own infrastructure, deliberately not the platform’s — that publishes, for customers and prospects:
- Instance status — the real-time health of every instance (operational, degraded, incident, or under maintenance), with history.
- An org’s own instance — every org can see which instance and region it lives on, so a customer knows exactly which status to watch, the way large platforms surface a customer’s instance rather than hiding it.
- Incident and maintenance history — a transparent, durable record of what happened and when.
- Certifications and compliance — the platform’s security and compliance posture. Certifications belong here, as the authoritative place a prospect or auditor confirms them.
Why it is hosted separately. A status page that shares infrastructure with the thing it reports on goes dark exactly when it is needed most. The trust site must survive a full platform outage, so it runs on independent infrastructure and reads health through a resilient channel that does not depend on the platform being up. This is non-negotiable for a status surface: a trust site that fails with the platform is worse than none, because it also reports “all clear” while the platform is down.
Where this sits
Section titled “Where this sits”This plane is beneath the engine, not part of it. It connects to three things already described: provisioning places an org on an instance; the domain-resolves-the-org routing carries the request to that instance; and the infrastructure tier of licensing is denominated in the instances and services an org’s shape consumes. The engine is indifferent to which instance runs an org — that indifference is what makes the fleet possible.
Open decisions
Section titled “Open decisions”- The concrete per-instance sizing limits — storage, connections, tenant count — at which a database instance is “too large,” and the placement and rebalancing policy that follows.
- The mechanics of migrating an org between instances with no data loss and minimal downtime.
- The region strategy and the data-residency guarantees offered per region.
- Redundancy: which failover and recovery guarantees are delegated to the managed database provider versus built above it.
- The intra-org sharding strategy for whale tenants, and the threshold that triggers it.
- The trust site’s hosting and the resilient health channel it reads from so it survives a platform outage.
- The certification and compliance set the trust site publishes.