OpenAI Scales Unsharded PostgreSQL to 800 Million Users
OpenAI says one unsharded PostgreSQL primary and ~50 read replicas now serve millions of QPS for 800 million ChatGPT users, with five-nines availability and just one SEV-0 in a year.

Updated
Why it matters
- OpenAI's PostgreSQL load grew more than 10x over the past year, serving 800 million users via a single Azure PostgreSQL primary and nearly 50 read replicas.
- The setup delivers five-nines availability, low double-digit millisecond p99 latency, and only one SEV-0 incident in 12 months — during the ChatGPT ImageGen launch, when write traffic surged 10x amid 100 million signups in a week.
- New tables are banned from the PostgreSQL deployment; shardable write-heavy workloads migrate to Azure Cosmos DB, and OpenAI is testing cascading replication with Azure to scale beyond 100 replicas.
OpenAI says a single unsharded PostgreSQL primary and nearly 50 read replicas now handle millions of queries per second for ChatGPT's 800 million users — a scale many engineers assumed the 30-year-old database could not reach. Over the past year, the company's PostgreSQL load has grown by more than 10x, and it continues to rise, according to a detailed engineering post published by the OpenAI infrastructure team.
The setup runs on Azure PostgreSQL Flexible Server, with the primary in High-Availability mode alongside a hot standby and replicas spread across multiple geographic regions. OpenAI reports five-nines availability in production, low double-digit millisecond p99 client-side latency, and replication lag near zero. The company recorded only one SEV-0 PostgreSQL incident in the past 12 months — during the viral launch of ChatGPT ImageGen, when write traffic suddenly surged by more than 10x as over 100 million new users signed up within a week.
The disclosure matters beyond OpenAI. PostgreSQL, first created by a team of scientists at UC Berkeley, is the default database for a huge share of the software industry, and the received wisdom has long been that workloads of this magnitude require sharded or distributed databases. OpenAI's experience shows that a single-primary architecture, aggressively optimized, can serve one of the largest consumer applications on the internet — while also documenting precisely where the approach strains.
Cracks in the initial design
After ChatGPT launched, traffic grew at what OpenAI describes as an unprecedented rate. The team scaled up instance sizes, scaled out read replicas, and optimized both the application and database layers. That architecture "has served us well for a long time," the team writes, and with ongoing improvements it "continues to provide ample runway for future growth."
SEV incidents followed a familiar pattern: an upstream issue — widespread cache misses from a caching-layer failure, a surge of expensive multi-way joins saturating CPU, or a write storm from a new feature launch — spikes database load. Query latency rises, requests time out, and retries amplify the load, creating "a vicious cycle with the potential to degrade the entire ChatGPT and API services."
The deeper constraint is PostgreSQL's multiversion concurrency control (MVCC) implementation, which the team says makes the database less efficient for write-heavy workloads. When a query updates a tuple or even a single field, the entire row is copied to create a new version. Under heavy write loads this produces significant write amplification and read amplification, since queries must scan through dead tuple versions to find the latest one. MVCC also brings table and index bloat, increased index maintenance overhead, and complex autovacuum tuning. The post links to a deep-dive co-written with Prof. Andy Pavlo of Carnegie Mellon University, "The Part of PostgreSQL We Hate the Most," which is cited on the PostgreSQL Wikipedia page.
Why not shard?
OpenAI's answer to write pressure is not sharding PostgreSQL but moving shardable, write-heavy workloads to sharded systems such as Azure Cosmos DB. New tables are no longer allowed on the current PostgreSQL deployment; new workloads default to the sharded systems.
The primary rationale for keeping PostgreSQL unsharded is engineering economics: sharding existing application workloads would require changes to hundreds of application endpoints and could take months or even years. Since the workloads are primarily read-heavy and heavily optimized, the current architecture retains ample headroom. Sharding PostgreSQL is not ruled out for the future, but it is not a near-term priority.
Load reduction on the primary
The team minimizes both reads and writes on the single writer. Read traffic goes to replicas wherever possible; reads that must stay on the primary because they sit inside write transactions get scrutinized for efficiency. On the write side, OpenAI fixed application bugs that caused redundant writes, introduced lazy writes to smooth traffic spikes, and enforced strict rate limits on backfilling table fields — a process that can take over a week but avoids production impact.
Query optimization
Expensive queries are treated as a systemic threat. One query OpenAI identified joined 12 tables, and spikes in that query were responsible for past high-severity SEVs. The team's guidance: avoid complex multi-table joins, break queries down, and move complex join logic into the application layer. Many problematic queries come from Object-Relational Mapping frameworks, so the SQL they generate gets carefully reviewed. Long-running idle queries are managed with timeouts such as idle_in_transaction_session_timeout to keep them from blocking autovacuum.
Mitigating the single point of failure
A single writer is a single point of failure, so OpenAI offloaded most critical read-only requests to replicas. If the primary goes down, writes still fail, but the impact drops below SEV-0 because reads remain available. The primary runs in HA mode with a hot standby that can be promoted quickly, and the Azure PostgreSQL team has done significant work to keep those failovers safe and reliable under very high load. Each region runs multiple replicas with sufficient capacity headroom, so a single replica failure does not cause a regional outage.
Workload isolation
To counter the "noisy neighbor" problem — where a new feature launch introduces inefficient queries that eat PostgreSQL CPU and slow other critical features — OpenAI splits requests into low-priority and high-priority tiers routed to separate instances. The same isolation strategy applies across products, so activity from one product does not affect the performance or reliability of another.
Connection pooling
Each Azure PostgreSQL instance caps connections at 5,000, and OpenAI has had incidents from connection storms that exhausted the pool. The fix is PgBouncer, deployed as a proxy layer in statement or transaction pooling mode. In OpenAI's benchmarks, average connection time dropped from 50 milliseconds to 5 milliseconds. Proxy, clients, and replicas are co-located in the same region to minimize network overhead, and each read replica runs its own Kubernetes deployment of multiple PgBouncer pods, load-balanced behind a Kubernetes Service. Idle timeouts and careful configuration remain essential.
Caching with lock-and-lease
When cache hit rates drop unexpectedly, the burst of misses hits PostgreSQL directly and can saturate CPU. OpenAI implemented a cache locking and leasing mechanism: only one reader that misses on a given key fetches the data from PostgreSQL, while all other requests waiting on the same key block until the cache is repopulated. This eliminates redundant database reads during cache-miss storms and protects against cascading load spikes.
Scaling toward a hundred replicas
The primary streams Write Ahead Log data to every read replica. Nearly 50 replicas already strain network bandwidth and CPU, producing higher and less stable replica lag. OpenAI is collaborating with the Azure PostgreSQL team on cascading replication, where intermediate replicas relay WAL downstream. The approach could scale the system to over a hundred replicas without overwhelming the primary, though it adds operational complexity around failover management. The feature is still in testing and will only roll out to production once it fails over safely.
Rate limiting and schema discipline
Rate limiting is layered across the application, connection pooler, proxy, and query tiers, with deliberate spacing of retry intervals to avoid retry storms. The ORM layer supports rate limiting and, when necessary, full blocking of specific query digests — targeted load shedding that enables rapid recovery from surges of expensive queries.
Schema changes follow strict rules. Even a small change such as altering a column type can trigger a full table rewrite, so only lightweight operations are permitted, with a hard 5-second timeout enforced. Index creation and drops must run concurrently. New features needing new tables must use alternative sharded systems such as Azure Cosmos DB rather than PostgreSQL.
The road ahead
OpenAI frames the results as proof that, with the right design and optimizations, Azure PostgreSQL can handle the largest production workloads, with sufficient capacity headroom for continued growth. The remaining write-heavy workloads are harder to shard, but the team is actively migrating them to sharded systems to further offload the PostgreSQL primary, and work with Azure on cascading replication continues. Looking ahead, OpenAI says it will explore additional scaling approaches — including sharded PostgreSQL or alternative distributed systems — as its infrastructure demands keep growing.
Original: learn.microsoft.com
More from Rebecca Stone
Show full bio
Correspondent covering consumer brands and retail at AI In Context.
135 articles
Related articles
- OpenAI Says a Quarter of U.S. Workers Now Use ChatGPT on the Job
- OpenAI Launches ChatGPT for Financial Services With Built-In Data
- OpenAI: Enterprise ChatGPT Messages Up 8x as AI Use Deepens
- OpenAI Launches ChatGPT Go Worldwide With GPT-5.2 Instant Access
- OpenAI Expands Data Residency Options for Enterprise Customers