OpenAI's Habitat Storage Layer Hits 70 Million Requests a Second
Habitat serves 1 billion weekly users at 70M requests per second. OpenAI details its Python-to-Rust journey, asyncio bottlenecks, and a metastable failure traced to aiohttp's LIFO connections.

Updated
Why it matters
- Habitat handles 70M+ requests per second, serves 1B+ weekly users across nearly 40 regions, and stores over 500 petabytes
- In Q2 2026, 2 engineers using Codex and GPT-5.5 rewrote the service in Rust, now serving 95% of production traffic with 6x CPU and 15x memory efficiency gains
- A metastable failure traced to aiohttp's default LIFO connection reuse was fixed by patching the connection pool to FIFO, breaking the feedback loop concentrating traffic on degraded pods
OpenAI's internal storage platform Habitat now handles more than 70 million requests every second, serving products used by over 1 billion people each week across almost 40 geographic regions and storing more than 500 petabytes of data. The company detailed the system's evolution in an engineering blog post published as the first installment of a two-part series.
The numbers matter beyond OpenAI's walls. Habitat sits under ChatGPT, Codex, and every other OpenAI product, and every user action — logging in, checking Codex settings, starting a conversation — can trigger many separate data lookups. As OpenAI puts it: "If those requests are slow, the product feels slow. If those requests fail, the product stops working entirely." The post offers a rare look at how infrastructure must be rebuilt in real time under hypergrowth, a challenge every company chasing AI-driven traffic spikes now faces.
From Python library to distributed service
Habitat launched at DevDay 2023 to support GPTs. It started as a simple Python client-side library connected to a single database, mapping a small set of operations to Azure Cosmos DB. The idea was straightforward: product engineers shouldn't need to think about database management. The library handled schema lookup, routing, authorization, encryption, serialization, request shaping, and connection pooling — and even shielded engineers from knowing whether data came from Cosmos DB, caches, or other storage.
The library saw rapid adoption. But by mid-2025 it had reached its limits. Backward-compatible protocol changes had become infeasible as OpenAI's service count grew. One migration to regionally distributed Cosmos DB accounts, intended to shrink the blast radius of a regional outage, illustrates the pain: coordinating rollouts across dozens of services took days, adding shadowing took another couple of days, a bug fix took another couple of days — and then one team rolled back their service to a previously buggy client for unrelated reasons, "causing the outage we had worked so hard to avoid."
OpenAI decided to pull Habitat into its own service. The move created a single point of control for deployments, observability, and platform enhancements. It also created a single chokepoint for security: the service centrally enforces access control policies, performs audit logging, and limits access to underlying storage. "Habitat plays a critical role in protecting user data and preventing unauthorized access from external, internal, and agent actors."
The scale of the growth was unusual even by infrastructure standards. "Often, system engineers build for 10x scale, and hope for it to hold for a few years while preparing for the next 10x," the post notes. "In our case, we've grown more than 10x year-over-year for the last three years."
Deliberate technical debt, and a bet on AI
OpenAI kept the service in Python despite knowing the trade-offs. Python increased network latency and added substantial CPU and memory scaling costs, and the team acknowledged that "the inefficiencies of Python would not be acceptable at 100x scale, making an eventual rewrite almost certain." The company viewed this as a "strategic incursion of technical debt" — the priority was unblocking product developers and achieving platform stability, not cost optimization.
The team also made a calculated wager that its own coding models would ease a future migration. "We bet that by the time a full migration off Python was required, Codex and GPT would make that migration achievable," OpenAI writes. "That bet eventually proved correct."
In the meantime, the engineering team fought tail latencies on several fronts. Because the average user request results in hundreds of database calls, "the slowest database call is the one the user feels."
One root cause was asyncio scheduling delay. Python's asyncio enables concurrent I/O-bound work but not CPU parallelism, and Habitat's CPU-heavy responsibilities — routing, compression, encryption, checksumming, health checking, request shadowing, and hedging — meant that even modest numbers of concurrent requests per process produced significant scheduling jitter, "up to hundreds of milliseconds and in some edge cases several seconds." The team began empirically measuring event loop delay in real time and responded by keeping each process serving only a small number of concurrent requests while massively scaling out the number of Python worker processes.
Live CPU profiling found another culprit: Statsig, the feature flag tool, was polling for refreshed configs every minute with no jitter, and the config included every production rule across every service. With up to 8 Python processes per pod, every minute each pod hit a moment where all workers stalled to parse a giant configuration file. The fix: a smaller targeted config, a longer refresh interval, and jitter on background tasks.
A metastable failure hidden in aiohttp
The most instructive incident involved connection pooling. Before load-balancing adjustments, some tail processes served 5-10x the number of concurrent requests as the average. During one incident, despite stopping the client that was overloading part of the service, a subset of processes kept degrading until restarted — a metastable failure, a class of failures some team members knew from prior work at scale.
The cause was Python's aiohttp TCPConnector defaulting to LIFO connection reuse. During request bursts, requests to slower, overloaded servers returned connections to the pool later and were therefore selected more frequently by subsequent requests, gradually concentrating traffic on struggling pods. "Patching the connection pool to use FIFO reuse broke this feedback loop and even reduced our steady state request variance as well." Today, OpenAI mostly relies on Istio and Envoy for connection pooling and server-load-aware balancing across its infrastructure.
Envoy also solved another problem: with so many Python processes, it became easy to flood downstream dependencies in a "thundering herd" — a daily deployment could cause significant CPU churn from connection cycling, and a connection leak could saturate the NAT gateway. OpenAI uses Envoy to upgrade Python's HTTP/1 connections to HTTP/2, multiplex them, extend connection lifetimes, and centrally implement rate limits and circuit breakers.
Deliberately weak API, deliberately strong isolation
A key reason Python scaled as far as it did was Habitat's constrained API. Rather than arbitrary SQL, Habitat exposes a simple NoSQL API modeled around client-defined object and edge types, inspired by Meta's TAO. "We aim to optimize for simple, predictable, constant-work requests," the post explains, because requests with unpredictable fanout "complicate isolation, load balancing, and introduce latency cliffs that are hard to scale for both the service and its clients."
The lesson came from experience. Before Habitat and Cosmos DB, most of OpenAI's online data lived on Postgres, where "a single expensive new query on a hot path took out the database" often enough that query review became unmanageable. The problem, OpenAI says, is cost imbalance: "it is cheap and easy to write SQL queries that are expensive and hard to run." Habitat makes expensive queries exceedingly obvious client-side; there are no unbounded queries, and complex joins and graph traversals require product teams to do the heavy lifting.
For complex querying needs, Habitat streams changes via change data capture to isolated Rockset instances in near-real-time, with each client team scaling their own instance. This isolates online storage from read-heavy analytical and search workloads.
The Rust rewrite
At its Python peak, Habitat served more than 20 million requests per second as the second largest service by core count at OpenAI, and fourth for Envoy footprint. In Q2 2026, a team of just 2 engineers — with Codex and GPT-5.5 — rewrote the entire service in Rust. The Rust service now handles 95% of production requests, with Python deprecation due within weeks. OpenAI's data shows the Rust version is 6x more CPU efficient and 15x more memory efficient than Python, with significantly lower average and tail latencies.
The second post in the series will cover the storage layer itself: multi-tenancy reliability, layered read performance optimization, and the scaled partnership with Azure Cosmos DB. OpenAI also says the AI-assisted rewrite proves out the bet it made years earlier — that its own models would make the eventual migration achievable.
Original: engineering.fb.com
More from Sophie Lindqvist
Show full bio
Staff writer covering marketplaces and e-commerce at AI In Context.
139 articles
Related articles
- OpenAI Scales Unsharded PostgreSQL to 800 Million Users
- AWS and OpenAI Sign $38 Billion Multi-Year Compute Deal
- OpenAI Brings GPT-5.5, Codex, and Managed Agents to AWS Bedrock
- AWS and OpenAI Sign $38 Billion Multi-Year Compute Deal
- OpenAI ships a model-native harness and native sandboxes for its Agents SDK