NEWS

Meta cuts ZippyDB connections 19x with the ZGateway proxy

The stateless layer built by Meta absorbs traffic from more than one million client hosts and sustains over 1 billion operations per second, without database servers seeing that load directly.

Meta detailed, in a technical post cited by InfoQ, the architecture of ZGateway, a stateless proxy layer built to solve a problem that only shows up at hyperscale: more than one million client hosts trying to talk directly to ZippyDB, the distributed key-value store that holds product metadata, counters, and configuration across the company's entire infrastructure. Today ZGateway already processes more than 1 billion operations per second and accounts for around 40% of ZippyDB's total traffic.

The problem: one million hosts hitting the database directly

In the direct-access model, each client connects to the replicas that serve the shards it needs to reach. Because a single client can touch tens of thousands of shards spread across hundreds of thousands of database hosts, the result is a many-to-many connection mesh too dense to scale safely. Meta reports that these "connection storms" contributed to file descriptor exhaustion and out-of-memory conditions on the servers.

This isn't a problem exclusive to hyperscale infrastructure. It's the same dynamic any team faces when scaling a service with a poorly sized connection pool: every new pod or instance opens its own connections, and the database becomes the bottleneck by managing sockets instead of serving queries.

How ZGateway works

ZGateway sits between ZippyDB clients and servers as a managed proxy. Clients keep "sticky" connections to regional gateway hosts, while database servers only receive connections coming from the controlled gateway fleet. Discovery of which gateway to use happens via ServiceRouter, Meta's internal service mesh, which keeps each client close to its regional gateway.

Architecture diagram showing ZippyDB clients connecting to ZGateway hosts, which act as a stateless proxy distributing requests to the ZServer fleet
Architecture diagram showing ZippyDB clients connecting to ZGateway hosts, which act as a stateless proxy distributing requests to the ZServer fleet. Reprodução: infoq.com.

Internally, the gateway uses the same C++ ZippyDB client as its request engine and supports two modes: pure proxy and a read-through cache layer. Because ZGateway sees traffic from multiple clients at once, it can coalesce requests that isolated client libraries would never see across different processes, on top of authenticating and authorizing requests, applying per-tenant admission control, resolving shards, and using a local cache before forwarding everything to the ZServer replicas.

It's worth noting that this architecture doesn't replace what ZippyDB already did: managed sharding, replication, failure detection, and capacity management remain the database's responsibility. ZGateway adds a shared traffic management layer in front of those capabilities, not a rewrite of them.

The numbers Meta measured

Meta's model estimates that per-host connection counts drop between 97% and 98%, and that the total number of persistent connections across the infrastructure decreases by roughly 19 times. It's an order-of-magnitude difference: fewer open sockets, less overhead from repeated TLS handshakes, less pressure on file descriptor limits on each database server.

Md Shuvo, an architect at Sherbrook, summed up the trade-off in a LinkedIn post cited by InfoQ: "Adding an extra network hop actually improves overall latency, by freeing database nodes from the brutal overhead of connection management." It's counterintuitive at first glance, because every dev learns that fewer network hops is always better, but the point is that the cost of keeping tens of thousands of connections open per host outweighs the cost of a well-sized extra hop.

The overload test: who takes the hit when CPU runs out

Meta ran a controlled test pushing ZGateway above 90% CPU usage. Of roughly 1,350 tenant buckets, only six had traffic shed, while the rest processed 99.9% of requests without rejection. This shows the gateway working as a central overload protection point: instead of the entire database degrading, the system isolates the damage to a small, predictable fraction of tenants.

The rollout is also controlled with fine granularity: Meta can route traffic progressively by service, by shard prefix, by percentage, or by region, with a global kill switch to revert everything at once if something goes wrong. This kind of blast radius control is what separates a well-executed critical infrastructure migration from a risky bet.

What this changes for those building with less traffic

No one outside Meta is going to operate 1 billion operations per second, but the architectural pattern is the same one that already exists in the day-to-day tools of anyone running Postgres, MySQL, or Redis in production: PgBouncer and PgCat in front of Postgres, ProxySQL in front of MySQL, or a managed proxy like AWS's RDS Proxy. All of them solve exactly the problem Meta describes, at a smaller scale: dozens of Kubernetes pods scaling horizontally, each opening its own connection pool against the database, until the database runs out of memory managing sockets instead of answering queries.

The interesting angle of ZGateway isn't the existence of the proxy, but what it adds beyond forwarding bytes: batching and coalescing requests coming from different clients (something a simple pooler doesn't do), per-tenant admission control, and a configurable read-through cache. For anyone designing an internal service with multiple consumers hitting the same data, it's worth asking whether the current pooler only limits connections or also sees access patterns across clients to save real round-trips to the database.

Using a service mesh (ServiceRouter at Meta, Envoy or Linkerd in more common stacks) to keep client and proxy geographically close is also replicable: it reduces network latency and avoids an entire region depending on a gateway on the other side of the world.

What's still left open

Meta still doesn't route 100% of ZippyDB's traffic through ZGateway: that's a stated goal, not an accomplished fact. The company also signaled that it's exploring agent-operated controls, selective co-location of the gateway with ZServer to further reduce latency, and a multi-process architecture for stronger failure isolation. None of these items has a public timeline, and InfoQ doesn't detail how ZGateway compares in absolute latency against direct access beyond the aggregated connection-reduction numbers.

Translated from the Brazilian Portuguese original · Read the original

Read also