NEWS

Cursor uses write-ahead log on S3 to scale Git to more than 300 pushes per second

The company behind the AI code editor created Continuity, an architecture that trades coordination between replicas for a durable write log on S3, solving a bottleneck that affects any platform with many repositories and many simultaneous pushes.

Cursor, the company behind the AI code editor of the same name, published details of a new Git storage architecture called Continuity. According to a report from InfoQ on September 30, 2026, the system uses a write-ahead log (WAL) backed by S3 as the source of truth, instead of the coordination between replicas that traditional Git hosting architectures use to ensure consistency.

The reported numbers stand out: in synthetic tests, Cursor recorded linear read scaling with up to 100 replicas and more than 300 pushes per second using S3 Express One Zone. For any team that has watched a monorepo stall under CI load or a code agent system generating dozens of small repositories per hour, the problem Cursor solved is recognizable.

The bottleneck that drove the redesign

Git hosting at scale has historically relied on local Git repositories and replicated packfiles. GitHub's Spokes architecture maintains multiple NVMe replicas and uses a three-phase commit to update references. The model works: replicas stay in sync and reads can be served from any of them.

The problem appears as the number of replicas grows. Three-phase coordination means that the more replicas participate in a write, the greater the overhead to confirm they all agree before acknowledging the push. This cost scales poorly precisely in the scenario that AI tools and intensive CI push toward: many small repositories or a monorepo with hundreds of pipelines running at the same time.

How Continuity changes the model

Instead of replicating data across nodes and coordinating every write, Continuity treats S3 as the source of truth and local NVMe repositories as a warm cache. A push's data goes to S3 and the corresponding reference update is recorded in the WAL; the push is only acknowledged after that data is persisted, which guarantees durability before the client receives confirmation.

The engineering behind this combines several specific techniques:

  • Operation batching: groups writes to reduce the impact of S3 PUT latency on overall throughput.
  • Rendezvous hashing: selects the preferred nodes for each repository, without requiring centralized coordination.
  • Atomic compare-and-swap on S3: allows any server to accept a push, since consistency is guaranteed by the storage itself rather than by a replica coordinator.
  • Gossip over UDP: propagates WAL updates between nodes asynchronously, with conditional reads on S3 serving as verification, taking, according to Cursor, less than 10 milliseconds on average.

A repository can be rebuilt (materialized) directly from the WAL when the local copy is not available. And, according to the company, the loss of gossip messages does not compromise the correctness of the system, because S3 remains the source of truth, regardless of what nodes know about each other in real time.

Git treated as a database

The design drew direct comparisons to database systems. Maksim Al Dandan, a senior software engineer, described the approach as treating Git storage like a database, pointing to push consistency, force-push transactions, and point-in-time recovery as relevant issues for this model, according to InfoQ.

Vicent Martí, an engineer at Cursor, described local NVMe repositories as warm caches rather than authoritative copies, reinforcing that the real source of truth is S3 and not the local disk of any specific server.

Casey Lee, CTO of Liatrio and a former AWS engineer, highlighted the architectural shift of placing the WAL on S3 as the source of truth, with local NVMe serving only as cache. Lee also noted, according to InfoQ, that the performance numbers published by Cursor have not yet been independently verified by third parties. It's a point worth keeping in mind: the benchmarks come from the company itself, tested on Cursor's internal monorepo (called everysphere), and not in an externally audited environment.

The numbers in context

In synthetic tests with the everysphere monorepo, Cursor reports up to 120 pushes per second using S3 Standard and more than 300 pushes per second with S3 Express One Zone, an AWS storage tier optimized for lower latency. At the higher rate, the bottleneck stopped being storage and became Git compaction, which according to the company is done only by the primary node, with replicas downloading the resulting packs from S3.

This detail matters: Cursor traded replica coordination for extra bandwidth and asynchronous validation against storage, but it did not eliminate bottlenecks, only moved them. The company states that the tested pushes were linearizable and persisted to external storage before confirmation, and that clones remained fully consistent.

What changes for those building infrastructure

The Continuity case is a case study of a pattern that appears increasingly in distributed systems: moving the consistency boundary from the replica layer to durable object storage, such as S3, and letting local nodes converge independently on that shared state. This trades synchronous coordination for asynchronous propagation, storage validation, and higher bandwidth consumption.

For teams operating platforms with many small repositories, whether because of code agents automatically generating projects or monorepos with hundreds of simultaneous CI pipelines, the practical lesson is not necessarily to replicate Continuity, but to recognize the same kind of design decision: exactly where the source of truth lives, and what can be treated as disposable cache that is reconstructible from it. It's the same question that distributed database systems have been asking for years, now applied to the Git storage layer.

Casey Lee's caveat about the lack of independent verification of the numbers serves as a caution for anyone evaluating this approach for their own use: the results exist, but they have not yet gone through the external scrutiny that benchmarks of this kind usually receive when adopted outside the company that produced them.

Translated from the Brazilian Portuguese original · Read the original