NEWS

Cloudflare's K2 builds event streaming on top of R2 object storage

In public beta, K2 stores events directly in R2 instead of a dedicated log, promises 1 second of production latency, and has already sparked debate over pricing and real-world performance on Hacker News.

What K2 Is

Cloudflare recently launched K2 in public beta: a serverless event streaming service whose storage is not a dedicated log, but rather R2 itself, the company's object storage, according to reporting from InfoQ. Under the hood, K2 is a durable, partitioned log that lives inside R2, with event production (produce) at around 1 second of p99 latency, a figure published by Cloudflare itself.

The project was born as internal plumbing, not as a standalone product. Basin Pipelines, Cloudflare's pull-based processing engine, needed a durable place to park events before reading and transforming them. The obvious answer would be Apache Kafka, but Kafka assumes a fixed cluster with replicated local disk, and Pipelines runs on an edge spanning more than 335 cities, in ephemeral slices of small machines, connected over the public internet rather than an internal data center network.

R2 as a Log: The Architecture Behind K2

Cloudflare's answer was to stack the service on top of something that already solves durability: R2 itself, which offers eleven nines of durability with strongly consistent APIs. Pushing replication and consensus down to the storage layer keeps the application layer simpler, and allows compute and storage to scale independently.

Object stores don't support append, which is a problem when the entire abstraction is a log. K2 works around this by holding writes in memory in an edge service, waiting a moment to accumulate more events, and then writing the batch as a segment file in R2. R2's atomic operations guarantee ordering and incrementing offsets, so there's no separate coordination service (no Zookeeper or equivalent). That brief wait to accumulate the batch is exactly where the production latency comes from.

The Numbers: What Cloudflare Promises and What the Beta Showed

The reaction on Hacker News was, in general, favorable to the architectural pattern and pointed in its criticism of the numbers. Commenter psanford summarized the bigger thesis behind K2:

Object store is quickly becoming the new core data substrate. Lets build kafka, but on s3. Lets build github, but on s3.

psanford, commenter on Hacker News

Commenter necubi, who identified themselves in the thread as the author of the original post and tech lead for K2, agreed with that reading: every data system that doesn't require sub-100ms latency is migrating to object storage, and said their team works alongside the R2 team and can co-evolve both products. The original post is authored by Micah Wylde and Marc Selwan.

The sharpest part of the conversation came from people already running the beta. Commenter e1g reported sub-second writes, but end-to-end delivery latency of p95 at 2.5 seconds and p99 at 7.5 seconds, and asked whether that was expected. Someone responding on behalf of the team confirmed that tail latency is higher than expected, that the Consume API still has performance work ahead of it, and that read-latency improvements should land in about a week.

A competitor brought the technical counterpoint. Commenter sensodine, who identified themselves as a developer at s2.dev, argued that the real opportunity lies beyond 100ms, citing faster object storage tiers such as S3 Express and GCS rapid buckets (both single-zone, and therefore still dependent on write quorum for regional durability). They described the central tension of the design:

One of the tensions of course is how long to linger before flushing to object storage - you have to trade off directly between latency and cost of your API ops for PUTs.

sensodine, developer at s2.dev

According to him, s2.dev itself uses stateful backend processes, continuously flushing multi-tenant objects with records from multiple streams, reaching around 50ms of p99 acknowledgment latency within the same region, without destroying the service's unit economics.

Pricing: Producing Is Cheap, Consuming Doubles the Bill

Pricing also became a point of discussion. Commenter nnx found the $0.04 per GB produced reasonable compared to other cloud event-streaming services, but pointed out that charging the same rate for consumption quickly drives up real-world usage costs:

This means actual usage is $0.08/GB in the simplest case (one consumer) but fan-out consumer strategies get very expensive very fast.

nnx, commenter on Hacker News

That matters because fan-out, that is, multiple consumers reading the same stream, is exactly one of the use cases Cloudflare uses to position K2. Retention is billed separately, at $0.02 per GB per month, and nothing is charged during the beta. The same commenter made the cost case against the obvious alternatives: K2 would be cheaper than Google Pub/Sub, especially with long retention, and considerably cheaper than self-hosted or managed Kafka (such as Amazon MSK or Confluent), since it requires no cluster and keeps performance consistent as data volume grows.

Where K2 Fits In: Kafka, AutoMQ, WarpStream, s2.dev

Several questions in the thread were about how K2 differs from AutoMQ and WarpStream, which apply the same object-storage pattern on top of Kafka, and noted that Kafka itself already has experimental support for diskless topics. One answer, coming from another commenter rather than from Cloudflare, was blunt: K2 is built for the Cloudflare ecosystem, and using it in isolation, outside of Workers and Pipelines, doesn't make much sense today.

Kafka API compatibility is on the roadmap, but K2's tech lead was direct about their reservations:

I'm not personally a huge fan of the kafka API, I think it's simultaneously too low level for normal users and too high level to deeply integrate into other systems (like stream processing engines), and requires a complex client library to use effectively.

necubi, tech lead for K2 at Cloudflare

That's why K2 uses a lease-based consume API: the consumer polls a subscription, receives a batch with a five-minute lease, and then acknowledges (ack), rejects for redelivery (nack), or extends the lease. According to them, this design allows for greater read parallelism, important when consumers are Workers, which parallelize well but individually have little processing capacity.

Producing Events: The Code

On the production side, a Worker binding sends batches and exposes whether a failure is worth retrying:

js
const result = await env.EVENTS.send([
 {
 content: new TextEncoder().encode(
 JSON.stringify({
 event: "page_view",
 path: new URL(request.url).pathname,
 timestamp: Date.now(),
 }),
 ),
 headers: { "content-type": "application/json" },
 },
]);

if (!result.success) {
 console.error(`Produce failed: ${result.error.message}`);
 return new Response("Failed to record event", {
 status: result.error.retryable ? 503 : 500,
 });
}

K2 represents data as raw bytes, leaving encoding (JSON, protobuf, whatever) up to the application. Streams can be created through the dashboard, the API, Wrangler, or cf, the command-line interface Cloudflare launched the same week.

Beta Limits and What's Missing

The beta has clear limits: 10GB of storage and 30 MB/s of production per stream, with a form available to request an increase. Ordering by message key, write parallelism in the gigabytes-per-second range, push-based Worker consumers (instead of polling), a lower-latency express tier, and support for Kafka clients are all listed as future work, with no date attached yet.

What Changes for Builders in Brazil

For those already using Cloudflare Workers, Pipelines, or R2 in production in Brazil, K2 opens up a use case that previously required running a separate Kafka cluster, or paying for a managed service billed on fixed throughput. The pitch is clear: trade dedicated coordination and replicated disk for object storage and a wait of about one second before each batch becomes a segment.

But the numbers from the Hacker News thread matter more than the announcement: the end-to-end tail latency measured by a real user (p99 of 7.5 seconds) is considerably higher than the 1-second production latency Cloudflare announced, because producing isn't the same as delivering to the consumer.

Anyone evaluating K2 for a pipeline that can't tolerate minutes of delay, but also doesn't need sub-100ms latency, should test consume latency with their own workload before committing an architecture to it, and should pay attention to the fan-out cost: charging the same rate for producing and consuming means every additional consumer of the same stream doubles the cost, something pipelines with multiple teams reading the same topic feel quickly on the bill.

Translated from the Brazilian Portuguese original · Read the original