Understanding Go 1.26's Green Tea GC

Go 1.26 brought one of the biggest changes to the runtime in recent years: the Green Tea garbage collector is now enabled by default. In this post, we'll look at what changed, why it matters, and what to expect in practice.
What is the Green Tea GC?
The Green Tea GC is a redesign of Go's garbage collector. It keeps the same mark-sweep approach as the previous GC, but fundamentally changes how objects are tracked and scanned.
The main difference: instead of operating object by object scattered across the heap, Green Tea works at the level of memory pages. It groups objects into contiguous 8 KiB blocks called spans and scans several objects at once within the same span.
How it works in practice
The focus is on small objects (up to 512 bytes), which are the most common and the most expensive to scan individually.
When the GC finds an object that needs to be scanned, it doesn't do so immediately. Instead, it marks the object's position within its span and waits until several objects accumulate in the same span. When it processes the span, it scans all of them at once.
This approach significantly improves cache locality: accessing contiguous memory is much faster than jumping between addresses scattered across the heap.
Work distribution across cores
The previous GC used a global queue to distribute work among goroutines. This created contention when multiple cores tried to access the queue at the same time.
Green Tea solves this with per-worker local queues. Each worker has its own queue of spans to process. When a worker becomes idle, it “steals” tasks from other workers, a pattern known as work stealing. This eliminates the central bottleneck and scales better with more cores.
Performance results
The numbers are encouraging:
-
A 10% to 40% reduction in GC overhead in real-world programs with heavy collection use.
-
An additional ~10% on modern AMD64 platforms, thanks to the use of vector instructions (AVX-256 and AVX-512) to speed up scanning of small objects.
This is not a synthetic benchmark: these are gains observed in production applications.
The trade-off: more resident memory
It's not all good news. Initial benchmarks show an 8% to 15% increase in RSS (Resident Set Size). This is the cost of the span-based scanning strategy: keeping more metadata per page consumes more baseline memory.
For most applications, this trade-off is worth it. But if your use case requires the smallest possible memory footprint, it's worth monitoring.
How to disable it
If you need to go back to the previous GC, just compile with: GOEXPERIMENT=nogreenteagc go build ./...
But the recommendation is to test with Green Tea enabled first: the gains are real for most scenarios.
Relationship with other Go 1.26 optimizations
The Green Tea GC complements speculative stack allocation, which had already been available since Go 1.25. While stack allocation reduces the number of objects that reach the heap, Green Tea improves the collection efficiency of the objects that inevitably end up there.
Combined, these two optimizations attack the GC performance problem from two angles: less garbage to collect and faster collection of the remaining garbage.
Conclusion
The Green Tea GC is the kind of improvement that demonstrates the maturity of the Go runtime. Without changing the API and without breaking compatibility, the Go team delivered significant performance gains that you get simply by updating the toolchain.
Translated from the Brazilian Portuguese original · Read the original



