NEWS

WSL3 Is Up to 61% Faster Than WSL2 in Memory, But Only 4% in Real-World Builds

An independent benchmark compared WSL's kernel 6.6 and kernel 6.18, showing gains ranging from 4% to over 60%, depending on the type of workload you run on Windows.

What Changed Under the Hood

The WSL 3.x stack brought two changes that stay hidden behind the version number: the virtualization runtime was updated, and the embedded Linux kernel jumped from the 6.6 LTS branch to 6.18. That's a whole kernel generation change, not an incremental patch.

Microsoft tends to announce generic performance gains with every WSL update, but that kind of marketing rarely says what really changes for people who compile code, run a local database, or spin up containers every day. That's the gap developer Tony Metzidis tried to close with a battery of side-by-side benchmarks, published on his blog and discussed on Hacker News.

How the Test Was Set Up

Metzidis compared WSL 2.6.3.0 (Linux kernel 6.6.87.2) against WSL 3.0.2.0 (kernel 6.18.40.1), using an Alpine Linux 3.23.0 rootfs with Go 1.27.1 installed on both sides. The host machine was an HP ProDesk 400 G4 Desktop Mini with an Intel Core i5-8500T (6C/6T, 2.10 GHz base) and 16 GB of DDR4 RAM, running Windows 11 Pro.

The .wslconfig file fixed the resources available to the VM in both tests:

  • processors=2 (2 fixed vCPUs)
  • memory=4GB
  • networkingMode=nat
  • firewall=true

With identical hardware and configuration in both runs, any difference in results comes from the kernel and the virtualization runtime, not the machine.

Where the Gain Is Biggest: Memory and Context Switching

The kernel microbenchmarks, run with the perf bench suite, isolate scheduler latency, IPC throughput, and memory bandwidth without relying on userspace tools:

Table of kernel microbenchmarks comparing syscall latency, context switch, hackbench, and memory bandwidth between WSL 2.6.3 and WSL 3.0.2
Table of kernel microbenchmarks comparing syscall latency, context switch, hackbench, and memory bandwidth between WSL 2.6.3 and WSL 3.0.2. Reproduction: tonym.us.
BenchmarkWSL2 (kernel 6.6.87)WSL3 (kernel 6.18.40)Variation
getppid() throughput1,474,691 ops/s1,533,399 ops/s+3.98%
Syscall I/O latency0.6781 µs/op0.6521 µs/op-3.83%
sched/pipe (100k ping-pong)48,445 ops/s54,589 ops/s+12.68%
Context switch latency20.64 µs/op18.32 µs/op-11.25%
sched/messaging (hackbench, 20 groups, 800 tasks)14.507 s13.047 s-10.06%
mem/memcpy (bandwidth, 1 GB)7.84 GB/s12.64 GB/s+61.35%

The memory bandwidth jump is the biggest number in the entire table. According to Metzidis, a direct contributor is the page_reporting.page_reporting_order=5 flag present in the kernel 3.x command line: increasing page reporting granularity reduces the frequency of traps to the hypervisor and the overhead of walking page tables when the guest interacts with Hyper-V's dynamic memory reclamation.

The 11% drop in context switch latency and the 10% drop in hackbench, meanwhile, come from scheduler refinements in kernel 6.18 combined with tighter handling of synchronous Hyper-V VMBus interrupts, which reduces the penalty of waking one vCPU from another.

The Test That Matters: Compiling a Real Go Project

A kernel microbenchmark is one thing; a real build is another. To measure the impact on an actual workflow, Metzidis compiled the GoReleaser codebase with go build -a -o /dev/null, with the compiler cache cleared via go clean -cache and the module cache already warm:

MetricWSL2WSL3Difference
Wall-clock time3m 11.27s3m 03.80s-7.47s (-3.91%)
User (CPU) time4m 58.99s4m 49.38s-9.61s (-3.21%)
Kernel (sys) time0m 57.39s0m 48.22s-9.17s (-15.98%)

Time spent inside the kernel dropped almost 16%, because the Go compiler fires repeated openat, newfstatat, mmap, and futex calls while scanning the package graph and coordinating compilation workers. Kernel 6.18 reduced the overhead of these calls and improved slab caching, which shaved more than 9 seconds off kernel time alone.

But total wall-clock time barely moved, around 4%. The reason, according to the benchmark's author, lies in the type of work that dominates compilation:

The reason total real time only moved by ~7.5 seconds is straightforward: compilation is predominantly a user-space compute workload. Lexing, AST construction, type checking, and SSA optimization account for over 80% of the total clock cycles.

Tony Metzidis, author of the WSL2 vs WSL3 benchmark

With the VM locked to two vCPUs at a 2.10 GHz base clock, the real bottleneck is processor frequency, not the kernel underneath.

What This Means for Windows Developers

The practical takeaway depends on the type of workload you run inside WSL. Anyone maintaining local environments with Redis, SQLite, or any workload that exchanges payloads heavily via Unix domain sockets feels the gain almost immediately: these are exactly the scenarios that benefit from the +61% memory bandwidth and the 10 to 11% lower context-switch latency.

Those who spend the day waiting on go build, cargo build, tsc, or any toolchain that spends most of its time processing code in user space will notice less difference on the wall clock. WSL3 cuts kernel time for these builds substantially, but the ceiling remains the number of cores and the frequency allocated to the VM via .wslconfig.

It's worth noting that the test used a 2018 i5-8500T with only 2 vCPUs allocated to the VM, hardware considerably more modest than that of a current development machine. On more recent processors, with more cores available to WSL, the compute-bound ceiling cited in the benchmark tends to recede, which could make the percentage gain in builds even larger than the ~4% measured here.

What Remains Open

The benchmark comes from a single machine, with a specific set of tests, and the author himself acknowledges that results vary by workload profile. The publication includes no data on disk I/O, Docker container usage inside WSL, or a comparison on more recent hardware with more cores allocated to the VM.

For those who want to reproduce or dig deeper, the author also published related benchmarks on the security cost of kernel mitigations in WSL3 and on how to recover performance in older WSL2 configurations, listed on the original page itself.

Translated from the Brazilian Portuguese original · Read the original