PostgreSQL with direct I/O keeps backups from worsening query latency
ClickHouse detailed how it swaps page cache for direct disk reads in managed Postgres backups, to stop stealing memory and CPU from production queries.
The backup reads the same disks that serve the query
In a post published on October 2, 2026 on ClickHouse's engineering blog, engineer Kaushik Iska describes a problem any DBA recognizes: backing up a production database doesn't wait for a quiet moment. It reads the entire data directory while the database is serving traffic, on the same NVMe disks, through the same kernel, that serve every query.
In ClickHouse Managed Postgres, when the instance has more than one local NVMe disk, they're combined with mdadm --level=0 into a RAID0 volume mounted at /dat, where Postgres lives. The backup agent, wal-g, points at the same directory and ships the files to an object storage bucket per timeline. One disk, one set of drives, two workloads competing at the same time.
Buffered reads are a silent tax on Postgres
By default, every file read on Linux goes through the page cache: the kernel keeps a copy in memory in case someone asks for it again soon. That helps almost every program except a backup, which reads each file exactly once, hands it to the uploader, and never touches it again. Because the page cache is shared and finite, the backup's stream of hundreds of gigabytes pushes out data that Postgres had in memory.
Postgres's shared_buffers stays protected, because it lives in pinned huge pages. But everything that depends on the operating system's cache (relation files that don't fit in shared_buffers, freshly written WAL, temporary files, visibility maps) is exposed. In ClickHouse's test, a buffered-read backup evicted the entire 40 GiB of a table that had been read minutes earlier and was merely idle, not cold.
Direct I/O is the flag that tells the kernel to skip the cache for those reads.
Kaushik Iska, engineer at ClickHouse
The obvious fix creates a new problem
Turning on direct I/O in wal-g is one configuration line: WALG_DIRECT_IO=true. The process opens the file with the O_DIRECT flag, data goes straight from the disk block to the process's buffer, and nothing stays in the page cache. The eviction problem disappears.
But direct I/O also turns off the kernel's readahead. In buffered reads, the kernel notices a sequential pattern and reads ahead in large blocks, keeping the device's queue full. Under O_DIRECT, the process receives exactly the bytes it asked for and nothing more. On a single disk this is manageable; on a RAID0 it's a throughput cliff.
ClickHouse's RAID0 splits data into 512 KiB chunks, distributed across the member disks. A small direct read fits inside a single chunk, which lives on a single disk: the others sit idle for that request. wal-g's default read size is 32 blocks of 4 KiB, that is, 128 KiB, smaller than the chunk itself. Turning on direct I/O with nothing else makes the backup slower, not faster.
Sizing the read to the whole stripe
The fix is to make each direct read cover the array's entire stripe. If a read spans one chunk on each member disk, all disks take part in the request and the array behaves like the parallel device it is. The variable that controls this is WALG_DIRECT_IO_BLOCK_COUNT, the number of 4 KiB blocks per request:
| Disks in RAID0 | WALG_DIRECT_IO_BLOCK_COUNT | Read size |
|---|---|---|
| 1 | 256 | 1 MiB |
| 4 | 1024 | 4 MiB |
| 8 | 2048 | 8 MiB |
The disk count comes from the actual list of storage devices the server discovers, excluding the boot volume. A direct read should be at least as wide as the array: anything smaller leaves part of the stripe idle on every request.
How many parallel readers, and why that changes with direct I/O
Read size is one parameter; the number of parallel readers (WALG_UPLOAD_DISK_CONCURRENCY) is another, and the right value depends on where the bottleneck is. Under buffered reads, ClickHouse uses half the vCPUs, because the kernel's readahead already keeps the device busy on its own. Under direct I/O there's no readahead: each reader submits a request, waits, submits the next, and disk utilization depends directly on how many readers are in flight.
That changes the math in two cases. In dense-NVMe instance families (such as AWS's i8g, i8ge, i7i, and i7ie instances, which carry a lot of local storage relative to compute), each device only hits its ceiling with many requests outstanding, and ClickHouse uses the full vCPU count. On very small servers, with two vCPUs or fewer, half would mean a single synchronous reader, unable to keep the device busy, so the full count is used there too. Outside those two cases, doubling the number of readers would cost Postgres CPU for little gain in backup throughput.
What the numbers show
The test ran on an i8ge.12xlarge instance in AWS's us-east-1 region: 48 vCPUs, 384 GiB of RAM, four local NVMe drives in RAID0 with a 512 KiB chunk, Postgres 18.6 with shared_buffers = 96GB, wal-g v3.0.9, and a 467 GB pgbench database larger than the machine's RAM. Sixteen clients run continuous point queries, and a separate 40 GiB table is read into cache and left idle before each round, simulating data that was hot minutes earlier.
| Configuration | Backup time | Disk read | Idle table evicted | p99 during backup |
|---|---|---|---|---|
| A, buffered, 24 readers | 70 s | 6.0 GB/s | 40 GiB | 0.18 ms |
| B, default direct I/O (128 KiB) | 96 s | 5.8 GB/s | 0 GiB | 0.06 ms |
| C, production configuration (4 MiB, 48 readers) | 71 s | 7.7 GB/s | 0 GiB | 0.06 ms |
| D, wide read, 24 readers | 75 s | 7.4 GB/s | 0 GiB | 0.06 ms |
With no backup running, the baseline p99 was 0.04 ms. The buffered backup pushed that number to 4.7 times the baseline and left a trail of cold reads for minutes after it finished, because the cache had to rebuild. The three direct I/O setups reached 1.5 times the baseline and returned latency to normal as soon as the backup ended. CPU usage also dropped: from 65% busy with buffered reads to 60% with direct I/O sized to the stripe, about 14% less for the same backup.
Three limits, not just one
A backup competes with the database for three resources: CPU, memory, and the disks. ClickHouse treats the three separately. The backup runs in its own cgroup, started with systemd-run --scope at CPUWeight=25 against the default weight of 100, so the scheduler gives Postgres four times more CPU when the machine is busy. wal-g's upload buffers are sized to stay under 5% of RAM. And direct I/O is the limit for the page cache and the disks: nothing the backup reads enters the cache, so whatever Postgres already had there stays intact.
What's left open
Every backup described in the post is a full one: it reads all 467 GB and ships about 54 GB compressed, regardless of whether one row or a billion rows changed since yesterday. For a multi-terabyte database, that's the largest cost line in keeping backups running. ClickHouse says it's prototyping incremental backups, which would read the data directory, identify the pages changed since the last backup, and ship only those, leaving the read path (direct I/O, stripe-sized reads, cgroup weights) unchanged.
Why this matters for anyone running Postgres
For anyone administering Postgres in production, the point isn't to copy ClickHouse's exact numbers, which depend on specific hardware, but to understand the reasoning behind each variable before flipping on O_DIRECT in any backup tool (pg_basebackup, wal-g, pgBackRest, or equivalent). Three questions are worth asking before touching this:
- Is the disk array a striped RAID0, or a single disk? That determines whether increasing read size is worth it.
- Does the backup tool offer control over read size and concurrency, or just an on/off switch for direct I/O?
- Does current monitoring capture page cache eviction and query p99 during the backup window, or does it only check whether the backup finished successfully?
If the answer to the third question is no, that's the first fix to make, regardless of any I/O configuration: without visibility into what the backup costs the rest of the database, any read optimization is a shot in the dark.
Translated from the Brazilian Portuguese original · Read the original
Thread pool in Percona Server: how to configure it to gain throughput under high concurrency
A Percona benchmark shows that thread pool can multiply throughput by nearly 18 times in high-concurrency scenarios, but it can also worsen performance if misconfigured. Understand the parameters that determine the outcome.