Dev & EngARTICLE

In PostgreSQL, a low min_wal_size can double tail latency

PostgreSQL's min_wal_size parameter doesn't retain WAL for a standby, it only avoids recreating segments on demand. A load test by consultant Christophe Pettus shows what this costs when the floor is set too low.

Few PostgreSQL parameters are as misleading by name as min_wal_size. In a post from the "All Your GUCs in a Row" series, published on Planet PostgreSQL, consultant Christophe Pettus (PGX Inc.) takes apart the most common reading of the parameter: that it controls how much WAL is available for a standby to reconnect. It doesn't. That's handled by wal_keep_size or a replication slot.

What the parameter actually does

The description in pg_settings is more precise than the one in the documentation: min_wal_size sets the minimum size that pg_wal can shrink to. It's a recycling floor, not a read reserve. Pettus shows a primary with min_wal_size = 1GB and a standby that's been down for 640 MB: the pg_wal directory has 1.2 GB and 76 files, but 75 of them are old segments renamed to future names, full of stale bytes waiting to be overwritten. When the standby comes back, it gets FATAL: could not receive data from WAL stream because the segment it needs was already removed long ago.

This means that setting min_wal_size high is not a high-availability strategy. There's a different mechanism for that, with a different responsibility. Confusing the two is the most common mistake among those who tune this parameter expecting slack for replication.

Where the floor comes from and how it fits into autotuning

min_wal_size arrived in PostgreSQL 9.5 along with max_wal_size, when Heikki Linnakangas replaced the old checkpoint_segments. The default is 80MB (five 16 MB segments); initdb writes that line into postgresql.conf and scales it with the segment size, so a cluster created with --wal-segsize=64 gets a 320MB floor. The context is sighup, so it changes on reload, without restarting the server.

The underlying mechanism was already described in the series' previous post about max_wal_size: at the end of each checkpoint, segments older than the new redo point are renamed to serve as future segments or removed. How many get renamed depends on a moving average of WAL written per checkpoint cycle, which jumps up at once when a cycle exceeds it and drops by a tenth per cycle when it doesn't. The server keeps enough files to cover 1 + checkpoint_completion_target cycles at that rate, plus 10%. max_wal_size is the ceiling for that number; min_wal_size is the floor. Setting them equal, as Heikki's original commit itself documents, turns off autotuning.

The pool only grows through recycling, never ahead of time

Two behaviors of the floor aren't in the name and surprise those who expect PostgreSQL to "prepare" the directory ahead of time. A freshly created cluster with min_wal_size = 4GB has a single 16 MB file in pg_wal, and the same holds for a replica coming out of pg_basebackup, which doesn't copy the recycling pool. A promoted standby recycles at most ten segments for the new timeline and discards the rest, so the server that just took over production starts with a 160 MB pool, regardless of the configured value.

Shrinking is lazy too: checkpoints never revisit segments ahead of the insert point, so the pool goes down by one file per WAL segment actually written through it. Pettus measured this after a 1.5 GB burst: five checkpoints with almost no writes in between lowered the internal estimate from 1.5 GB to 0.9 GB without removing a single file; it took another 1.5 GB written in a trickle, 16 MB per checkpoint, for the pool to reach five segments.

In practice, the floor only matters when the server writes a whole pool slowly, over enough cycles for the estimate to fall by half (which takes about seven checkpoints), and then suddenly gets busy again. The classic example is an overnight batch job followed by the morning traffic peak.

The measured cost: 21 ms per segment, 53 times in a row

When WAL goes past the end of the pool, the process that needs the next segment (almost always a regular backend, not the checkpointer) creates that file on the spot: zeroes out 16 MB, calls fsync, renames it, all while holding WALWriteLock. That blocks any other commit waiting in line. Pettus reproduced the scenario with pgbench (scale 50, eight clients, two minutes) after draining the pool down to the configured floor:

  • With min_wal_size = 80MB (5-segment pool): 3,396 tps, p50 of 2.1 ms, p99.9 of 22.8 ms, 53 segments created on demand, adding up to 1,103 ms of extra work.
  • With min_wal_size = 8GB (227-segment pool): 3,439 tps, p50 of 2.1 ms, p99.9 of 11.4 ms, zero segments created.

A second run confirmed the pattern (22.6 ms versus 11.2 ms). Throughput varied only 1% to 2%, within the normal noise between runs, and the median didn't move. What doubled was exactly the tail: zeroing out and syncing a new segment costs about 21 ms on the disk used in the test (0.3 ms fsync), not counting the fsync for the rename itself. On a network volume capped at 125 MB/s, the zeroing alone already costs 128 ms.

To see this in production, PostgreSQL 18 exposes segment creation in pg_stat_io, in the rows with object = 'wal' and context = 'init', with timing if track_wal_io_timing is on. Before version 18, the signal is indirect: the wait events WALInitWrite and WALInitSync (called WalInit… up through 17). log_checkpoints doesn't help: the "WAL file(s) added" line only counts the segment that the checkpointer sometimes pre-allocates for itself, not the ones created by backends under pressure.

When the floor becomes a risk: a full disk

The documentation calls the floor's space "reserved," and it really is, since recycled segments occupy actually allocated disk. Pettus tested the other side: filled a volume to 100% and kept writing. With the floor at 80MB, the server still wrote 78 MB of WAL before a PANIC: could not write to file due to lack of space; with the floor at 512MB, another 506 MB. Both cases failed crash recovery with the same message and went down. At the test's write rate, the difference between the two configurations was ten seconds versus one minute of margin before the collapse.

The validation that comes too late

pg_settings reports the accepted range as 2 to 2147483647, and the first number doesn't hold in practice: the real minimum is two segments, but that check only runs when the postmaster reads the control file, not when the value is applied:

sql
ALTER SYSTEM SET min_wal_size = '16MB';
SELECT pg_reload_conf();
SHOW min_wal_size; -- shows 16MB, server keeps running

The server operates normally with that value until the next restart, or until any backend crashes and the postmaster reinitializes everything:

LOG: all server processes terminated; reinitializing
FATAL: "min_wal_size" must be at least twice "wal_segment_size"
LOG: database system is shut down

This turns a simple OOM kill into downtime that only ends when someone manually edits the configuration file. Bharath Rupireddy proposed validating the limit at SET time, back in 2022; Tom Lane pointed out that a check hook can't validate one parameter against another, and the discussion stopped there. The behavior is the same through version 19 beta 4. With 64 MB segments, the real minimum rises to 128MB, above the compiled-in default of 80MB, which is why initdb writes the min_wal_size line explicitly into postgresql.conf on those clusters: commenting out that line stops the server from starting.

Two other combinations fail silently. With wal_recycle = off (the recommended setting for copy-on-write filesystems), min_wal_size becomes inert, because nothing gets recycled; Pettus shows a checkpoint removing 56 segments even with a 1GB floor configured. And a floor larger than max_wal_size is accepted with no error or warning: the ceiling simply prevails.

The side effect on pg_rewind

A high floor costs disk that max_wal_size would already claim in a checkpoint cycle anyway, that is, space that was normally already provisioned. The real cost Pettus found is elsewhere: pg_rewind copies the reserve segments from the source side along with the real ones (1,235 MB in his test, almost all pool). Version 19 learns to skip WAL that both sides already have, but still copies the entire pool, so every rewind gets slower in proportion to the floor's size.

The practical rule of thumb

For those running the database in production, the decision boils down to two situations. If the disk has already been sized for the worst case of max_wal_size, there's no reason to leave the floor low: setting min_wal_size equal to max_wal_size lets pg_wal grow to the peak and stay there, eliminating on-demand segment creation.

If max_wal_size is deliberately huge (to absorb rare bursts without forcing a checkpoint), the floor should reflect how much space the team is willing to permanently give up to pg_wal, not the entire ceiling. PGTune uses a quarter of max_wal_size as a reference in its server profiles; Pettus notes that nothing derives that proportion mathematically, but nothing invalidates it either.

Neither side is large, and knowing that is most of what there is to know about this parameter.

Christophe Pettus, PostgreSQL consultant at PGX Inc.

The only value that actually hurts is a floor below two segments, because then the server won't even come back up after a restart. Beyond that, the choice between a pg_wal that breathes (recycling down to the minimum) and one that stays parked at the ceiling is a small trade-off between idle disk and tail latency during post-batch bursts, not a setting that determines the database's overall health.

Translated from the Brazilian Portuguese original · Read the original