Dev & EngARTICLE

Why max_sync_workers_per_subscription doesn't speed up initial copy in PostgreSQL

The parameter controls how many tables a subscription copies at the same time, not the speed of each individual copy, and changing it without understanding the cost on the publisher can stall VACUUM and exhaust replication slots.

The parameter controls how many tables a subscription copies at the same time, not the speed of each individual copy, and changing it without understanding the cost on the publisher can stall VACUUM and exhaust replication slots.

The parameter controls how many tables a subscription copies at the same time, not the speed of each individual copy, and changing it without understanding the cost on the publisher can stall VACUUM and exhaust replication slots.

In a new chapter of the "All Your GUCs in a Row" series, published on Planet PostgreSQL, the post dissected a parameter that is often misunderstood by those who administer logical replication in PostgreSQL: max_sync_workers_per_subscription. The text goes straight after a common myth: the parameter does not determine the speed of a subscription's initial copy, but rather how many different tables can be copied in parallel. Whoever raises the value expecting the database's largest table to copy faster is solving the wrong problem.

The parameter has existed since logical replication arrived in PostgreSQL 10, with sighup context (it accepts a reload without a restart), a range of 0 to 262143, and a default value of 2. None of that has changed through version 19, currently in beta. Its relevance for those running PostgreSQL 17 or 18 in production doesn't come from any recent change to the parameter, but from the fact that, year after year, it keeps being poorly sized in migrations and high-scale subscriptions.

What the parameter actually buys you

Each table synchronization worker copies exactly one table, from start to finish, via COPY ... TO STDOUT on the publisher and COPY FROM on the subscriber, with row-by-row index maintenance during the fill. There is no parallelism within a single table: if the largest table in the schema takes forty minutes to copy, it will take forty minutes whether the parameter is set to 2, 8, or 200. The value only defines how many different tables can be copied at the same time, subject to the ceiling imposed by max_logical_replication_workers (the pool these workers are drawn from). And the limit is per subscription: three CREATE SUBSCRIPTION commands initializing together need three times that number of free slots in the pool.

A practical point the text highlights: the reload really is dynamic. The apply worker re-reads the configuration in its main loop, so raising the value while the copy is already underway takes effect immediately, without needing to recreate the subscription. In a test by the author with PostgreSQL 18.6, a subscription with six tables started with the parameter set to 1 and, a few seconds after a reload raising it to 3, three synchronization workers were already running simultaneously. That gives a real emergency exit during a migration: raising parallelism mid-process without tearing anything down.

The cost that stays hidden on the publisher

The most important part of the article, for whoever administers the publisher side, is the inventory of cost per worker. Each synchronization worker is a full logical replication client and consumes, on the subscriber, a slot from the worker pool and a replication origin (against max_active_replication_origins from PostgreSQL 18 onward, or max_replication_slots before that). On the publisher, the cost is higher: each worker shows up as a WAL sender (consuming max_wal_senders), creates a permanent replication slot named following the pattern pg__sync__ (counting against the publisher's max_replication_slots), and opens a REPEATABLE READ transaction that stays open for the entire copy.

It's that open transaction that demands extra attention from anyone who treats the database as a critical asset. While the copy runs, that transaction's backend_xmin prevents VACUUM from removing any dead row in any table in that database, not just the table being copied. Eight simultaneous synchronization workers on a busy publisher mean eight long transactions holding back the vacuum horizon's advance at the same time. In databases with a high write rate, that's a recipe for bloat silently accumulating throughout an entire migration.

The order you don't choose, and the failure that restarts from scratch

The apply worker scans pg_subscription_rel in physical order and dispatches a worker for each table that isn't ready yet, whenever a slot is free. That means the copy order is neither by size nor by name: it's the order in which the publisher listed the tables at the time of CREATE SUBSCRIPTION, subject to drift as state changes turn into catalog updates.

The behavior in the face of failure is more serious. A copy that fails, for instance due to a unique key violation, restarts from scratch, and the new worker discards the old slot and creates another one on the publisher. If the interval between retries (wal_retrieve_retry_interval) has already elapsed, which happens when the copy took longer than five seconds, the restart is nearly immediate. With the default value of 2, a single problem table stuck in a retry loop ties up half the available parallelism for the rest of the migration, and the only visible sign is the sync_error_count counter in pg_stat_subscription_stats climbing with each attempt. Since PostgreSQL 15, there's disable_on_error = true on the subscription, which turns that silent loop into a single error and a disabled subscription, an option the text recommends enabling by default.

The pause nobody expects at the end of each copy

When a copy finishes, the worker still needs to catch up, starting from the copy's snapshot, to the leader apply worker's current point. While that happens, the apply worker enters the LogicalSyncStateChange wait event and stops applying changes to all tables, not just the one that just finished syncing. That worker's WAL sender decodes everything the publisher wrote during the copy, including changes from other published tables, only to discard whatever doesn't belong to the table just copied. In a test by the author with two concurrent update workloads during a 34-second copy, the synchronization worker's WAL sender ended up 681 MB behind the publisher, and the leader's own slot fell 40 MB behind while waiting, producing a pause of roughly two seconds in applying changes. On busy publishers, with copies that take the whole afternoon, that pause grows proportionally, once per table.

The two extremes: zero and PostgreSQL 19

The value 0 is a silent trap. With max_sync_workers_per_subscription = 0, CREATE SUBSCRIPTION is accepted, the apply worker comes up, but no table ever leaves the srsubstate = 'i' (initial) state, and nothing gets logged because nothing is actually failing: the apply worker simply never requests a worker. Rows inserted on the publisher in the meantime are discarded by the leader, which only applies changes to tables that are already ready. The post's author found no legitimate use for this value.

PostgreSQL 19, still in beta, introduces sequence synchronization, done by a single worker per subscription for all sequences, and that worker is drawn from the same per-subscription pool. The documentation recommends adding one extra worker to the sizing, and the author's test confirms it: with the ceiling at 1, the sequence worker competed for the single slot between two table copies and finished in ten milliseconds. That's not a reason to raise the default value, just a detail to consider when planning the exact number of free slots.

Practical recommendation for whoever administers the subscription

The author's guidance, worth reinforcing for whoever operates logical replication in production in Brazil, today under PostgreSQL 17 or 18: leave the parameter at 2 (the default) on a subscription that will keep replicating on a permanent basis. For a one-off migration, raise the value to the number of tables the two servers can copy at the same time without saturating cores and disks, since each copy is a COPY occupying a core on the publisher, holding an xmin, feeding a COPY FROM with index maintenance on a core of the subscriber. In the author's test, on a two-core machine, six workers finished the same six tables in 14.5 seconds versus 13.3 seconds with only two workers, a picture of CPU contention when parallelism exceeds the available hardware.

After the migration, the value needs to go back to normal. An ALTER SUBSCRIPTION ... REFRESH PUBLICATION run by anyone on a Tuesday afternoon will trigger exactly that number of simultaneous copies against production, without asking first. Reserving the necessary WAL senders and replication slots on the publisher ahead of time, and reviewing disable_on_error on the subscription, are the two precautions that separate a controlled migration from a stalled-VACUUM incident discovered too late.

Source

Planet PostgreSQL, "All Your GUCs in a Row: max_sync_workers_per_subscription (The Build)" (https://postgr.es/p/9wl).

Translated from the Brazilian Portuguese original · Read the original