Dev & EngARTICLE

How Postgres decides how many workers to use in a parallel scan

A technical article by Christophe Pettus details min_parallel_table_scan_size and min_parallel_index_scan_size: the two PostgreSQL parameters that define the floor for a parallel scan, and why zeroing them out carelessly can hurt performance.

PostgreSQL has two parameters with similar names and functions that are easy to confuse: min_parallel_table_scan_size and min_parallel_index_scan_size. A recent article by Christophe Pettus, a consultant at PGX Inc., published in the "All Your GUCs in a Row" series on Planet PostgreSQL, takes both apart with practical tests. The central conclusion matters to anyone who has ever thought about "turning on parallelism" for a slow query without understanding what that actually changes in the execution plan.

The names say "minimum," and both do define a floor: a scan that the planner expects to read less than min_parallel_table_scan_size of a table (default 8MB) or less than min_parallel_index_scan_size of an index (default 512kB) doesn't even make it onto the list of possible parallel plans. That part is documented. The part the article lays bare is that the same value is also the first step of the ladder the planner climbs to decide how many workers to request, so changing the minimum changes every step above it.

The worker ladder

The calculation is simple to describe and easy to get wrong in your head: the planner starts at min_parallel_table_scan_size, triples that value successively until it exceeds the table's estimated size, and counts how many steps it climbed. Each step is one more worker, up to the cap of max_parallel_workers_per_gather. Pettus tested this on a 116MB table in version 18.6, with the per-gather limit set to 16 and parallelism costs zeroed out to isolate just the ladder's effect:

Table shows how many workers the planner grants for a 116MB table as min_parallel_table_scan_size rises from 1MB to 128MB, including the case with the value zeroed out
Table shows how many workers the planner grants for a 116MB table as min_parallel_table_scan_size rises from 1MB to 128MB, including the case with the value zeroed out. Reprodução: postgr.es.
min_parallel_table_scan_sizePlanned workers
1MB5
8MB (default)3
64MB1
128MBnone (serial plan)
09

The last row is the article's warning. Zeroing out the parameter doesn't mean "no minimum, and everything else stays the same": the code has a floor of one block, so the ladder starts running from 8kB, and a 116MB table ends up eight triplings above that, granting nine workers where the default would grant three. PostgreSQL's own regression tests zero out this parameter (along with the parallelism costs) just to get parallel plans on tiny test tables. In production, with max_parallel_workers_per_gather also raised, this hands out seven workers for a 10MB table.

The author measured the effect: on a filtered count(*) over a 9.6MB table, the serial execution took about 10ms; with seven workers granted by the zeroed-out minimum and a limit of eight per gather, the time rose to 33 to 37ms. The overhead of coordinating workers far outweighed the gain from splitting the work. The source code comment about this heuristic dates to 2015, for version 9.6, and remains unchanged through 19 beta 4:

we need something here for now

comment in the PostgreSQL planner's source code, quoted by Christophe Pettus

Eligible is not the same as chosen

Clearing the minimum only makes a parallel plan eligible. Whether it's actually chosen is a cost comparison, mainly against parallel_setup_cost (default 1000), and that comparison looks at rows and work per row, while the minimum looks at pages. Pettus built tables of increasing size and ran count(*) with a cheap filter on each one.

With narrow rows (81 per page), the plan stayed serial until around 204,000 rows, or 20MB: the minimum had already been satisfied since 8MB, but it made no difference, because the cost model still said no. With two-integer rows (226 per page), the cost started accepting around 200,000 rows too, except that equals 7MB, so on that table the plan turned parallel at exactly the default's 8MB, and that's where the minimum did the deciding.

The default, therefore, sits close to the point where the cost model already agrees for narrow rows and cheap filters. For ordinary rows, the cost usually agrees well before the minimum does. Lowering the minimum on small tables mostly works through a different path: lower steps on the ladder offer more workers, and more workers make the parallel plan look cheaper on paper, even with no real change in the work per row.

What escapes the calculation

Two situations ignore this calculation entirely, and both are worth knowing:

  • The per-table storage parameter parallel_workers overrides the minimum and ladder calculation.
  • Children of an append (partitions, inherited tables, arms of a UNION ALL) get a parallel path regardless of individual size, under the logic that many small pieces add up.

The author tested sixteen 2MB hash partitions: they produced a Parallel Append even with the minimum set to 1GB, while the same rows in a single table stayed serial. A small table can still show up inside a parallel plan without the scan itself being split across workers: a 48kB lookup table used in a join with a large table is read in full by each worker, which is exactly the desired behavior in that case.

The index side

min_parallel_index_scan_size compares against the estimated index pages the scan will read, not the index's total size. On a table with five million rows and a 107MB primary key, a range of 20,000 ids (estimated at 54 index pages) stayed serial, and a range of 30,000 (estimated at 80 pages) already picked up a worker, against the default of 64 pages.

A regular index scan has to clear both minimums, table and index, and uses the smaller worker count of the two calculations. The table check uses a pessimistic estimate of heap pages, so it's rarely the one that blocks the plan: the range of 30,000 ids read less than 3MB of heap and still passed the 8MB table minimum.

An index-only scan is checked only against the index minimum, because it can access almost none of the heap; in the tests, a range covering four million rows planned four workers as a regular index scan and five as an index-only scan. A parallel bitmap heap scan flips the logic: only the table minimum counts, against estimated heap pages, and the index parameter is ignored, because the underlying bitmap index scan isn't the parallelized part. B-tree is the only access method with a parallel index scan, so, for query purposes, this parameter is specifically a B-tree parameter.

Beyond the query planner

CREATE INDEX reads min_parallel_table_scan_size from the session and climbs the same ladder to decide how many workers to request for the index build. That means raising the value in postgresql.conf to keep small queries serial also makes index builds serial on any table below the new value, including indexes rebuilt by a pg_restore. With the minimum set to 1GB, B-tree and BRIN builds on a 482MB table stopped requesting a parallel worker.

VACUUM, in turn, uses min_parallel_index_scan_size to decide which indexes are large enough to get a parallel worker, but here the comparison is against the actual on-disk size, with no estimate involved. Three 456kB indexes didn't get workers; after SET min_parallel_index_scan_size = 0, the following VACUUM (PARALLEL 2) launched both workers.

PostgreSQL 19, still in beta at the time of the article, adds two new readers to these rules: the parallel TID range scan, checked against the table's total size regardless of the requested range, and parallel autovacuum, which now applies the same index cutoff as manual VACUUM.

What to do in practice

In short, the article's recommendation is to leave both parameters at their defaults. If parallel plans are showing up on queries too small to make up for the overhead, the right adjustment is parallel_setup_cost, not the minimum. If large tables are getting too few workers, the adjustment is max_parallel_workers_per_gather.

The one job only this parameter solves is stretching the ladder: anyone who raised the per-gather limit to 8 with large tables in mind, and now sees every 1GB table requesting five workers, can use min_parallel_table_scan_size = '64MB' to bring that down to three without touching the cap.

The price is that every scan and every index build under 64MB also turns serial, so the change should go on the role that received the higher limit, not into postgresql.conf globally. And if you find either of the two parameters zeroed out outside of a test suite, it's a sign someone copied the configuration from there without understanding the side effect.

Translated from the Brazilian Portuguese original · Read the original