Dev & EngARTICLE

PostgreSQL treats kill -9 on one connection as if it were a full crash

A post published on October 9, 2026 by Shridhar Khanal explains why a signal sent to a single PostgreSQL child process brings down every connection and triggers WAL recovery, even though no one restarts the service.

The symptom: a crash recovery with no apparent cause

In an article published on October 9, 2026 on Planet PostgreSQL, DBA Shridhar Khanal, from Stormatics, describes a scenario familiar to anyone who runs Postgres in production: a log line like server process (PID 48213) was terminated by signal 1: Hangup, followed by automatic recovery. Every connection drops, the application throws errors for half a minute, and then everything comes back as if nothing had happened.

The strange detail, according to Khanal, is that signal 1 is exactly SIGHUP, the same one PostgreSQL uses to reload configuration without bringing anything down. Understanding why it shows up in a crash log requires understanding what operating system signals are and how PostgreSQL reacts to each one.

What a signal actually is

A signal is a notification sent by another process, or by the operating system itself, to a running process. It carries no data at all: it just announces that something happened, such as "stop", "cancel whatever you're doing", or "the terminal hung up". Every signal has a name and a number, and the two are interchangeable: commands use the name, while the PostgreSQL log shows the number followed by a short description.

The signals most relevant to anyone operating PostgreSQL, according to the article, are these:

NumberNameOriginal meaningPostgreSQL's response
1SIGHUPterminal disconnectedreloads the configuration
2SIGINTinterrupt (Ctrl+C)cancels the current query, via pg_cancel_backend()
3SIGQUITimmediate exit (Ctrl+\)ends the session without cleanup
9SIGKILLkillnobody handles or ignores it; the process simply dies
13SIGPIPEbroken pipeignored; PostgreSQL handles the lost connection on its own
15SIGTERMtermination requestends the session cleanly, via pg_terminate_backend()

A signal does nothing on its own: what happens depends on the program that receives it. A process can run its own code for that signal (a signal handler), ignore it, or do nothing special, in which case Linux applies the default action. For SIGHUP, the default action is to terminate the program.

Why PostgreSQL survives SIGHUP and most programs don't

Anyone who has run a long pg_dump over SSH and watched the connection drop along with the Wi-Fi knows the practical effect of this default action: when the session closes, the system sends SIGHUP to the processes of that terminal, and since pg_dump has no handler for that signal, Linux simply terminates it. That's why DBAs usually run long jobs like this:

nohup pg_dump mydb > mydb.sql &

nohup stands for "no hang-up": it tells the job to ignore SIGHUP, so that it keeps running even after the session closes. PostgreSQL, on the other hand, has its own handler for signal 1 built into its source code: when it arrives, the process rereads the configuration and carries on instead of dying. That's the behavior you see every time someone runs SELECT pg_reload_conf();: the signal reaches the main process, which passes it on to the others, and none of them dies.

The architecture behind the scare

PostgreSQL doesn't run as a single process. It runs as a family of processes, with the postmaster as the parent: it starts first, accepts connections, spawns the other processes, and keeps watch over them. Each connection gets its own backend, and background processes handle tasks such as checkpointing and autovacuum.

All of these processes share an area of memory, the shared memory, where cached data, locks, and other shared state live. In short: this shared memory is part of what makes PostgreSQL fast, but it's also the reason why a single process failing can affect every connection at once.

Two log lines that look alike but aren't

PostgreSQL writes two very different lines involving SIGHUP, and mixing them up is what creates the mystery described by Khanal:

  • received SIGHUP, reloading configuration files: is a receipt. The postmaster received signal 1 and is reloading. Everything is fine.
  • server process (PID 48213) was terminated by signal 1: Hangup: is a death certificate. A child of the postmaster died, and the cause was signal 1.

The label "server process" doesn't mean the process belonged to PostgreSQL: it just means it was a child of the postmaster. In newer versions, that death certificate may show a more specific label, such as client backend, instead of server process.

Why the death of a child process resets everything

When a process dies, Linux cleans up its memory, closes its files, and notifies the parent. What Linux can't clean up is PostgreSQL's shared memory, because it doesn't know what's inside it. A process that dies suddenly might have been holding a lock, or in the middle of writing a data page, and the postmaster has no way to check, so it assumes the worst case.

That's why, when the postmaster discovers that a child was killed by a signal, it: terminates every other process; resets the shared memory; and replays the write-ahead log (WAL) to bring the database back to a consistent state. The postmaster process itself never stops, so its PID stays the same and the server never actually appears to have restarted. This automatic behavior is controlled by the restart_after_crash parameter, which is on by default.

The central point of Khanal's article is that the postmaster reacts to how a child died, not to what that child was. If, for whatever reason, a process that isn't part of PostgreSQL ends up as a child of the postmaster and is killed by a signal with no handler, the same rule applies: the postmaster sees a child killed by a signal, assumes the shared memory may be compromised, and resets everything, even if that process never touched shared memory at all.

The everyday version of the problem: kill -9

It doesn't take a strange process to trigger this reset: the most common cause, according to the article, is a well-intentioned kill -9. Someone finds the PID of a slow query in top and kills the process. The query stops, but every other connection to the server goes down along with it, and the log shows terminated by signal 9: Killed, followed by the same crash recovery. Linux's OOM killer uses that same signal 9, so running out of RAM triggers exactly the same reset.

The safe path is to let PostgreSQL itself handle the stop. First, identify the process:

sql
SELECT pid, usename, state, now() - query_start AS running_for, query
FROM pg_stat_activity
WHERE state = 'active'
ORDER BY running_for DESC;

Then, cancel the query before ending the session:

sql
-- Step 1: cancel the query, keep the connection
SELECT pg_cancel_backend(48213);

-- Step 2: if that's not enough, end the session
SELECT pg_terminate_backend(48213);

Both functions go through PostgreSQL's internal signal handling, so the process cleans up after itself and only that session is affected, with no full reset. To use them, you need to be a superuser, be connected with the same role as the target session, or belong to the pg_signal_backend role.

What remains open

Khanal's article is the first in a series and deliberately leaves two questions unanswered: how a process PostgreSQL never started ends up as a child of the postmaster, and who actually sent signal 1 in the case described. The second part, promised by the author, promises a reproducible test and recommendations for avoiding this in production.

For anyone operating PostgreSQL in a real environment, the immediate lesson is already worth the read: before reaching for kill -9 under pressure, it's worth stopping to check pg_stat_activity. A well-applied pg_cancel_backend() costs one extra command and avoids half a minute of cascading errors for every connected user, not just the session that was causing trouble.

Translated from the Brazilian Portuguese original · Read the original

Read also
Dev & Eng

PostgreSQL's numeric is slow because of design decisions from 1998

A study published on October 8 on Planet PostgreSQL shows that aggregations with numeric can take almost three times longer and consume nine times more memory than the equivalent in double precision. The reason lies in four architecture decisions made more than 25 years ago.

Roberto Diniz··1 min