Percona ClusterSync for MongoDB 1.0 gets automatic active-standby failover
PCSM 1.0.0 eliminates the single point of failure that stalled multi-day migrations between MongoDB clusters: now multiple instances compete for a lease and automatically take over replication when one goes down.
The bottleneck that forced manual recovery
Percona ClusterSync for MongoDB (PCSM) is not a replacement for MongoDB's native replica set: it's a migration tool that clones an entire cluster and then keeps both sides synchronized via change streams while the team decides on the final cutover. According to Percona's post from September 30, 2026, written by Adnan Supic, this continuous synchronization process usually takes days, not minutes.
Throughout that time, up to version 0.9.0, PCSM was a single process. If the host, container, or pod running the synchronization went down, replication stopped right there, with no automatic alert. Someone had to notice the interruption and manually recover the process before synchronization could resume. For anyone in the middle of a vendor lock-in migration with a defined cutover window, that's pure operational risk.
PCSM version 1.0.0 solves exactly that gap: not the high availability of the target database, which MongoDB already handles with its own replica set, but the high availability of the synchronization agent. It's a distinction worth noting before moving on, because confusing the two leads to the wrong expectations about what the tool covers.
Lease-based election, not naive heartbeat
The solution avoids any external coordinator (no Zookeeper, etcd, or Consul): the target MongoDB cluster itself holds the coordination state, in the percona_clustersync_mongodb database. Multiple PCSM instances can point to the same source and target; only one of them holds the lease and becomes ACTIVE, running the replication. The others remain STANDBY, monitoring.
The lease document is small and straightforward:
{
"_id": "lease",
"term": 7,
"instanceId": "b3f1c2a4-9d7e-4c11-8a2f-1e6b0d5c9a77",
"electionDate": { "$date": "2026-07-17T09:14:02.190Z" },
"expiresAt": { "$date": "2026-07-17T09:20:41.882Z" }
}The election happens as a single atomic conditional write: an instance only takes over the lease if the current one has already expired, and expiration is evaluated by the MongoDB server's clock, not by the clock of each PCSM host. That's the detail that matters for anyone who distrusts distributed solutions by default: clock skew between the machines running PCSM becomes irrelevant, because only the database's clock counts. Two standbys competing can't both win at the same time.
Term-based fencing: the problem of the zombie that keeps writing
The classic distributed-systems scenario is this: the ACTIVE instance freezes (a GC pause, a network partition, a pod restarted before the lease expires), the lease expires, a standby is promoted, and then the old instance wakes up and keeps writing as if nothing had happened. Without protection, this corrupts the new active's state.
PCSM avoids this risk with a classic fencing token: the term field is monotonic, increments on every switch from ACTIVE to STANDBY, and is stamped on every checkpoint write made by the active instance, in percona_clustersync_mongodb.checkpoints. When the zombie instance tries to write with an outdated term, the target rejects the write, the deposed instance notices the rejection and self-demotes to STANDBY. The new active's state is never corrupted by a late write.
Checkpoint-based recovery: what's covered and what isn't
Upon promotion, the new active instance resumes replication from the last persisted checkpoint, the same failure-recovery mechanism PCSM already used, now triggered automatically. The timings are fixed in this version: the lease has a TTL of 10 seconds, the active renews it every 3 seconds, and every instance emits a heartbeat every 3 seconds as well. In practice, a forcibly killed active is replaced in about the time of the lease's TTL.
There's an explicit limit every operator needs to know before blindly trusting the HA: automatic coverage applies to the continuous replication phase via change streams. The initial clone, as of now, is not resumable. If the active dies in the middle of the clone, a standby is still promoted, but it detects the interrupted clone during recovery and fails explicitly, with the reason exposed in the logs and in the /status endpoint:
initial clone interrupted by failover and is not resumable; start a new run to re-clone from scratch
The recovery, in that case, is a simple /start on the new active, which restarts synchronization from scratch. Automatic failover, in fact, only kicks in after the clone finishes and continuous replication begins.
Operating the cluster: API with envelope and three metrics
Each instance maintains a liveness document in percona_clustersync_mongodb.members, and API responses now include a cluster envelope: who is me, what the role is, and the full list of live members with host, port, and role. Sending an operational command (/start, /pause, /finalize) to a standby returns HTTP 409 with error: "not_active", already carrying the same envelope in the body, so the client can locate the active instance and retry without guesswork or outdated DNS tricks.
Two operational details are worth noting:
- The envelope only appears when more than one live member is observed; a lone instance responds identically to the previous API, so existing tools keep working without changes.
/metricsandpprofare served regardless of role, whether active or standby.
For state cleanup, two direct commands act straight on the target, without needing a running server: pcsm reset members and pcsm reset lease.
Three Prometheus metrics tell the full story of a failover:
| Metric | Type | What it shows |
|---|---|---|
percona_clustersync_mongodb_ha_active | gauge | 1 if the instance is active, 0 if it's standby |
percona_clustersync_mongodb_ha_term | gauge | current lease term (fencing) |
percona_clustersync_mongodb_ha_role_transitions_total | counter | role switches on the instance |
Percona distributes a ready-made Grafana dashboard in the project's repository, using these three metrics: the ha_active timeline per instance works as a deployment map and shows who's active and since when, ha_term shows the current fencing term, and the rate of ha_role_transitions_total works as a flapping detector (zero is stable; a rising number signals instability).

Migrating from the previous version
A point of attention for anyone already running PCSM in production: the replication state from version 0.9.0 is not compatible with 1.0.0. You need to run pcsm reset against the target before starting replication with the new version. After that, spinning up a second instance pointing to the same source and target is enough:
pcsm --source "mongodb://src-mongos:27017" \
--target "mongodb://tgt-mongos:27017"Running the same command on a second host, /status on either of the two shows one ACTIVE and one STANDBY. Killing the active makes the standby take over replication from the last checkpoint within seconds.
Where this matters and where it doesn't solve everything on its own
In short: PCSM 1.0.0's HA is worth it for anyone in the middle of a long migration between MongoDB clusters, whether to switch providers or to escape cloud lock-in, who doesn't want to rely on an operator watching the process 24 hours a day. For that case, the gain is real: zero configuration change to enable it, and automatic failover in the phase that matters most, continuous replication.
What the feature doesn't replace is operational discipline around the migration. The initial clone remains a point requiring manual attention, the final cutover (/finalize) still requires consistency validation from whoever is operating it, and nothing here removes the need to test the cutover plan before running it in production with real data. HA on the synchronization agent reduces a specific risk; it doesn't eliminate the need to understand the migration's execution plan as a whole before trusting it.
Translated from the Brazilian Portuguese original · Read the original
PostgreSQL 19 Will Allow Forcing the Execution Plan with pg_plan_advice
Two extensions announced for PostgreSQL 19, pg_plan_advice and pg_stash_advice, let you lock join, scan, and parallelism behavior when the optimizer gets it wrong. The community resisted hints for years: understand why it gave in and when it's worth using.