Dev & EngARTICLE

Azure disk over 4 TiB disables cache and misleads Postgres tuners

Azure disk over 4 TiB disables cache and misleads Postgres tuners

Stormatics documented a case in which a week of tuning shared_buffers, work_mem, and effective_cache_size did not fix slow reads: the data disk had grown past 4 TiB, and Azure disables host caching at that threshold without warning.

Stormatics documented a case in which a week of tuning shared_buffers, work_mem, and effective_cache_size did not fix slow reads: the data disk had grown past 4 TiB, and Azure disables host caching at that threshold without warning.

In a post published on September 25, 2026, Umair Shahid, of Stormatics, describes a diagnosis that should sound familiar to any team that has fought with Postgres in the cloud: a full week reviewing shared_buffers, effective_cache_size, work_mem, and the rest of the classic tuning checklist, everything set sensibly, and reads stayed slow. The cause wasn't in any database parameter. The client's data disk had grown to 4 TiB, and at the exact moment it crossed that mark, Azure turned off the host cache that sat in front of it.

The case illustrates the angle that matters to anyone building on Postgres, not just to whoever administers the server: the database blindly trusts the storage layer beneath it. When that layer is throttled, limited, or lacking cache, Postgres has no way to warn about it. It just waits. And waiting on a slow disk, from inside the database, looks exactly like a poorly configured database, so that's where diagnosis time tends to be spent, even when the real problem sits one layer down.

The 4 TiB trap

On Azure, host caching is a real performance feature: the VM keeps a cache built from its own memory and a local SSD, and serves a good share of reads from it without those reads having to travel to the remote disk. For a read-heavy workload, that cache does quiet, important work. The problem is that Azure only supports host caching on managed disks smaller than 4 TiB (up to 4,095 GiB). Once a disk reaches 4 TiB or more, the cache disappears and every read goes straight to remote storage.

The treacherous detail, according to Shahid, is that the Azure portal keeps showing a cache option available to select even on a disk that size, except it has no effect at all. You can look at the configuration, see "ReadOnly cache" checked, and reasonably assume the cache is active, while every read is actually going straight to the remote disk.

The fix Stormatics recommended for that client was to swap a single giant disk for several smaller disks, each under 4 TiB, joined into a logical volume via LVM or software RAID. Each individual disk keeps its host cache, the combined volume delivers the needed capacity, and reads go back to being served from cache. Same data, same Postgres, different layout.

The VM has its own ceiling too

The second trap explains why a bigger, faster disk sometimes delivers no gain at all. The disk has a performance limit, but the VM has its own too, and the two are independent. The virtual machine caps how much IOPS and throughput it lets through, regardless of what the disk underneath is capable of delivering. Attaching a top-tier disk to a small VM makes the VM's ceiling the real ceiling: the disk sits partly idle while the VM simply won't let more traffic through.

Because of that, when investigating an I/O bottleneck, you need to check both numbers: what the disk is rated for and what the VM is rated for. If measured throughput is up against the VM's uncached limit, a faster disk fixes nothing. The fix requires a larger VM size or a family built for heavy storage workloads, and Shahid points to this as one of the most common places where money gets spent in the wrong direction: upgrading the disk when the VM was the wall the whole time.

Capacity and performance tied together

In the classic Premium SSD tiers, performance scales together with capacity: a small disk gets a small IOPS and throughput ceiling, a large disk gets a large one, in fixed bands. The trap here is subtle: sometimes a disk gets provisioned for the space it needs and the result is a lack of IOPS, because the chosen size falls into a modest performance band. In other cases, capacity that will never be used gets provisioned just to buy the IOPS that comes bundled with a larger band.

Premium SSD v2 loosens that constraint by letting you configure capacity, IOPS, and throughput independently, buying exactly the performance needed without inflating the disk size to get it. Shahid is clear that this flexibility comes with its own trade-offs around cache and configuration, so it isn't a free upgrade: it's a deliberate choice.

WAL and data call for different disks

Postgres has two quite distinct I/O personalities living in the same system. Data files are heavy on random reads. The write-ahead log is heavy on sequential writes. Treating both the same way at the storage layer leaves performance on the table: a read cache is a gift to data files but does nothing useful for the WAL, which is almost entirely sequential writes.

A clean layout separates data files, on a disk with read caching, from the WAL, on its own disk optimized for sustained writes, with no read cache in the path. This also keeps a WAL activity spike during a heavy-write window from competing for the same disk with read traffic.

How to confirm before touching a parameter

Before touching any Postgres configuration, Shahid recommends confirming where the time is actually going, which takes a few minutes and avoids tuning the wrong layer. The first step is to look inside Postgres itself, at pg_stat_activity, watching the wait events: sessions stacked up on I/O waits indicate the database is waiting on storage, not on locks or CPU. The second step is to leave Postgres and go to the operating system, with a tool like iostat, to check disk utilization, average wait time per request, and the actual volume of reads and writes. Utilization pinned at the top with rising wait times points to a saturated disk.

After that, what was measured gets compared against the disk's and the VM's rated limits: is throughput close to the disk's ceiling or the VM's? Is host caching actually active, or did the disk cross the 4 TiB mark without anyone noticing? Do WAL and data share the same disk? According to the author, these questions almost always point to which of the four traps is in play.

What changes for those building on Postgres

For those writing applications and standing up infrastructure on Azure, the practical point is this: a poorly calibrated index or a bad execution plan are still the first things to rule out when a query is slow, and nothing in the post replaces that work. But once you've confirmed the plan is correct and the indexes exist, the next place to look is no longer a postgresql.conf parameter, it's the disk layout underneath it. That changes the order of investigation for anyone running databases on their own VM in Azure, with managed disks and VM size configurable directly. The post deals specifically with that self-managed infrastructure scenario, and does not describe how fully managed services handle the disk layer.

The point Shahid makes a point of stating clearly is that these fixes cost no license, no code rewrite, and no new product: they're decisions about layout, VM size, and cache configuration. The expensive alternative is never finding the real bottleneck and continuing to scale up the instance, paying more every month for a problem that a few smaller disks, arranged differently, would solve. The post does not cover managed databases outside Azure, nor does it set equivalent numbers for AWS or GCP: every cloud imposes similar disk and instance limits, but the exact thresholds vary and aren't covered in the source.

Translated from the Brazilian Portuguese original · Read the original