Skip to main content
Metrics type: Key MetricsCategory: Executive Overview

At a glance

The percentage of provisioned data volume currently consumed by the PostgreSQL instance: table data, indexes, WAL, temporary files, and the catalogue. For a platform team this is the single most unforgiving capacity number on the board. PostgreSQL does not gracefully degrade when the data disk fills: once the volume hits 100%, the database refuses new writes, autovacuum cannot reclaim space, and on many managed services the instance is forced into a read-only or recovery state. This card is the early-warning gauge that keeps you ahead of that wall.

Calculation

The gauge is used_bytes / total_provisioned_bytes * 100, sampled on the real-time cycle. On a self-managed instance the engine derives used_bytes two ways and reconciles them:
  1. Filesystem-level: statvfs on the PGDATA mount point gives total and available; used = total - available. This is authoritative because it captures everything on the volume, including temp files and any non-PostgreSQL data sharing the mount.
  2. PostgreSQL-level: the sum of pg_database_size(datname) across all databases, plus the size of pg_wal, plus temporary file usage from pg_stat_database.temp_bytes. This is what PostgreSQL itself believes it is using.
The headline gauge uses the filesystem figure because that is what actually fills. The PostgreSQL-level figure is retained so the drill-down can attribute growth to heap, indexes, WAL, or temp files. On managed services there is no filesystem access, so the engine uses the provider’s own storage metric:
  • RDS / Aurora: 100 - (FreeStorageSpace / AllocatedStorage * 100). Aurora auto-scales storage, so the gauge there reflects used against the current allocated ceiling, which is itself elastic.
  • Cloud SQL: database/disk/bytes_used / database/disk/quota * 100.
WAL deserves special attention. A stuck replication slot, a failing archive_command, or a long-running base backup can cause pg_wal to grow without bound while the rest of the database is quiet. Because WAL lives on the data volume by default, a WAL blow-up shows up here first and can be the difference between 70% and 100% within the hour.

Worked example

A platform team runs a self-managed PostgreSQL 15 primary on a 500 GB gp3 volume backing an order-management service. Snapshot taken on 14 Apr 26 at 02:10 BST during the overnight batch window. The gauge reads 95.6% in red and the alert has already paged the on-call DBA, because 90% was crossed at 01:54. Reading the drill-down, two things stand out:
  1. pg_wal at 64 GB is roughly 4x its steady-state size of around 16 GB. A quick check of pg_replication_slots shows a slot named analytics_cdc with a restart_lsn far behind the current WAL position: the downstream change-data-capture consumer has been down since 22:30, so PostgreSQL is retaining every WAL segment since then.
  2. Temp files at 19 GB come from an overnight reporting query spilling a large sort to disk because work_mem is too small for it.
After dropping the orphaned slot and the next checkpoint, pg_wal returns to 17 GB and the gauge falls to 84.3%, back under the alert threshold. The platform team then files a follow-up to add a separate monitor on replication-slot lag so an idle consumer never silently fills the data disk again. Three lessons platform teams should carry from this:
  1. Disk usage is not just table growth. The scary, fast movements come from WAL retention and temp-file spill, not from the heap creeping up. When this gauge jumps suddenly, check pg_wal and temp files before assuming you simply need more storage.
  2. 90% is a real deadline, not a vanity threshold. A write-heavy primary can close the last 10% in well under an hour. The aggressive alert exists so you have time to act, not to nag.
  3. Reclaim before you resize. Expanding a volume is the right medium-term fix, but the immediate move is almost always to release retained WAL (orphaned slots, broken archiving) or kill a runaway temp spill. Those reclaim space in minutes; a resize plus rebalance can take longer.

Sibling cards

Reconciling against the source

Where to look in PostgreSQL and the host:
Filesystem truth (self-managed): df -h $PGDATA on the host shows the volume the gauge tracks. This is the number that matters when the disk is about to fill. PostgreSQL’s own view: SELECT pg_size_pretty(sum(pg_database_size(datname))) FROM pg_database; for total logical size, and SELECT pg_size_pretty(sum(size)) FROM pg_ls_waldir(); for the WAL directory. Per-table attribution: SELECT relname, pg_size_pretty(pg_total_relation_size(relid)) FROM pg_stat_user_tables ORDER BY pg_total_relation_size(relid) DESC LIMIT 20; Managed services: the RDS / Aurora console CloudWatch tab shows FreeStorageSpace; the Cloud SQL console shows storage usage under the instance overview.
Why our number may legitimately differ from the native tooling: Cross-source reconciliation:

Known limitations / FAQs

The gauge says 95% but sum(pg_database_size()) only accounts for 70% of the volume. Where is the rest? Almost always WAL and temp files, neither of which is counted in pg_database_size(). Run SELECT pg_size_pretty(sum(size)) FROM pg_ls_waldir(); to size pg_wal, and check pg_stat_database.temp_bytes for active spill. Filesystem reserved blocks (around 5% on ext4) and any non-PostgreSQL files on the mount make up the remainder. The gauge tracks the filesystem because that is what actually fills. Why is the alert at 90% and not 95%? It feels early. Because the last 10% can vanish faster than you can respond. A runaway replication slot or broken archive_command retains WAL on the data volume, and on a busy primary that can add several GB per hour with no warning. The 90% page buys you the time to reclaim space before the database stops accepting writes at 100%. What actually happens when PostgreSQL hits 100% disk? The instance can no longer write WAL, so it refuses new write transactions and may panic-shutdown to protect data integrity. Autovacuum cannot run (it needs to write), so you cannot reclaim space the easy way. On RDS the instance enters a storage-full state and may become unavailable. Recovery usually means expanding the volume or, on self-managed hosts, manually freeing space (dropping orphaned slots, clearing temp files) before PostgreSQL will restart cleanly. Avoiding 100% is far cheaper than recovering from it. On Aurora the gauge looks low even though my workload is huge. Is it broken? No. Aurora separates compute from a distributed, auto-scaling storage layer that grows in 10 GB increments up to a large ceiling. The gauge shows used against the current allocated amount, which keeps expanding, so it rarely approaches 100% the way a fixed gp3 volume does. On Aurora, watch the absolute storage cost and growth rate rather than the percentage. Does dropping a large table immediately free disk? DROP TABLE and TRUNCATE release space back to the filesystem promptly. DELETE does not: it only marks rows dead, and the space is reused by future inserts only after autovacuum processes the table. To return space to the operating system after large deletes you need VACUUM FULL (which takes an exclusive lock and rewrites the table) or an online repack tool such as pg_repack. This is why a table can stay large on disk long after its logical row count drops. Can I move WAL off the data volume to protect against this? Yes. Mounting pg_wal on its own volume (via a symlink or the --waldir option at initdb) isolates WAL growth from data growth, so a stuck slot fills the WAL volume rather than stopping all writes on the data volume. Many production deployments do this. The headline gauge then tracks the data volume; the WAL volume is surfaced separately. It is a sound mitigation but does not remove the need to monitor replication slots. The number jumps up and down by a few percent within minutes. Why so noisy? Temporary files. Large sorts, hash joins, and index builds spill to disk and release when they finish, so a snapshot taken during a heavy query reads higher than one taken a moment later. If the noise is large, raise work_mem for those workloads so they sort in memory, or schedule heavy reporting off the primary. Persistent growth (not transient spikes) is the signal to act on.

Tracked live in Vortex IQ Nerve Centre

Database Disk Usage % is one of hundreds of KPI pulses Vortex IQ tracks across PostgreSQL and 70+ other ecommerce connectors. Nerve Centre runs the detection layer; Vortex Mind investigates the cause when something moves; Ask Viq lets you interrogate any number in plain English. Start for free or book a demo to see this metric running on your own data.