Skip to main content
Metrics type: Key MetricsCategory: Nerve Centre

At a glance

A real-time alert that fires when the live connection count climbs above 90% of max_connections and stays there for at least one minute. This is the early-warning band before total exhaustion. At 90% you still have headroom, but you are close enough that one traffic burst, one slow query holding threads, or one runaway batch job can push you to the wall, at which point MariaDB returns ERROR 1040: Too many connections and refuses new clients outright. For a DBA this is the “act now, do not wait” signal: you have minutes, not hours, to shed load or add capacity before connections start being rejected.

Calculation

The card computes a live saturation ratio on every sample:
Threads_connected is the number of currently open connections (both running queries and idle-but-open sessions), read from SHOW GLOBAL STATUS LIKE 'Threads_connected'. max_connections is the hard ceiling configured for the server. The “sustained 1m” requirement means a single sample above 90% does not raise an alert; the ratio must remain above 90% across consecutive samples spanning at least a minute. This deliberately ignores brief, self-correcting spikes (a deploy, a cache flush, a momentary thundering herd) and only escalates genuine pressure. Note that MariaDB silently reserves one extra connection above max_connections for a user with the SUPER/CONNECTION ADMIN privilege, so an administrator can always log in to remediate even when ordinary clients are being refused. The card measures against the ordinary ceiling, which is what application traffic competes for.

Worked example

A retail platform runs MariaDB 10.11 with max_connections = 500 behind an application tier that auto-scales during promotions. Snapshot taken on 22 Apr 26 during a flash-sale window. At 20:01 the alert fired: saturation had held above 90% for a full minute. The on-call DBA had roughly three minutes of headroom before 100%. Two parallel actions:
The processlist showed 140 connections in Sleep state from the reporting application (idle pool connections holding slots) and a cluster of long-running Sending data threads from an unindexed analytics query launched at 19:55. The DBA killed the analytics query (KILL <id>), which freed threads as the pool recycled, and the ratio fell back below 90% by 20:05. Three takeaways:
  1. Idle connections count too. Threads_connected includes sessions in Sleep state. A leaky or oversized application pool can hold MariaDB near the ceiling even when almost nothing is actually executing. Pair with Connections In Use to separate active from idle.
  2. Slow queries pull saturation up indirectly. A handful of long-running queries hold their threads, so new requests stack up behind them. Always check Query Latency p95 (ms) and the processlist together; the fix is often killing one query, not adding capacity.
  3. Raising max_connections is the last resort, not the first. Each connection costs memory (thread stack plus per-connection buffers). Bumping the ceiling without checking RAM can trade connection refusals for an OOM kill. Confirm Memory Usage % headroom before increasing the limit.

Sibling cards

Reconciling against the source

Where to look in MariaDB’s own tooling:
SHOW GLOBAL STATUS LIKE 'Threads_connected'; for the live connection count. SHOW VARIABLES LIKE 'max_connections'; for the configured ceiling (compute the ratio yourself). SHOW GLOBAL STATUS LIKE 'Max_used_connections'; for the high-water mark since startup, and Max_used_connections_time for when it occurred. SELECT * FROM information_schema.PROCESSLIST; (or SHOW FULL PROCESSLIST) for who is holding each connection right now.
Why our number may legitimately differ from a raw SHOW STATUS: On managed services: Amazon RDS / Aurora for MariaDB sets max_connections from a formula based on instance memory (DBInstanceClassMemory), so confirm the effective value before reasoning about the ratio; CloudWatch exposes DatabaseConnections for the live count. SkySQL and Azure Database for MariaDB surface connection metrics in their own consoles. Once you align on the same max_connections value, the ratio should match.

Known limitations / FAQs

Q: The alert fired but I could still connect fine. Was it a false alarm? No. 90% is the warning band, not the failure point. You hit refusals at 100%. The alert exists precisely so you act before clients are turned away. You could connect because you used a SUPER-privileged account, which gets the one reserved connection above the ceiling, or because saturation dipped between the alert and your test. Treat a fired alert as a near-miss to investigate, not a non-event. Q: Most of my connections are in Sleep state. Should I still worry? Yes, idle connections still occupy slots and count toward Threads_connected. An oversized or leaky application pool can pin the server near the ceiling with almost no real work happening. Tune the pool’s maxPoolSize and idle-timeout so it returns connections, and consider a connection proxy (MaxScale, ProxySQL) to multiplex. Pair with Connections In Use to see the active-versus-idle split. Q: Should I just raise max_connections to make the alert stop? Only after checking memory. Every connection consumes a thread stack plus per-connection buffers (sort_buffer_size, join_buffer_size, read_buffer_size, and so on), so raising the ceiling raises peak memory. If you bump it past what RAM allows, you trade Too many connections for an OOM kill, which is far worse. Check Memory Usage % first, and prefer fixing the source of the connections (pool sizing, slow queries) over raising the limit. Q: Why the one-minute sustained requirement? I want to know about every spike. Brief spikes above 90% are common and self-correcting: a deploy reconnecting pools, a cache stampede, a momentary burst. Alerting on every single sample would bury you in noise. The one-minute hold ensures the alert reflects genuine, persistent pressure. If you genuinely want a tighter trigger, the Connection Pool Saturation % gauge shows every sample without the sustain filter. Q: On a Galera cluster, does this measure one node or the whole cluster? This card measures the selected MariaDB instance. In a Galera cluster each node has its own max_connections and its own Threads_connected, and traffic is rarely balanced perfectly, so one node can saturate while others have headroom. For the cluster-wide picture use Pool Saturation Across Galera Nodes vs Traffic, which compares saturation across all nodes against incoming traffic. Q: What happens at exactly 100%? New non-SUPER connections are refused with ERROR 1040: Too many connections, and Connection_errors_max_connections increments. Existing connections keep working. The application sees connection failures and, if it does not degrade gracefully, user-facing errors. This is why the 90% alert matters: it gives you the lead time to avoid 100% entirely.

Tracked live in Vortex IQ Nerve Centre

Connection Pool at >90% Saturation is one of hundreds of KPI pulses Vortex IQ tracks across MariaDB and 70+ other ecommerce connectors. Nerve Centre runs the detection layer; Vortex Mind investigates the cause when something moves; Ask Viq lets you interrogate any number in plain English. Start for free or book a demo to see this metric running on your own data.