Skip to main content
Metrics type: Key MetricsCategory: Backup

At a glance

The age, in hours, of your most recent verified-successful backup of this MongoDB deployment. This is the single most important number for a DBA to know is healthy, because it sets the floor on how much data you can lose. A reading of “2h” means a catastrophic failure right now would cost you at most two hours of writes. A reading of “97h” means your backup pipeline has been silently broken for four days and a failure would be a four-day data loss event. The card turns red at >72h: at that point you are operating without a usable recovery point.

Calculation

The card resolves the timestamp of the most recent successful backup and subtracts it from the current time:
How last_successful_backup_completed_at_utc is resolved depends on how the deployment is backed up:
  • Atlas Cloud Backups: the engine queries the snapshot list for the cluster and takes the newest entry where status == "completed", using its createdAt (snapshot completion) timestamp. Continuous Cloud Backup also exposes an oplog window; when continuous backup is active the effective recovery point is near-real-time, so the card reflects the latest snapshot marker rather than the oplog tail.
  • Self-managed mongodump: the engine reads the success marker your backup job records (a sentinel object in object storage, the archive file’s modification time, or a status row your job writes). Only runs that exited 0 with a non-empty archive count.
  • Ops Manager / Cloud Manager: the engine reads the latest snapshot metadata for the deployment and uses the completion timestamp of the newest snapshot in a complete state.
  • Filesystem / volume snapshots: the completion time of the newest snapshot tagged for this deployment.
All timestamps are normalised to UTC before subtraction, then the result is rendered in the merchant’s display time zone for any chart axes. The headline is a single duration in hours.

Worked example

A platform team runs a 3-node MongoDB 6.0 replica set on Atlas (M30) backing an order-processing service. Daily Cloud Backup snapshots are scheduled for 02:00 UTC, with continuous backup enabled for a 24-hour oplog window. Snapshot taken on 14 Apr 26 at 09:15 UTC. The card shows 7h in green. The on-call DBA reads this as healthy: a total-loss event right now would cost at most the writes since 02:04, and because continuous backup is on, the real recoverable point is within minutes, not hours. The 7h figure simply reflects the last full snapshot marker. Now contrast a failure scenario two weeks later. The scheduled snapshot job started failing on 26 Apr 26 because the Atlas project’s backup storage quota was exhausted, but nobody was watching the Atlas alert. Snapshot taken on 29 Apr 26 at 10:00 UTC.
The card now reads 80h in red, having crossed the >72h threshold at roughly 02:00 on 29 Apr. This is exactly the signal the card exists to surface: three consecutive snapshot failures that the team would otherwise only discover when they tried to restore. The DBA’s response is, in order: (1) confirm the deployment itself is healthy and writes are still landing, (2) find the root cause of the failed snapshots (here, the quota), (3) clear the blocker and trigger an on-demand snapshot immediately rather than waiting for the next 02:00 window, (4) once the on-demand snapshot reaches completed, the card drops back to single digits. Three things worth remembering:
  1. A low number is necessary but not sufficient. “2h ago” only means a backup completed 2 hours ago. It does not prove the backup is restorable. Pair this card with periodic restore tests; a backup you have never restored is a hypothesis, not a recovery point.
  2. The threshold is a ceiling, not a target. >72h is the alert line, but if your business can only tolerate one hour of data loss, your real target RPO is one hour and you should be on continuous backup, not daily snapshots. Configure the alert threshold to match your actual RPO.
  3. Watch the trend, not just the value. A backup age that climbs smoothly from 2h to 26h over a day and then snaps back to 2h is a healthy daily cycle. A backup age that climbs past one cycle boundary without resetting is the early sign of a broken job, visible hours before it crosses the red line.

Sibling cards to read alongside

Reconciling against the source

Where to confirm the number in MongoDB’s own tooling:
Atlas: the Cloud Backups dashboard for the cluster lists every snapshot with its status and completion time; the newest completed row is the basis for this card. Atlas also exposes the continuous-backup oplog window here. Ops Manager / Cloud Manager: the Backup tab for the deployment shows the snapshot schedule and the latest snapshot’s completion time. Self-managed mongodump: check your backup job’s logs and the archive’s timestamp directly, for example the LastModified on the S3 object or the file mtime, and confirm the run exited 0.
Why our number may legitimately differ from the native view: Cross-connector reconciliation:

Known limitations / FAQs

My backup completed an hour ago but the card still shows the old age. Why? The card refreshes on a 60-second cycle and, for Atlas, depends on the snapshot reaching a completed status in the Cloud Backups API. A snapshot that is still finalising or replicating shows as in-progress and does not reset the age until it terminates successfully. Allow one refresh cycle after the native console shows completed. Does a low backup age guarantee I can restore? No. This card proves a backup finished, not that it restores cleanly. The only way to prove restorability is to actually restore, ideally on a schedule into an isolated environment. Treat a green reading as “a recovery point exists” and back it with periodic restore tests for “the recovery point works”. I have continuous backup enabled, so why does the card sometimes read several hours? Continuous (point-in-time) backup gives you a recoverable point within the oplog window, often minutes, but this card reports the last full snapshot marker, which still follows your snapshot schedule. The headline being a few hours old is normal and healthy when continuous backup is on; your effective RPO is much smaller than the number shown. Why is the alert at 72h rather than 24h? 72h is a deliberately conservative default so it does not cry wolf on weekly or every-other-day schedules. It is the line past which most teams have no usable recovery point. If your RPO is tighter, lower the alert threshold to one or two backup intervals so you are warned after a single missed run, not three. We back up from a secondary. Does that affect this card? Not directly: the card reports completion age regardless of which member the backup ran against. But a backup taken from a heavily lagging secondary can complete successfully while capturing stale data. Pair this card with Replica Lag (seconds) so a fresh-looking backup is not quietly behind the primary. The card shows no value at all. What does that mean? A blank or null reading means the engine found no successful backup record for this deployment: either backups have never been configured, the connector cannot see the backup metadata (missing Atlas backup read scope, or a self-managed job that records no success marker), or every recorded run has failed. Treat an empty value as more urgent than a high value: it usually means there is no backup at all.

Tracked live in Vortex IQ Nerve Centre

Last Successful Backup (hours ago) is one of hundreds of KPI pulses Vortex IQ tracks across MongoDB and 70+ other ecommerce connectors. Nerve Centre runs the detection layer; Vortex Mind investigates the cause when something moves; Ask Viq lets you interrogate any number in plain English. Start for free or book a demo to see this metric running on your own data.