> ## Documentation Index
> Fetch the complete documentation index at: https://docs.vortexiq.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Databricks on Vortex IQ

> Monitor Databricks health, cost and reliability signals, and catch incidents and runaway spend early.

Monitor Databricks health, cost and reliability signals, and catch incidents and runaway spend early.

[Connect or manage this source](https://app.vortexiq.ai/workbench/settings/sources) · [How connecting works](/integrations/connector-catalogue) · [Create a workflow](https://app.vortexiq.ai/workbench/flows/create?connector=databricks)

<CardGroup cols={5}>
  <Card title="31">
    performance signals
  </Card>

  <Card title="7">
    automated checks
  </Card>

  <Card title="Build your own">
    automated fixes
  </Card>

  <Card title="Ready to build yours">
    workflows
  </Card>

  <Card title="6">
    API operations
  </Card>
</CardGroup>

<Tabs>
  <Tab title="Overview">
    ### What you can achieve

    Capabilities are grouped around merchant outcomes, not API terminology.

    <CardGroup cols={2}>
      <Card title="Customer experience">
        Find storefront, speed, accessibility and journey problems.
      </Card>

      <Card title="Run operations">
        Monitor orders, fulfilment, delivery and settlement.
      </Card>

      <Card title="Protect revenue">
        Find failures, leaks and risks before they cost sales.
      </Card>

      <Card title="Control risk and change">
        Keep tracking, access and change under governed control.
      </Card>
    </CardGroup>

    ### From connection to verified outcome

    The controlled sequence every capability follows. Nothing changes a connected system without the approval step.

    <Steps>
      <Step title="Connect">
        Authorise the source. Scopes are shown before access is granted.
      </Step>

      <Step title="Monitor">
        Watch the signals against your own baselines, not universal defaults.
      </Step>

      <Step title="Detect">
        Run checks and gather evidence specific to your store.
      </Step>

      <Step title="Recommend">
        Explain what happened, why it matters and the proposed action.
      </Step>

      <Step title="Approve">
        You review scope, risk and reversibility before anything changes.
      </Step>

      <Step title="Execute">
        Apply through governed connector operations.
      </Step>

      <Step title="Verify">
        Confirm the intended result and keep the receipt.
      </Step>
    </Steps>

    No changes are made without the configured approval policy. Read-only operations do not modify the connected system; schedules, access scopes, API usage and data handling remain governed by Vortex IQ controls.
  </Tab>

  <Tab title="Monitor (31)">
    ### Monitor performance

    31 performance signals. Open an outcome to see its signals and how each one alerts. Read-only operations do not modify the connected system.

    <AccordionGroup>
      <Accordion title="Run operations (15 signals)">
        | Signal                                | Alert behaviour    | What it tracks                                                                                                                                         |
        | ------------------------------------- | ------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------ |
        | **Active Clusters**                   | Watch only         | Description pending editorial review; the signal is live.                                                                                              |
        | **Active SQL Sessions**               | Watch only         | Description pending editorial review; the signal is live.                                                                                              |
        | **Active SQL Warehouses**             | Watch only         | Description pending editorial review; the signal is live.                                                                                              |
        | **Avg DBU per Job Run**               | Merchant rule      | Description pending editorial review; the signal is live.                                                                                              |
        | **DBU Burn +50% Week-over-Week**      | Merchant rule      | Alerts for DBU Burn +50% Week-over-Week.                                                                                                               |
        | **DBU Burned (24h)**                  | Merchant rule      | From billable usage API. Databricks-defining cost metric. Includes job compute + SQL warehouse + interactive.                                          |
        | **DBU by Cluster (7d)**               | Watch only         | DBU by Cluster (7d).                                                                                                                                   |
        | **Databricks Health Score**           | Merchant rule      | Description pending editorial review; the signal is live.                                                                                              |
        | **Idle Cluster DBU Wasted (24h)**     | Merchant rule      | DBU spent while cluster had no active jobs. Auto-termination tuning signal.                                                                            |
        | **Last Delta Lake Vacuum / Optimize** | Alert band 24 / 72 | Delta Lake doesn't have traditional backup , Time Travel (default 7d retention) + table OPTIMIZE/VACUUM stand in. This tracks last successful OPTIMIZE |
        | **Long-Running Jobs (>1h)**           | Merchant rule      | Jobs still running past expected duration , common cost-runaway pattern.                                                                               |
        | **Pipeline Lag (since last success)** | Alert band 1 / 10  | Description pending editorial review; the signal is live.                                                                                              |
        | **SQL Queries per Hour (live)**       | Watch only         | Description pending editorial review; the signal is live.                                                                                              |
        | **SQL Warehouse Saturation %**        | Alert band 70 / 90 | Description pending editorial review; the signal is live.                                                                                              |
        | **Top 10 Failing Workflows (7d)**     | Watch only         | Top 10 Failing Workflows (7d), broken down by row.                                                                                                     |
      </Accordion>

      <Accordion title="Protect revenue (8 signals)">
        | Signal                                      | Alert behaviour | What it tracks                                                                                                |
        | ------------------------------------------- | --------------- | ------------------------------------------------------------------------------------------------------------- |
        | **DBU Burn vs Ecom Order Volume**           | Merchant rule   | Databricks-distinctive XC , DBU burn should track ecom volume. Divergence = inefficient pipelines.            |
        | **DLT Pipeline Status Distribution**        | Watch only      | Running / Idle / Failed / Stopped , from /pipelines/list.                                                     |
        | **Databricks SQL Spike vs Ecom Order Rate** | Merchant rule   | Description pending editorial review; the signal is live.                                                     |
        | **Failed Job Burst (>5 failures in 1h)**    | Merchant rule   | Databricks-distinctive , pipeline failures cascade fast across dependent jobs.                                |
        | **Failed Jobs (24h)**                       | Merchant rule   | result\_state=FAILED OR TIMEDOUT. Triage queue.                                                               |
        | **Job Success Rate (24h)**                  | Merchant rule   | From /jobs/runs/list , Databricks-distinctive defining metric. Failed scheduled runs = data pipelines broken. |
        | **Pipeline Lag vs Ecom Order Flow**         | Merchant rule   | Description pending editorial review; the signal is live.                                                     |
        | **Slow SQL Queries During Checkout Window** | Merchant rule   | Slow SQL Queries During Checkout Window, broken down by row.                                                  |
      </Accordion>

      <Accordion title="Customer experience (6 signals)">
        | Signal                            | Alert behaviour      | What it tracks                                                                           |
        | --------------------------------- | -------------------- | ---------------------------------------------------------------------------------------- |
        | **Avg Cluster CPU Utilisation %** | Merchant rule        | Databricks-distinctive , drives right-sizing decisions for cost + performance.           |
        | **SQL Query Latency p50 (ms)**    | Watch only           | Description pending editorial review; the signal is live.                                |
        | **SQL Query Latency p95 (ms)**    | Alert band 50 / 200  | Lakehouse SQL p95 measured in seconds typically , threshold reflects warehouse workload. |
        | **SQL Query Latency p99 (ms)**    | Alert band 100 / 500 | Description pending editorial review; the signal is live.                                |
        | **Slow-Query Rate %**             | Alert band 1 / 5     | Description pending editorial review; the signal is live.                                |
        | **Top 10 Slowest SQL Queries**    | Watch only           | Top 10 Slowest SQL Queries, broken down by row.                                          |
      </Accordion>

      <Accordion title="Control risk and change (2 signals)">
        | Signal                                     | Alert behaviour    | What it tracks                                            |
        | ------------------------------------------ | ------------------ | --------------------------------------------------------- |
        | **SQL Query Error Rate %**                 | Alert band 0.1 / 1 | Description pending editorial review; the signal is live. |
        | **SQL Query Error Rate Spike (>1% in 5m)** | Alert band 0.1 / 1 | Alerts for SQL Query Error Rate Spike (>1% in 5m).        |
      </Accordion>
    </AccordionGroup>
  </Tab>

  <Tab title="Audit (7)">
    ### Audit risks and opportunities

    A fix status appears only where the action, inputs, approval, verification and recovery controls are mapped. Candidate remediations are never executable. Open a check for the detail.

    <AccordionGroup>
      <Accordion title="Connection pool saturation above 90%">
        **Severity** critical · **Outcome** Customer experience · **Fix status** Report only

        At 90% of the connection pool in use, the database is close to refusing new connections outright. Once it does, every part of the application that needs a fresh database connection, including new customer sessions and checkout, starts failing, not just slowing down.

        Vortex IQ detects and explains this; resolution is manual, with evidence and recommended steps.

        Reference: `DB-CAP-001`
      </Accordion>

      <Accordion title="Disk usage above 90%">
        **Severity** critical · **Outcome** Run operations · **Fix status** Report only

        A database that runs out of disk stops accepting writes entirely, which for most stores means orders, inventory updates and customer records stop being saved, not just that the database gets slower. There is very little runway left at 90%.

        Vortex IQ detects and explains this; resolution is manual, with evidence and recommended steps.

        Reference: `DB-CAP-002`
      </Accordion>

      <Accordion title="Query error rate above 1% in last 5 minutes">
        **Severity** critical · **Outcome** Run operations · **Fix status** Report only

        More than 1 in 100 queries is failing right now. Depending on what those queries do, this can mean orders not saving, pages failing to load product or customer data, or background jobs silently dropping work, and a rate this high in a 5-minute window is an active incident, not background noise.

        Vortex IQ detects and explains this; resolution is manual, with evidence and recommended steps.

        Reference: `DB-ERR-001`
      </Accordion>

      <Accordion title="Last successful backup older than 72 hours">
        **Severity** high · **Outcome** Run operations · **Fix status** Report only

        If something goes wrong with this database right now, the most recent point it can be restored to is over 3 days old. Every order, customer record and inventory change since that backup would be unrecoverable in a real incident, not just delayed.

        Vortex IQ detects and explains this; resolution is manual, with evidence and recommended steps.

        Reference: `DB-BAK-001`
      </Accordion>

      <Accordion title="Replication lag above 10 seconds">
        **Severity** high · **Outcome** Run operations · **Fix status** Report only

        Anything reading from the replica, reports, dashboards, or read traffic split off the primary for capacity, is now up to 10+ seconds stale. If the primary fails while lag is this high, the replica is also that far behind on failover, which is a bigger problem than the staleness alone.

        Vortex IQ detects and explains this; resolution is manual, with evidence and recommended steps.

        Reference: `DB-REP-001`
      </Accordion>

      <Accordion title="Slow-query rate above 5% of total">
        **Severity** high · **Outcome** Customer experience · **Fix status** Report only

        More than 1 in 20 queries is landing in the slow bucket. That is frequent enough to be a pattern, not noise, and it means a meaningful share of every page load or job that touches this database is paying the slow-query cost, not just an unlucky occasional request.

        Vortex IQ detects and explains this; resolution is manual, with evidence and recommended steps.

        Reference: `DB-PERF-002`
      </Accordion>

      <Accordion title="p95 query latency above 200ms sustained 15m">
        **Severity** high · **Outcome** Customer experience · **Fix status** Report only

        One in twenty queries against this database is taking over 200ms, sustained for at least 15 minutes, not a brief spike. Any storefront page, checkout step or order sync that depends on this database inherits that slowness directly, and a sustained p95 this high is usually already visible to customer

        Vortex IQ detects and explains this; resolution is manual, with evidence and recommended steps.

        Reference: `DB-PERF-001`
      </Accordion>
    </AccordionGroup>

    #### Build your own automated fixes

    Turn any finding into an automated fix with a Vortex IQ workflow: **over 13,000 read and write operations across more than 200 connectors** are available as building blocks, with approval, verification and rollback on every change.
  </Tab>

  <Tab title="Automate">
    ### Automate approved work

    Vortex IQ is integrated with **6 read** and **0 write** operations across billableusagedownloads, clusterlists, jobrunlists, pipelines, sqlhistoryquerys, sqlwarehous on Databricks. Combine them with anything from the **over 13,000 operations across more than 200 connectors** to automate the work in your own words.

    Changes follow your configured approval policy: the target, proposed change, affected records, risk, reversibility and verification plan are shown before execution.

    [Create a workflow](https://app.vortexiq.ai/workbench/flows/create?connector=databricks)

    #### Ready to build your first Databricks workflow

    Pick a trigger, add the operations above as steps, and every step that changes data pauses for your approval. Monitoring and audits are live now and can start any workflow you build.

    <Accordion title="Browse the operations you can build with">
      | Resource               | Read operations | Write operations |
      | ---------------------- | --------------- | ---------------- |
      | billableusagedownloads | 1               | 0                |
      | clusterlists           | 1               | 0                |
      | jobrunlists            | 1               | 0                |
      | pipelines              | 1               | 0                |
      | sqlhistoryquerys       | 1               | 0                |
      | sqlwarehous            | 1               | 0                |

      Signed-in users see the full catalogue in the workflow builder, filtered to the sources they have connected.
    </Accordion>
  </Tab>
</Tabs>
