Omnitrace
Databricks agent platform

Your lakehouse spends money while you sleep.

Omnitrace watches 24/7, reasons in dollars, and ships the fix — before you open your laptop.

app.omnitrace.io / agent-actions
live
VERIFIED OT-MAT-487
$2,568/yr

Idle cluster auto-terminated

0428-194212-violet42

Applied 14:23 → Verified 14:27

APPLIED OT-WHO-203
$22,080/yr

Warehouse auto-stop enabled

Northwind-ETL · serverless

Applied 14:21 · verifying...

INVESTIGATING OT-OPT-994
$5,040/yr est.

OPTIMIZE proposed

main.analytics.events

Awaiting blast-radius approval

50+detector types
19auto-fix paths
5 minapply → verify
3-tierguardrails

Omnitrace is an agentic observability platform for Databricks that helps enterprises reduce cloud costs, improve reliability, and automate operational remediation.

The FinOps Agent identifies hidden spending and optimization opportunities, while the Reliability Agent detects performance bottlenecks and operational risks across clusters and workloads. With autonomous remediation and human-in-the-loop escalation, Omnitrace helps engineering teams spend less time firefighting and more time building.

FinOps Agent — for Databricks Reliability Agent — for Databricks

Enterprise proof

A closed loop your platform team can audit.

Omnitrace is built for teams that need more than alerts. Every recommendation is tied to evidence, every action is governed, and every fix is verified after it runs.

No customer table contents

No query result sets

No business records

Cloud-bound model option

Scoped credentials

Human approval controls

Step 1

Connect

Read-only operational metadata

Databricks, cloud cost, system-table, configuration, owner, and workflow signals enter the agent loop without customer rows, files, or query results.

Step 2

Reason

Cost and reliability evidence graph

The agent correlates spend, workload behavior, policy drift, failures, and ownership to decide what matters and why.

Step 3

Act

Guarded MCP tool calls

Approved actions move through scoped tools, Jira workflow, autonomy policies, blast-radius checks, and savings thresholds.

Step 4

Verify

Read-back proof

Omnitrace checks the target state after every action and keeps verification evidence with the finding and action record.

Architecture

From lakehouse telemetry to verified action.

Omnitrace connects Databricks and cloud signals to detectors, Atlas Playbook agents, governed remediation, and verified outcomes.

View architecture

01

Lakehouse signals

Databricks, cloud cost, jobs, SQL, Spark, ownership, and workflow metadata.

02

Omnitrace agents

Detectors and Atlas Playbook agents reason over metadata and operational telemetry.

03

Governed action

Policy gates, approvals, scoped tools, and autonomy levels control every change.

04

Verified outcomes

Read-back checks prove what changed and keep evidence with the action record.

The problem

Your lakehouse leaks ~$100K/year.

Dashboards see it. Tickets pile up. Nobody fixes it.

Untracked annual waste, per workspace
Idle DBUs · forgotten clusters $38,400
Warehouses without auto-stop $22,200
Photon-eligible on STANDARD $28,800
Slow queries on stale tables $11,520
Total per year $100,920

How it works

Not a dashboard. A teammate.

Every minute, on every workspace. Continuously.

1 Sense

Continuously senses your lakehouse

Agent ingests rich telemetry across compute, queries, tables, billing, and infrastructure — building a live model of your environment.

2 Reason

Reasons over evidence

AI weighs signals, quantifies dollar impact, and writes natural-language narration explaining why action is justified.

3 Act

Acts through MCP-orchestrated tools

Within your autonomy budget, the agent calls the right write-tool — adjusting cluster, warehouse, or table state. Auditable, reversible, sandboxed.

4 Verify

Verifies its own work

Two-tier verification reads back the post-action state. If the change didn't stick, the agent flags REGRESSED and notifies your team.

Continuous loop — every minute, the agent re-evaluates and re-acts

Why Omnitrace

Outcomes, not dashboards.

Native tools Generic obs Manual Omnitrace
Lakehouse-aware
Quantifies in dollars
Autonomous remediation
Verifies own work
Multi-tier cost attribution
Bring-your-own MCP tools

Capability matrix

Built for FinOps, SRE, and data platform operators.

The product surface connects cost evidence, reliability symptoms, workflow ownership, governed action, and verification in one operating loop.

Cost attribution

DBU, cloud cost, workspace, owner, and workload evidence

Detector coverage

50+ implemented detector types across FinOps, SQL, Spark, storage, and governance

Remediation coverage

19 auto-fix paths with manual, approval-required, and low-risk autonomy modes

Workflow integration

Jira-ready context, owner routing, approval state, and audit trail

Enterprise boundary

Metadata-only design with cloud-provider-hosted model options

Verification

Scheduled polling, targeted read-back checks, and outcome evidence

What the agent fixes

One agent. Fifty-plus ways to find waste.

56 detector implementations across FinOps, SQL, Spark reliability, table hygiene, and storage. 19 auto-fix paths are already wired into the apply-and-verify loop.

Idle cluster termination Warehouse auto-stop Photon enablement OPTIMIZE & ZORDER VACUUM & snapshot cleanup Cluster policy attach Small-files compaction Auto-termination defaults Spark AQE enablement Shuffle partition tuning Spill-heavy stage tuning GC thrash mitigation Executor OOM recovery Executor right-sizing Data skew mitigation Cluster mode correction SQL anti-pattern triage S3 request hotspots Abandoned table cleanup Cost spike triage Multi-platform compute Tool marketplace
Live Shipping next On the roadmap

Roadmap

Where the agent goes next.

Real-time Spark telemetry. Multi-platform compute. Spark config auto-tuning. Tool marketplace.

See the full roadmap

Q3 · Now

Real-time Spark telemetry

Q4 · Next

Multi-platform: EMR, Starburst, Snowflake

2026

Spark config auto-tuning + tool marketplace

Ready to put the agent to work?

Connect operational metadata, prioritize verified savings, and move approved Databricks fixes through the agent loop.