Skip to content
Lakehouse Freeway unified data platform connecting enterprise sources, canonical data, applications, analytics, and AI
Shawn DavisonOct 2, 20267 min read

The Lakehouse Freeway, Part 1: Your AI Strategy is a Data Architecture Problem

A case for building a Unified Data Platform (UDP) on Databricks, leveraging both Lakehouse (OLAP) and Lakebase (OLTP) engines to curate and share data products across disparate teams and systems.

Most enterprises don't have an AI model problem. They have a data access problem that AI makes impossible to ignore.

Gartner predicts that through 2026, organizations will abandon 60% of AI projects that aren't supported by AI-ready data. The same research found 63% of organizations either lack, or aren't sure they have, the right data management practices for AI (Gartner).

The pattern is familiar. A pilot works on a hand-curated extract. Then it meets production: customer records split across five products, three database engines, and a decade of point-to-point integrations. Every agent, copilot, and model needs the same answer to "who is this customer?" – and nobody can give it with confidence.

That's the job of a unified data platform (UDP): shared infrastructure for the data work that applications, analytics, and AI depend on.

The 80/20 rule of enterprise AI

In DevIQ's delivery experience, about 80% of the effort in an AI program is data work and 20% is the AI itself. That held for classic machine learning (ML), and so far, it appears to be holding for agentic AI.

"Then" (AI/ML Wave)

"Only a small fraction of real-world ML systems is composed of the ML code… The required surrounding infrastructure is vast and complex."
– Sculley et al., Google, NeurIPS 2015

"and Now" (Agentic AI Wave)

50% of data leaders name data quality as the top challenge in deploying agentic AI, and 57% see data reliability as a key barrier to moving AI from pilot to production.

– Informatica CDO Insights 2026

Agentic AI changes the shape of the data work. Agents need consistent entities, permissions enforced at query time, fresh operational state, and durable memory. A UDP provides a governed foundation for those requirements.

The implication for leaders: you can't skip the data work, but you can make it reusable. A UDP turns it into shared infrastructure, so each new AI use case starts from governed data instead of a fresh data project.

What a unified data platform is

At DevIQ, we describe a UDP as a lakehouse freeway: one governed road that every product, analyst, and AI workload travels to reach enterprise data – instead of a tangle of private side streets between every application and every source.

Three design principles define it:

  • Operational systems stay the writers of their own records. The CRM owns customer edits; billing owns invoices. The UDP never writes back into a product's OLTP database.
  • The platform ingests, curates, and publishes. Raw data comes in via well-defined lanes based on business cadence (streaming or batch), is cleaned and subjected to governance policies, and then published as trusted data products.
  • Consumers share one governed foundation. Applications, analytics, and AI use defined read paths appropriate to their workloads rather than becoming writers of each other's stores.

This one-way design is deliberate. When a stakeholder asks to "sync UDP data into my product's database," we remodel it as a read-path integration. That constraint is what keeps the platform simple, auditable, and safe to scale.

Why it matters for AI specifically

AI multiplies the cost of fragmented data. A scheduled dashboard and an agent acting on a customer record have different freshness needs, but both need consistent entities, governed access, known lineage, and data fresh enough for the task. A UDP establishes those controls in shared infrastructure rather than rebuilding them inside each AI project.

Lakehouse Freeway unified data platform connecting enterprise sources, canonical data, applications, analytics, and AI

Inside the architecture: DevIQ UDP Pattern 1

DevIQ builds UDPs around three related patterns. Pattern 1 – multi-source ingest, centralized SQL read – establishes the ingest and curation foundation. Pattern 2 adds a shared business model. Together, they form a useful architecture with operational systems staying where they are. Pattern 3 is an optional extension for selected workloads.

UDP Pattern 1 architecture: multiple data sources feed Medallion curation and a centralized Lakebase SQL read surface

Data flows left to right, one way: many sources in, curated data products out. The diagram shows the Lakebase SQL serving path for applications; analytical and conversational workloads can read Gold directly.

1. Multi-source ingest, declared as code

A versioned ingestion manifest and per-table schema contracts declare what's in scope. Adding a table is a pull request with a visible scope diff, not a click in a connector UI. Three acquisition methods land side by side in Bronze:

Method Use it when
Change data capture (CDC) High-change operational databases like SQL Server, Postgres, or MySQL
External tables Upstream ETL already lands Iceberg or Delta in object storage
Query federation Read-only or reference sources where CDC isn't justified

Pattern 1 Opening the Freeway: governed ingestion, Bronze Silver Gold curation, and Lakebase read access

2. Medallion curation

  • Bronze is the raw landing zone, with idempotent upserts keyed on source primary keys.
  • Silver normalizes entities and masks PII at write. Data quality expectations quarantine bad rows, and CI fails a pipeline that ships without them.
  • Gold holds curated data products, ready for applications, analytics, and AI.

3. Lakebase as the application-serving destination

For application workloads, Gold is projected into Lakebase, Databricks' managed Postgres, through metadata-driven synced tables. Product teams connect with standard Postgres drivers and SELECT-only credentials. Freshness is measured as the lag from a Gold change to Postgres visibility, with alerts against an SLA.

Lakebase is an opinionated serving choice, not a requirement for every consumer. Analytics and Genie can query governed Gold directly; retrieval workloads use governed indexes built from curated data. The common asset is the foundation, not one mandatory interface.

The quiet superpower: decoupling from source engines

Because consumers read curated data products, they aren't coupled to the source database. A product can migrate its OLTP store from SQL Server to PostgreSQL alongside its app upgrade. The platform retargets the ingest path; downstream consumers keep their read contracts, provided the data product's keys, schema, and meaning stay stable.

Why Databricks fits the foundation

A UDP is an architecture, not a product. But the platform underneath decides how much of it you build versus configure. Pattern 1 brings ingestion, curation, governance, and a standard SQL read surface together.

Unity Catalog: governance across the read paths

Unity Catalog governs lakehouse data through grants, column masks, tags, and lineage. Those controls support pipelines, SQL warehouses, dashboards, and Genie.

Lakebase projections and application gateways need matching Postgres permissions, tenant boundaries, and masking controls on their own read paths. A shared governance design does not mean every source policy automatically carries through every interface. Each consumer must receive only the data its identity is allowed to read.

Lakeflow Connect: ingestion without the integration tax

Lakeflow Connect provides 100+ built-in connectors for databases, SaaS applications, and file sources (Databricks). In our pattern, managed database CDC connectors move changed data into Bronze, driven by the ingestion manifest. Teams can focus on source contracts and pipeline behavior rather than hand-building connectors.

Lakehouse + Lakebase: OLAP and OLTP on one platform

The Lakehouse handles analytical scale: Medallion pipelines on Delta, SQL warehouses, and AI/BI. Lakebase adds managed Postgres and synced tables for operational serving.

Having analytical and operational capabilities on one platform makes the application read surface practical. Gold data products project into Lakebase via synced tables, while application-serving permissions remain an explicit part of the design.

What the first lane unlocks

You can’t skip the data work, but you can make it a shared asset. Pattern 1 establishes ingestion contracts, quality controls, and governed read paths that keep consumers independent of source engines.

  • Reusable integration. Governed data products replace repeated source-specific integrations as the portfolio grows.
  • Trusted answers. Quality expectations and quarantine keep bad rows out of the data your models and agents act on.
  • Freedom to modernize. Source systems can change engines on their own schedule while preserving the contracts AI and analytics depend on.

These are familiar lakehouse building blocks. The architectural contribution is the discipline around them: declared scope, quality controls, write ownership, and stable serving contracts.

The next challenge is meaning: if every product defines “account” differently, shared access still isn’t shared understanding. Part 2 builds that business model on the foundation.

Next: The Lakehouse Freeway, Part 2: Shared Business Meaning for Apps and Agents

Ready to see where your data estate sits on the path? Talk to DevIQ about a UDP readiness assessment. We’ll map your sources, identify your first high-value data products, and outline a Pattern 1 roadmap you can put into production.

avatar
Shawn Davison
Shawn is a serial entrepreneur and software architect with 25+ years of experience building innovative technology used by millions worldwide. Beyond transforming concepts into realities, Shawn is an IRONMAN Triathlete, loves photography, the outdoors, snowboarding, and kitesurfing.