Skip to content
Shawn DavisonOct 9, 20266 min read

The Lakehouse Freeway, Part 2: Shared Business Meaning for Apps and Agents

A case for building a Unified Data Platform (UDP) on Databricks, leveraging both Lakehouse (OLAP) and Lakebase (OLTP) engines to curate and share data products across disparate teams and systems.

In Part 1, we opened the freeway: multi-source ingest, Medallion curation, and governed data products, with Lakebase as the application SQL read surface. Operational products kept ownership of their writes.

Pattern 2 keeps that foundation and addresses the next problem: the same business fact can mean something different in every product.

From governed data to shared business meaning: Pattern 2

Pattern 1 makes data accessible through governed read paths. Pattern 2 makes its business meaning explicit. A common structure establishes consistent representation; shared definitions explain how to interpret it.

UDP Pattern 2 architecture: canonical Unified Data Model and tenant schema registry with GraphQL access for apps and AI

Pattern 1 stays the ingest and curation substrate. Pattern 2 turns Gold into a governed canonical model and puts GraphQL in front of it as the application read path; analytics tools still read Gold directly.

Pattern 2 Lakehouse Freeway: canonical customer, bill-to party, and support concepts connected through a governed UDM and read-only GraphQL access.

Pattern 2: a canonical Unified Data Model (UDM)

In a multi-product enterprise, "account," "owner," and "invoice" mean slightly different things in every system. Tenants add their own custom attributes on top. Pattern 1 preserves those extras in an opaque variant column; Pattern 2 governs them.

  • Canonical entities. Accounts, products, opportunities, and invoices have stable keys, types, and defined relationships across source products.
  • A tenant schema registry. Custom attributes are declared, typed, and mapped, with their business definitions and domain or tenant scope. Unmapped keys are quarantined with lineage rather than silently dropped.
  • Graph-based access. A GraphQL gateway sits over the UDM. Applications request the fields and relationships they need without embedding source-specific SQL. Access rules and masking are enforced along that read path.

For AI, the UDM and its accompanying definitions are the payoff. An agent, retrieval pipeline, or Genie experience needs to know what a customer represents, not just which table contains one.

One customer question, three product vocabularies

Consider an illustrative portfolio. The CRM calls a customer an “account.” Billing uses “account” for the party responsible for payment. Support associates an organization with its users and an owner. Those records can describe related parts of the same customer relationship without representing the same entity.

Now ask: “Which active customers using this product have unpaid invoices, and who owns the relationship?”

One SQL read surface makes the records accessible. It doesn't decide what “customer” means, whether “active” refers to a commercial relationship or recent product use, or whether the support owner is also the commercial owner.

Pattern 2 makes those decisions explicit. A canonical customer has a stable key. Source records map to that identity where the relationship is established; invoices, product relationships, and owners connect through defined relationships. A bill-to party that serves several customers stays a related entity rather than being forced into one customer record.

In this example, billing is authoritative for invoice status and the CRM for commercial ownership. “Active” needs a stated definition and time window. If sales and product operations use different definitions, preserve both with domain-qualified names rather than force them into one ambiguous field.

Canonical does not mean flatten everything into one table. Agree on the entities and relationships the use case needs, and make differences explicit. Ambiguous matches need resolution, not a confident guess from an agent.

Where the meaning lives

The UDM carries stable identities, relationships, and curated values. Accompanying definitions capture entity meaning, important states, business rules, authoritative sources, and domain or tenant scope. Rules that calculate a shared value belong in governed transformations or metric definitions, not in a different prompt or report for every consumer.

Give those definitions an accountable domain owner and treat them as versioned contracts. A change to “active customer” or the source of an attribute is reviewed alongside its effect on curated values, APIs, metrics, and AI context. Consumers should be able to trace the answer to the definition and source that produced it.

This is the useful role of an ontology: describe the business concepts, their relationships, and the meanings the model preserves. It does not require a separate graph database or an exhaustive enterprise model before the first use case.

Tenant attributes become part of the contract

The standard model won't cover every tenant's business. Suppose one tenant stores customer segment as segment_code in its CRM and customer_band in another product. If both describe the same business fact, the registry declares a canonical customerSegment attribute, its definition, type, authoritative source, and mappings. Different value sets need an explicit translation; similar names alone aren't enough.

If the fields represent different segmentation schemes, preserve that distinction. An undeclared key doesn't quietly become a shared field: it is quarantined with lineage for review. Changes to the mapping are changes to the contract consumers depend on.

Start with the domains and relationships needed for a useful question. Extend the model as new use cases arrive.

Shared meaning, defined consumption paths

Lakebase is standard Postgres. That makes it a practical distribution layer for application data through familiar drivers and SELECT-only credentials. It is not the required destination for every analytical or AI workload.

  • Applications and APIs. Lakebase serves curated UDM data. GraphQL resolves defined relationships and exposes fields through a shared application contract; field descriptions and resolver behavior must reflect the governed definitions.
  • Analytics and conversation. SQL and Genie use governed Gold data and metric definitions, with business terms available as context.
  • Agents and retrieval. Agentic tools (MCP servers, Agent Bricks, Genie spaces, Vector indexes, etc.) can use selected curated data with the definitions, source references, and tenant scope needed to interpret it.

GraphQL reduces repeated joins and source-specific code. It does not define business meaning by itself. Likewise, copying a table into Postgres or an index does not automatically carry its definitions with it. Publishing the data product includes making the relevant meaning available through each selected path.

Databricks semantics and ontology

Unity Catalog semantics provides governed metric views, domains, Pages for business concepts, and certification signals. These give the model's definitions and authority a concrete place alongside catalog data.

Genie Ontology combines that modeled context with context inferred from assets and usage for Genie One and Genie Code. It supports business-aware interpretation in those experiences. Application APIs and other agent tools still need explicit contracts for the definitions they consume.

Governance follows the read path

A shared graph is not permission to see every tenant’s data. The gateway needs authenticated tenant and user context, and each resolver must restrict the records and fields it returns. SELECT-only credentials limit writes; they do not establish which tenant's rows a consumer may read.

Underlying Postgres permissions, tenant boundaries, and masking must match the gateway's access contract. SQL, Genie, and retrieval paths need corresponding controls at their own entry points. The gateway cannot protect a query that bypasses it.

Pattern 2's graph is read-only. Product-owned APIs can support governed actions while operational databases remain where they are; migrating those databases is a separate architectural decision.

Shared meaning, reused by every consumer

For AI, the UDM is the payoff when its business meaning is explicit. Shared identities, relationships, definitions, and governed access give applications and agents a common starting point without forcing every domain into one definition.

Patterns 1 and 2 form a complete architecture with existing operational systems in place. Part 3 explores an optional extension: moving selected product workloads onto Lakebase to bring transactions, curated data, and AI into a tighter loop.

Previous: The Lakehouse Freeway, Part 1: Your AI Strategy is a Data Architecture Problem

Next: The Lakehouse Freeway, Part 3: Real-Time Data + AI on One Platform

Ready to see where your data estate sits on the path? Talk to DevIQ about a UDP readiness assessment. We’ll map your sources, identify your first high-value data products, and outline a Pattern 1 roadmap you can put into production.

avatar
Shawn Davison
Shawn is a serial entrepreneur and software architect with 25+ years of experience building innovative technology used by millions worldwide. Beyond transforming concepts into realities, Shawn is an IRONMAN Triathlete, loves photography, the outdoors, snowboarding, and kitesurfing.