Ask three teams in your company to define “active customer” and you’ll get four answers. Marketing counts anyone who opened an email this quarter. Finance counts paying accounts. Product counts monthly active users. And the executive dashboard shows a fifth number that nobody can explain but everyone quotes.
This is the metrics mess, and it’s the most expensive problem in enterprise data that nobody’s job title mentions. Billions get spent on lakes, warehouses, and pipelines, and then the business argues about what the numbers mean. The technology worked. The meaning didn’t.
The fix has a name, and it’s the most underrated architecture in the data world: the semantic layer.
What a semantic layer actually is
A semantic layer is a single, governed definition of every business metric and dimension, sitting between the raw data and everyone who consumes it. “Revenue” is defined once — the logic, the filters, the grain, the edge cases — and every dashboard, report, API, and AI agent uses that definition.
Not a wiki page describing revenue. Not a Confluence doc everyone ignores. The actual logic, executable, versioned, and shared. When someone asks “how is revenue calculated?”, the answer is the semantic layer, and it’s the same answer everywhere.
If your metrics are defined in forty dashboards instead of one place, you don’t have metrics. You have opinions with charts.
Why the old answers failed
We’ve tried to solve this before. The data warehouse was supposed to be the single version of the truth — but the logic leaked into BI tools, dbt models, spreadsheets, and one-off SQL. The “metrics layer” wave of a few years ago got the idea right but often landed as yet another tool to adopt rather than a discipline to enforce.
The failure pattern is always the same: the definitions live wherever it’s convenient in the moment, and convenience compounds into chaos. A semantic layer only works when it’s the only place metric logic is allowed to live. That’s a governance decision before it’s a technology decision.
Conceptual, logical, physical: the modeling discipline underneath
Under every good semantic layer is old-fashioned data modeling discipline — the kind that fell out of fashion and is now quietly saving AI projects.
Conceptual modeling answers “what are the things we care about and how do they relate?” — customers, orders, products, the entities and relationships, in business language. Logical modeling adds structure: attributes, keys, grain, the rules that hold regardless of technology. Physical modeling is the implementation: tables, partitions, indexes, the performance reality.
Most organizations skipped straight to physical — built the tables, worried about meaning later. Later never came. The semantic layer forces the discipline back in: you can’t define “revenue” once, governably, without agreeing on what a customer is and what counts as an order. The modeling conversation is the alignment conversation.
The business glossary as a contract
The glossary is the semantic layer’s public face: every metric and dimension, defined in plain language, with an owner, a lineage trail, and a version history. But the key word is contract. A glossary that nobody is obligated to use is documentation. A glossary wired into the platform — where the BI tool pulls definitions from it, where data contracts reference it, where new pipelines must map to it — is infrastructure.
The test is simple: can a new analyst find the definition of a metric, see who owns it, trace it to source, and trust it — in an afternoon? If yes, you have a contract. If it takes a quarter and five Slack threads, you have a wiki.
The semantic layer and AI
Here’s what makes this urgent instead of merely important: AI agents need shared meaning more desperately than humans do.
A human analyst who sees a weird revenue number will squint, ask around, and figure out the definition drift. An agent won’t. It will take the number, act on it, and compound the error into decisions. Worse, agents from different teams will each invent their own interpretation of the same metric — automated definition drift, at machine speed.
The semantic layer is what grounds AI in shared business meaning. Retrieval gives agents facts; the semantic layer gives them definitions. An agent that queries “revenue” through the semantic layer gets the governed answer. One that writes its own SQL gets whatever it feels like.
In the agentic era, the semantic layer isn’t a BI convenience. It’s the difference between grounded AI and confident fiction.
Where to start
Don’t boil the ocean. Pick the ten metrics the business argues about most — revenue, active customers, churn, cost, whatever sparks the weekly debate — and define them once, in one place, with owners and lineage. Wire one BI tool to consume them. Show the before and after: same question, one answer.
Then expand. The goal isn’t a perfect enterprise ontology on day one; it’s breaking the pattern where every new dashboard invents its own definitions. Ten governed metrics, actually used, beat a hundred documented ones nobody trusts.
The enterprises that get this right will discover something their competitors won’t: when the AI era demands grounded, consistent, explainable answers, they’ll already have the layer that provides them. Everyone else will be retrofitting meaning onto chaos — while their agents confidently report five different revenues.
The work is unglamorous. The payoff isn’t. Define it once.
— Jugal

Thanks for the comment, will get back to you soon… Jugal Shah