Obsessed With Data Lineage

Data Governance, Unfiltered. | Part 3 of 7

Our Data catalogs are filling up faster than they can answer for what they hold. We keep metadata harvesting, documenting, and documenting again, while the questions that matter most go unanswered.

Five of them stay with us:

  • What data matters most?
  • Can we trust it?
  • Who owns it?
  • How does it shape our AI use cases?
  • What business decisions depend on it?

Traditional Data governance had a good run, and it taught us a great deal. The horizon looks different now. The innocent question we used to chase, where did this data come from, is giving way to a harder one. Which dataset trained this model? The data quality monitoring we once set up for our pipelines is quietly making room for continuous anomaly detection.

No executive sponsor for data lineage writes a check for documentation.

Lessons From the Busiest Subway on Earth

The New York City subway opened in 1904 with 28 stations. It now spans 472 and carries nearly 1.3 billion trips a year, roughly 3.6 million on an average day, as if the entire city of Los Angeles boarded a train every morning.

Running a system that size is not one machine but many layers stacked on each other: train movement tracked in real time, signaling and switching, the physical infrastructure of track and power and stations, passenger flow, and a daily schedule that folds in maintenance without stopping the trains. Every layer has to be known accurately for the whole to hold.

From all that complexity, what reaches a commuter is remarkably simple, with the hard machinery kept under the hood. A clean map shows the line, a screen gives the next arrival, and a notification names a delay or a faster route. The signaling and the schedules and the live tracking stay out of view, and the rider sees only what matters, in the plainest form, right when it counts. That is why they trust and rely on the busiest rapid transit system in the Western Hemisphere.

The subway mastered two disciplines: knowing the whole system, and communicating only what each rider needs. Have we ever done the same for our business users, taken the complete picture of our data, and made it simple enough to trust?

Two Problems: Hard to Capture, Harder to Explain

Here, the comparison gets humbling. The subway runs on rails. Trains move along known tracks on known schedules, and a train’s position is a physical fact you can measure. Data movement offers no such mercy. It threads through pipelines and transformations and handoffs that change constantly, much of it unobservable unless someone instrumented it, and so capturing all of it is a genuine struggle. The lineage platforms hold a great deal of what was collected, but even the teams who built that record have lost the thread of how complete it really is.

That is the first hard problem, and it is the easier of the two. The harder one is the communication that the subway makes look simple. It means taking what we hold and rendering it back to a business user in terms they understand, in the context of how they use the data and the decisions they make, nurturing their trust and confidence in it. This is where most lineage quietly fails. We export the technical flow exactly as the tools recorded it, a knot of systems and columns I have come to call Enterprise Spaghetti, and hand it to someone who only wanted to know whether they can trust the number in front of them.

A business user wants more than the raw technical flow. They want to know, in their own language, whether the number on their report can be trusted. Which raises the question we have avoided for years. If our lineage cannot answer that in terms a business understands, what exactly have we been building?

Lineage: The Map the Business Actually Reads

Data lineage is the knowledge of a data element’s flow from its point of origin to its point of provision. It describes where data starts, how it moves, the transformations it goes through, and its quality along the way.

That knowledge can be told at two very different altitudes. Technical lineage describes how data moves through the machinery: system-level, where applications connect across the enterprise data landscape; table-level, where datasets feed one another; and column-level, tracing a single data element through every transformation. It is the view that the technology team needs to build, debug, and maintain the pipelines. Business lineage, in simpler terms, traces data through business concepts rather than tables and pipelines, showing how a report or a metric ties back to a trusted source, who owns it, and how current it is.

Good business lineage carries the things a business actually asks for: the KPI definition and its glossary term, the transformation rules in plain language, the authoritative system of record, the quality metrics, who is accountable, whether the metric is official or still experimental, and what breaks if it changes.

A Graph Database Is Not Data Lineage

Organizations reach for a graph database to illustrate data lineage because it is already in the stack. They use it, and they call that lineage.

This leads to confusion. Data lineage and graph databases both deal with nodes and relationships, so at first glance, they look like the same problem. They are not. A graph database is a technology built to store and traverse highly connected data. Data lineage is a business capability, built to explain how data moves, transforms, and eventually earns the trust of the people who use it.

Lineage can be drawn as a graph, but that does not make it a graph database problem. The technology user asks how lineage should be stored. The business user asks whether they can trust the KPI on their screen. The value is in impact analysis, trust, transparency, and accountability, and a graph database solves none of those.

Why Lineage Matters More Now Than Ever

Not far back, data flowed through a handful of databases and a few ETL jobs, everything curated and handed over ready-made by technology. Tomorrow, a single analytics dashboard will draw from many platforms at once, cloud warehouses, streaming services, APIs, SaaS applications, and data lakes, threaded together by hundreds of pipelines. A change in one upstream source can silently break dozens of reports downstream. Data lineage is the practical way to see those dependencies, and when a source breaks, it should alert everyone who relies on it. Don’t we get exactly that from the MTA?

GDPR, CCPA, BCBS 239, HIPAA, the European Banking Authority, the EU AI Act, every one of them wants to know where personal data came from, how it changed, and who transformed it. Data governance itself is shifting from documentation to accountability, and lineage is the traceability layer tying business terms to owners, quality and privacy controls, and the physical data beneath them.

AI’s urgent need for provenance is arguably the biggest driver today for the modern stack of data lineage platforms. As organizations deploy AI, these questions turn pressing.

  • What data trained this model, and was its source trustworthy?
  • Did sensitive data get swept in without anyone noticing?
  • What happens to the model when an upstream source changes?

The polish a New York MTA commuter gets from a system that answers their questions, flags trouble, and guides the next step is what earns their trust and reliance. Shouldn’t modern data lineage platforms deliver that same depth for the data behind our AI use cases and self-service analytics?

These four problems, complexity, regulation, AI, and self-service, all reduce to a single question the business keeps asking: Can we trust our data? Data lineage is the gateway through which that trust is earned, and the business value the executive sponsor was paying for all along.

So, where does your program stand, lineage a business leans on, or a diagram that gathers dust?

Ravi Sutar
Founder, 1lessclick® | Obsessions Series

1lesscli>k
It’s a questions about the “why” side of Data Catalog and what makes a Data Catalog a Good Data Catalog.
1lesscli>k
Metadata isn’t just descriptive; it informs where data lives and how it should be used.