While a Data Lake can easily turn into what many call a “data swamp.” As Yuval Noah Harari explains in Nexus: A Brief History of Information Networks from the Stone Age to AI, simply collecting more information does not automatically create truth or wisdom. He describes this as the naive belief that more data always leads to more insight or power.
While Data Lake can store large amounts of raw data from many sources and support analytics and machine learning, without strong governance, it can become a disorganized dumping ground where data is hard to find, trust, and use effectively.
I was working with the Enterprise Data Governance team, and they shared the challenges they were facing due to the lack of strong Data Governance for their Enterprise Data Lake. Although the Data Management team successfully built the Data Lake in support of analytics, it soon began to drift into an uncontrolled effort. The Platform Manager had limited visibility of the types of data being ingested into the Data Lake, with no clear ownership oversight, unclear requestors of new datasets, and little understanding of their intended purpose. Access controls were unclear, and concerns about data sensitivity, and classification continued to grow. At the same time, the compliance, regulatory, security, and privacy teams were calling for stronger oversight to ensure the platform remained secure, accountable, and compliant. Upcoming regulatory requirements, including the California Consumer Privacy Act (CCPA), were approaching and required proactive preparation.
In my first meeting with the Enterprise Data Governance team, it became clear that there was a lack of a centralized metadata management platform and formal ownership or accountability for datasets. Managing a Data Inventory is not just about meeting data requirements, but it also addresses business needs, compliance, data security, and privacy. Understanding the data quality and privacy needs is important for business ownership. Equally important is knowing who is using the data and how they are using it. The same approach can also be adapted to support their unstructured data requirements.
1. It’s All About People (Operating Model)
Data Governance starts with an operating model of people, metadata, and operating procedures. Using a RACI matrix, workflows for data access and change requests were designed, embedding Legal and Security teams to ensure compliance, quality, and privacy. Routing approvals to the appropriate stakeholders streamlined the process, so by the time requests reached the Platform Manager, they had already been thoroughly vetted.
The following roles were instrumental in bringing Data Lake assets (datasets) under effective Data Governance. Each role was embedded at the appropriate stages of the process to ensure structured oversight and accountability. True impact is achieved by aligning our objectives with what matters most to the people.
1 | 2 | 3 | 4 | 5 |
Data Provider Data Requester | Data Gov. Manager Data Gov. Admin Security Liaison Legal Counsel | Data Owner Data Trustee | Platform Manager Platform Admin | Data Consumers |
2. Data Processing Without Purpose
During the design phase, even before implementation, the team discovered that 21% of datasets had no assigned owner, and no business unit came forward to take responsibility. This revealed that we were processing, transforming, and storing data in the Data Lake that had no clear purpose. This reflects a common challenge with Data Lakes: a general lack of understanding about the data and its ownership. It turned out that if no one owns a dataset, it just ends up being part of the technology debt we have to manage. To me, it is just common-sense approach.
3. Proactive Privacy and Security Enablement
Often, privacy and security are added only after data is developed. Here, we had the opportunity to embed privacy from the start, allowing the Privacy Officer to define requirements as datasets entered the Data Lake. All customer-related PII fields were flagged, and legal counsel vetted the privacy rules. Similarly, for data access requests, the Security Liaison provided official approval and guidance on proper guardrails. Many Data Governance processes overlook security and privacy, but we integrated them seamlessly, which also supported implementing CCPA policies for California customers.
4. Change Requests with Audit Trail
All access and data change requests were now fully audited and tracked for operational metrics. All relevant parties were consulted and informed throughout the request’s lifecycle. No changes were applied to the Data Lake until all required approvals were obtained, and any rejected or canceled requests remained in the system with enough justification for auditing, ensuring a complete and traceable record of all actions.
5. Aligned Incentives, Accelerated Adoption
A key reason for adopting Data Governance is to simplify processes, ensuring every step provides clear incentives for the stakeholders involved. From the moment a change request was submitted, tasks for review and approval were assigned automatically, with reminders to ensure timely completion.
This level of automation, paired with a streamlined workflow, effectively connected metadata, processes, rules, data decisions, and people. The simplified process encouraged active participation, while the workflow ensured all tasks were assigned and nothing was overlooked. The Data Platform Manager no longer had to track requests manually; there were no more offline discussions or email follow-ups, and responsibility for moving requests forward was shared across the team. The focus on simplicity drove success by making processes both efficient to operate and easy to maintain.
|
|
Gradual removal of redundant datasets (21%) lacking clear business ownership. Best Practice: Regularly audit and decommission unused datasets to maintain a lean and accountable data environment. | Reduced storage costs, minimized data clutter, and improved overall data quality and accountability across the platform. |
Established an upfront validation process involving Governance, Data Owners, Privacy, and Security teams. Best Practice: Engage cross-functional stakeholders early to validate data and workflows, ensuring accountability and compliance. | Increased trust in platform management, strengthened compliance, and motivated the Data Lake Platform Administration team to follow best practices. |
Simplified data analytics with clear links between datasets and defined ownership. Best Practice: Map datasets to their business domain and their owners to maintain clear relationships between them to improve transparency and traceability. | Easier understanding of data, faster issue resolution, and more effective data-driven decision-making for the Data Analytics team. |
A streamlined process that is efficient, engaging, and clearly incentivized. Best Practice: Design workflows that save time, are simple to operate, and thus motivate stakeholder participation. | Increased satisfaction, higher engagement, and more effective collaboration across teams. |