Unstructured data governance is the operational practice of assigning ownership, enforcing data policies, controlling access, and managing the lifecycle of non-tabular data – documents, emails, PDFs, and images – across the enterprise. It is the layer of accountability that sits on top of data discovery and cataloging.
Most organizations stop at cataloging their unstructured data. They run the discovery scans, classify the files, push metadata into Collibra – and consider the job done. It isn’t.
Cataloging creates visibility. Governance creates accountability. These are not the same thing, and confusing one for the other is one of the most expensive mistakes we see in data programs. We’ve watched enterprises invest heavily in discovery and classification tooling, build out a data catalog of thousands of unstructured assets, and then leave it sitting there – correctly labeled, completely ungoverned.
Two tools are most commonly used alongside Collibra to reach the governance layer. Ohalo Data X-Ray handles the security and privacy side: scanning repositories to find sensitive data you did not know was there. Deasy Labs handles the AI readiness side: making unstructured content findable, classified, and usable for downstream AI and analytics. This article addresses both integration paths – but the governance layer itself (ownership, policies, access control, retention, audit trails) is the same regardless of which discovery tool you use.
Key takeaways
- Unstructured data governance is the operational layer that assigns ownership, enforces policies, controls access, and manages the lifecycle of files, emails, and documents. It starts where “conventional” cataloging ends.
- Governing unstructured data is fundamentally different from governing structured data: files don’t have schemas, Data Owner assignments don’t map cleanly to file shares, and retention can’t be set on a column.
- Effective unstructured data governance requires five operational elements: ownership, classification-to-policy mapping, access control, retention enforcement, and audit trails.
- Without governed unstructured data, AI systems pull from ungoverned file repositories – surfacing outdated content, leaking sensitive data, and producing outputs that can’t be traced or trusted.
- Collibra integrates with two complementary tools depending on your primary challenge. Ohalo Data X-Ray is a security and privacy lens: it identifies what sensitive or risky data you have and where it lives. Deasy Labs is an AI readiness lens: it makes unstructured data findable, governed, and usable for AI and analytics pipelines. Both require a defined operating model to deliver value.
- The clearest sign governance isn’t working: your catalog has hundreds of PII findings and nobody has acted on them.
What unstructured data governance actually means (and what it doesn’t)
Unstructured data governance is the operational practice of assigning ownership, enforcing data policies, controlling access, and managing the lifecycle of non-tabular data – documents, emails, PDFs, images, and files – across the enterprise. Unlike cataloging, which creates visibility, governance creates accountability.
That distinction matters more than it sounds. A file with a sensitivity label but no owner, no retention policy, and no access review is not governed. It’s tagged. The classification tells you what the file contains. Governance determines what you do about it.
At a practical level, unstructured data governance answers five questions that cataloging alone cannot:
- Who is responsible for this data?
- What policies apply to it based on its sensitivity?
- Who should and shouldn’t have access to it?
- How long should it be retained, and what triggers deletion?
- Can we prove, in an audit, what happened to this data over time?
If you can’t answer all five for your unstructured assets, you don’t have governance yet. You have a catalog.
Why governing unstructured data is harder than governing structured data
Structured data governance has been maturing for two decades. Relational databases have schemas, which means you know exactly what every column contains before governance work even begins. You can assign owners at the table or domain level. You can set data quality rules on fields. You can define retention based on a timestamp column.
Unstructured data offers none of that.
A PDF looks the same at the system level whether it contains a marketing brochure or a customer’s tax identification number. Classification can’t rely on column names – it depends on reading and interpreting the content itself. And the ownership question becomes genuinely complicated: who owns a SharePoint site? The department that created it? The IT team that manages the server? The data steward whose business domain it falls under?
The differences compound across every governance dimension:
| Dimension | Structured data | Unstructured data |
| Ownership | Assigned at domain or table level | Must be mapped to repository, folder, or content type |
| Classification | Based on schema and field names | Requires content scanning (OCR, NLP, AI) |
| Retention | Set on timestamp columns or partitions | Must be applied at file, folder, or tag level |
| Access control | Role-based on database permissions | Often inherited from folder hierarchy, inconsistent |
| Audit trail | Query logs, lineage graphs | Requires integration with file system and catalog |
We see organizations apply their existing structured data governance model to unstructured data and wonder why it doesn’t work. It doesn’t work because the problem is different. It’s not only the adaptation issue – it’s a ground-up rethink. Unstructured data at enterprise scale means petabytes of files with no owner, no schema, and no inventory. That’s not a governance model that needs tweaking – it’s a dark data problem that needs discovery before governance can even begin.
The 5 operational elements of unstructured data governance
Effective unstructured data governance isn’t a framework in the abstract sense. It’s five concrete operational capabilities. Miss any one of them and the others don’t hold.
1. Data ownership
Every unstructured data asset needs a named business owner. Not IT. Not “the data team.” A specific person or role in the business who is accountable for what that data contains, who accesses it, and how long it lives.
In practice, ownership for unstructured data is typically assigned at the repository or collection level – a SharePoint site, a network drive, a project folder – not at the individual file level (which doesn’t scale). That means your data governance roles design needs to account for how your unstructured data is actually organized, not just how your databases are structured.
2. Classification-to-policy mapping
Unstructured data classification is valuable only when it triggers a consequence. A “PII – High Risk” label that sits in a catalog field without generating a workflow, an access restriction, or a retention rule is theater.
The governance layer maps classification outputs to specific policies: if a file is classified as containing customer PII, it inherits the GDPR-relevant retention window. If it contains financial data subject to DORA, it enters a records management workflow. If it contains credentials or proprietary IP, access is restricted automatically. Classification without this mapping is where most programs stall.
3. Access control governance
Permissions for unstructured data are almost always messier than organizations admit. Files get shared via public links. Folder hierarchies inherit permissions in ways nobody remembers setting. Former employees retain access to SharePoint sites. Contractors have broader read rights than their scope warrants.
Access governance for unstructured data means establishing a baseline of who has access to what, reviewing that baseline on a defined cadence, and ensuring that sensitivity-based restrictions are enforced automatically rather than manually. This is where tools like Collibra, connected to your identity provider and file system, do heavy lifting – but only if someone has defined the rules they’re enforcing.
4. Retention and lifecycle management
GDPR’s right to erasure doesn’t stop at your databases. If a customer requests deletion of their personal data, that request extends to the PDF contracts, the email threads, and the scanned forms that live in your file shares. Most organizations cannot honor that request cleanly because they don’t know where all instances of that data live, and they have no automated deletion mechanism for unstructured sources.
A mature unstructured data retention policy defines how long each class of content is kept, what triggers the review, and what “deleted” actually means when a file exists in three systems simultaneously. This connects directly to GDPR, DORA, and sector-specific regulations like banking records retention requirements. The policy must be enforced through automation – manual deletion schedules don’t survive contact with petabyte-scale file stores.
5. Audit trail and lineage
When your AI model surfaces an output that turns out to be based on a three-year-old internal policy document that was superseded last quarter, can you trace that? When a regulator asks which documents containing customer PII were accessed over the past 12 months and by whom, do you have the answer?
Audit trails for unstructured data require that every significant event – classification, access, modification, retention action – is logged and connected back to the asset record in your catalog. Lineage goes further: if unstructured content feeds AI pipelines or analytics, you need to trace which source documents contributed to which output – and that requires data lineage infrastructure that extends beyond your databases. Without this, your unstructured data governance exists on paper but not in practice.
Unstructured data governance for AI: why it can’t wait
AI systems don’t evaluate the governance status of the data they consume. A retrieval-augmented generation (RAG) system will pull from a SharePoint folder whether or not the files in it are classified, owned, retained appropriately, or current. It doesn’t know the difference between a live policy and a deprecated one. It doesn’t know it just included a document containing personal data that should have been deleted 18 months ago.
The consequence isn’t hypothetical. Ungoverned unstructured data entering AI pipelines creates three distinct risks:
- Accuracy risk: AI systems surface outdated, conflicting, or superseded information as if it were authoritative.
- Compliance risk: Sensitive data – PII, PHI, confidential commercial terms – enters AI outputs without the controls required by GDPR, DORA, or sector regulations.
- Traceability risk: You cannot audit what the AI used or explain why it produced a specific output.
IBM’s 2025 research found that 74% of organizations have only limited or moderate coverage of their AI risks in governance frameworks. For most enterprises, unstructured data is the gap. You can’t fix AI governance without first governing the unstructured data AI systems depend on.
Collibra Unstructured AI addresses this directly – but the governance operating model must exist first. The tool enforces the rules. Someone has to write the rules.
The operating model: managing unstructured data governance in Collibra
Collibra, combined with Ohalo Data X-Ray’s unstructured data discovery, enables most of the automation that mature unstructured data governance requires. Ohalo scans and classifies. Collibra catalogs, assigns ownership, triggers workflows, enforces policies, and maintains audit trails. Together they can govern unstructured data at enterprise scale.
But the technology is only as effective as the operating model behind it. Here’s what that model needs to include:
- Policy definition (CDO / DG Program Lead) – Someone with authority must define the actual rules: what retention period applies to each data class, what sensitivity thresholds trigger access restriction, what constitutes a governed unstructured asset versus a transient working file. This is governance strategy, not technical configuration. In practice however it’s not simply IT vs business – the tension is often between the CDO office, Legal/Compliance, and the business domains.
- Data Ownership (Business Domain Leads) – For each major unstructured data repository – SharePoint sites, network drives, S3 buckets, email archives – a named business owner is registered in Collibra. Getting this name in Collibra is easy, but bear in mind that SharePoint sites and network drives often have no obvious business owner – they were created organically over years. That owner is accountable for their domain’s unstructured data cataloging completeness, classification accuracy reviews, and access decisions. IT manages the infrastructure. Business owns the content.
- Data Steward responsibilities – Stewards handle the ongoing workflow: reviewing Ohalo classification findings that require human judgment, approving or rejecting access requests for sensitive repositories, escalating retention decisions where policy is ambiguous, and maintaining the accuracy of metadata in Collibra. Stewards should not be reviewing thousands of findings per week – if they are, the classification rules or the routing logic needs to be fixed.
- Collibra workflows for unstructured data – they should be triggered by risk findings, not manually initiated. In the European bank engagement we documented in our unstructured data case study, file hierarchies were automatically built from the directory structure and synced into Collibra. Workflows were triggered when X-Ray surfaced new high-risk findings. Data Owners received structured review tasks, not ad hoc email chains. Metadata was continuously refreshed – not set once and left to drift. The result was a catalog that stayed accurate and stewardship that was operationally sustainable.
That’s the difference between a governance program and a governance project.
Unstructured data governance best practices that actually hold up
These are the ones we’ve seen work consistently, as opposed to the ones that look good in a slide deck and collapse within six months.
- Start with your highest-risk repositories, not all of them. The HR file share, the legal contract archive, the finance shared drive – these are where the compliance exposure lives. Govern them first. All-of-enterprise rollouts in year one don’t survive the politics and the scale.
- Define retention policy before you classify, not after. Classification without a retention policy produces labels that go nowhere. Know what you’ll do with a “PII – Medium Risk” finding before you generate ten thousand of them. Classification categories themselves should be designed around actionability – if two classification labels trigger the same downstream treatment, they’re not distinct categories, they’re noise.
- Never assign IT as Data Owner. IT manages the infrastructure. Business units own the content. An IT team assigned as Data Owner for a Legal SharePoint site will not know which documents are contractually binding, which are drafts, or what the business retention requirement is. The wrong owner is worse than no owner. Also, when no natural business owner exists for a repository, that’s a signal the repository itself may be ungoverned legacy content.
- Treat ROT data as a first-order governance priority. Redundant, obsolete, and trivial files inflate storage costs, complicate compliance, and add noise to classification results. Identifying and removing ROT before building out ongoing governance workflows reduces the surface area significantly and produces a quick, visible win.
- Steward workload is a design variable, not a consequence. If stewards have 300 pending issues on Monday morning, governance will fail by Wednesday. Routing logic, automatic resolution thresholds, and escalation paths must be designed so that stewards receive workable volumes of meaningful decisions – not everything that the scan surfaced.
- Review access quarterly for high-risk repositories. Annual access reviews are standard in structured data governance. Unstructured data moves faster and permissions drift faster – think a contractor whose access was never revoked, or a folder shared during a project that was never rescoped. Quarterly is the floor for repositories containing sensitive or regulated content.
Signs your unstructured data governance isn’t working
These are the patterns we encounter most often when organizations ask us to help recover a governance program that’s stalled.
- Your Collibra catalog has sensitivity classifications on thousands of unstructured assets, but nobody has acted on any of the findings in the past 90 days.
- “Data Owner” for your SharePoint environment is listed as IT or a generic team inbox.
- You have no retention policy documented for files, only for database records.
- Your last access review for file shares was more than a year ago – or you’ve never done one.
- Your AI or analytics team is using a file repository that has never been through a governance review.
- A recent audit or pen test surfaced PII in locations where it shouldn’t have been.
- Data Stewards describe their governance work as “reviewing a list” with no clear criteria for when to act.
- You can’t answer “what unstructured data did our AI systems consume last quarter and where did it come from?”
If three or more of these are accurate, the data governance strategy for your unstructured environment needs attention before the next audit surfaces it instead.
When to handle unstructured data governance in-house vs. bring in experts
Most organizations can handle some of this work internally – and should. Here’s an honest breakdown.
You can manage it in-house if:
- Your unstructured data environment is relatively contained (one or two primary repositories)
- You have internal Collibra expertise and a functioning data governance program for structured data
- Your regulatory exposure is limited and well-understood
- You have dedicated data steward capacity with defined, manageable workloads
- You’re willing to move slowly and accept a longer time-to-value
Bring in specialists when:
- You’re operating at scale (hundreds of TB, multiple cloud and on-premise repositories, hybrid environments)
- Your industry has specific compliance requirements – banking, pharma, energy – where getting governance wrong has regulatory consequences
- You have a hard deadline: an audit, a regulatory submission, an AI initiative that leadership has committed to
- Your internal Collibra configuration doesn’t yet support unstructured assets and you’re starting from scratch
- You’ve already attempted governance of unstructured data and it didn’t hold
At Murdio, we’ve built unstructured data governance programs in banking, energy, and life sciences – starting from exactly the point where most organizations get stuck: the moment after the catalog is built and nobody knows what to do next. Our 18 certified Collibra Rangers have built the operating models, configured the workflows, and established the stewardship structures that make governance sustainable rather than ceremonial.
If that’s the situation you’re in, let’s talk about your unstructured data environment.
Frequently Asked Questions
The data cataloging process creates visibility: it tells you what unstructured data you have, where it lives, and what it contains. Governance creates accountability: it determines who is responsible for that data, what policies apply to it, who can access it, and how long it should be retained. Organizations frequently complete the cataloging phase and assume governance is done. It isn’t.
Data ownership for unstructured data is typically assigned at the repository or collection level – a SharePoint site, a network drive, an S3 bucket – rather than at the individual file level. Ownership is assigned to a named business role or individual with subject matter knowledge of the content, not to IT. This assignment is registered and maintained in your data governance platform (such as Collibra).
A retention policy for unstructured data defines how long each class of content is kept (based on its sensitivity and regulatory category), what triggers a deletion or archival workflow, and what “deleted” means when a file may exist across multiple systems. Effective policies are enforced through automation – triggered by classification labels and retention metadata – rather than manual processes.
GDPR applies to all personal data regardless of format, which means obligations like the right to erasure, data minimization, and breach notification extend to documents, emails, and files – not just structured database records. Organizations must be able to locate, classify, and delete personal data in unstructured sources on request, and must be able to demonstrate that data retention policies are enforced. Without governed unstructured data, GDPR compliance is structurally incomplete.
Yes. Collibra data governance capabilities extend to unstructured data assets when integrated with tools like Ohalo Data X-Ray. Ohalo handles discovery and classification; Collibra catalogs the results, assigns ownership, triggers stewardship workflows, enforces policies, and maintains audit trails. The integration enables governance of unstructured data at enterprise scale with a level of automation that manual processes cannot match.
Ohalo Data X-Ray scans unstructured data sources – file shares, SharePoint, cloud storage, email archives – using AI and OCR to detect and classify sensitive content at scale. It feeds classification findings and metadata into Collibra, which then applies governance policies, generates stewardship workflows, and maintains a continuously updated catalog. Data X-Ray handles the intelligence layer; Collibra handles the governance layer. Together they cover what neither does alone.
They answer different questions. Ohalo Data X-Ray is built for the security and privacy challenge: it scans unstructured repositories to identify what sensitive or risky data exists and where, then feeds those findings into Collibra for policy enforcement and remediation workflows. Deasy Labs is built for the AI readiness challenge: it makes unstructured content findable, classified, and usable for analytics and AI pipelines. Both integrate with Collibra. Both are legitimate starting points depending on whether your primary driver is risk reduction or AI enablement – and in mature programs, organizations use both.
The most common failures we see: classification that produces findings with no downstream policy action; ownership assigned to IT rather than business units; retention policies that exist on paper but aren’t enforced technically; stewardship workloads that are unsustainably high because routing logic wasn’t designed; and AI teams consuming file repositories that were never put through a governance review. Most of these failures are design problems, not technology problems.
