The challenge
A specialized life-sciences division of a global pharmaceutical company was working with large volumes of data generated across multiple systems and sources.
One of the most complex areas involved animal-level data. Information about a single animal could originate from different monitoring systems and use different identifiers, naming conventions and structures. As a result, records that referred to the same animal could appear to represent different entities, while similar concepts could be interpreted differently depending on the source system.
As the organization expanded the scope from one species to another, the problem became more visible. Similar business concepts started to overlap, multiple types of animal identifiers appeared across different databases, and the existing structure was not designed to scale consistently.
The client struggled with:
- Fragmented data across multiple systems: Relevant information was distributed across different monitoring solutions, making it difficult for users to understand what data was available and where to find it.
- Inconsistent interpretation of the same concepts: Different identifiers and data structures could describe the same underlying business concept. Without an additional semantic layer, users had to understand the specifics of individual systems to interpret the data correctly.
- Limited data discoverability: Research and analytics teams needed access to the data, but the knowledge required to navigate the underlying databases was concentrated among a small number of subject matter experts.
- A growing risk of duplicated business definitions: When the governance scope expanded from a single species, similar concepts began to be defined independently. Continuing this approach for additional animal types would make the model increasingly difficult to manage.
- Strict confidentiality requirements: Some of the underlying information was commercially sensitive. The organization needed to make data discoverable without automatically exposing every physical table, column or technical detail to every user.
The goal was therefore not simply to catalog another database.
The client needed a scalable governance model that could connect physical data with consistent business logic and meaning and eventually serve as a controlled one-stop shop for the division’s critical data.
The solution
Murdio supported the client in designing and implementing this use case in Collibra.
A Collibra Ranger-certified Solution Architect worked with the highly motivated team of subject matter experts, translating the business requirements into a governance model that could support both the immediate use case and future expansion.
1. Connecting physical data with business context
The first step was to represent the existing data environment in Collibra and add a business layer on top of the technical structures.
Instead of exposing users only to tables and columns, the solution introduced definitions explaining what individual data elements actually represented.
A business glossary was created and the information was organized into relevant subject areas, such as animal identification, animal behavior and other business contexts.
These concepts were then connected to the corresponding physical data.
This allowed users to move from a business question to the data behind it without first having to understand the structure of individual databases.
2. Introducing a logical layer to manage multiple interpretations
The need for a more scalable model became clear when another animal species was added to the scope.
Although many concepts were similar, they were not always represented in exactly the same way. Animal identification was a good example: several different identifiers could exist across systems, each serving a slightly different purpose.
Creating separate business terms for every technical variation would quickly lead to duplication and an increasingly difficult governance model.
Murdio therefore introduced an additional logical data layer between high-level business concepts and physical source data.
For example:
Business concept
Animal identification
↓
Logical representation
Different types of animal identifiers and their meaning
↓
Physical data
Relevant columns across individual source systems
This separated the meaning of the business concept from the way it happened to be implemented in a specific system.
It also created a model that could be extended to additional animal categories without repeatedly rebuilding the same conceptual structure.
3. Creating a foundation for scalable data discovery
The resulting model gave users a clear path from a business concept to the relevant data.
Instead of asking:
What is the difference between Animal ID and Animal Number and where can I find the identification type I need?
users could start with what they were actually trying to understand and follow the relationships to the appropriate source.
The first implementation covered two animal types, while the structure was intentionally designed so that additional species and datasets could be incorporated later.
The same model also created a foundation for future Data Products and AI-supported discovery, where a user could search for information using a business question rather than having to understand the underlying technical landscape.
4. Making sensitive data discoverable without making everything visible
Discoverability was only one side of the use case.
The division’s datasets included highly sensitive information, and the organization did not want all technical metadata to be visible across the enterprise.
The solution therefore used Collibra domains and view permissions to create different levels of visibility.
Selected business information could be made broadly discoverable, while more detailed logical and technical metadata was restricted to authorized users.
This enabled the organization to share knowledge about its data without automatically providing direct access to the underlying database or exposing sensitive technical details.
For Data Stewards, this was particularly valuable: Collibra became a controlled publishing layer where they could document what data existed, explain its meaning and decide which audiences should be able to see it.
The results
The implementation created a consistent foundation for discovering and interpreting division’s data.
While the client did not provide quantitative adoption or efficiency metrics, feedback from the initial group of analysts and Data Stewards was strongly positive.
The key outcomes included:
- A single discovery point for critical data: Analysts could see what information was available, where it was located and what it represented without repeatedly relying on individual subject matter experts.
- Clearer interpretation of complex identifiers: The logical layer helped users distinguish between multiple identifiers and understand which data should be used for a specific analytical purpose.
- Reduced duplication in the governance model: Instead of recreating similar business concepts for every system or animal category, the client gained a structure that separated business meaning, logical representation and physical implementation.
- Controlled access to sensitive metadata: Collibra enabled the organization to make information discoverable while restricting more sensitive technical details to selected audiences.
- A better publishing model for Data Stewards: Data owners and stewards could document and expose information through Collibra instead of providing users with direct access to underlying databases.
- A scalable foundation for additional use cases: The model was prepared to accommodate additional animal categories and other division’s information without redesigning the governance structure from scratch.
- A foundation for Data Products and AI-enabled discovery: By connecting business concepts with logical and physical data, the organization created the semantic structure required for future Data Products and natural-language access to governed information.
What started as a problem of fragmented animal data ultimately became a broader governance use case: creating a controlled layer where users can understand what data exists, what it means, where it comes from and whether they are allowed to see it.
For a data-rich pharmaceutical organization, this is what turns a technical catalog into a practical data discovery environment.
