Data Tagging Inconsistencies
Confidential in the Source System. Public in the Analytics Platform. Same Data.
6 min read · 26 July 2026 · Privacy
A global healthcare vendor maintained a data classification program with four tiers. Patient clinical records were Tier 1 , most sensitive. Their data governance platform applied Tier 1 classification to clinical records in the source EHR system, with corresponding access controls, encryption, and audit logging. When those records were replicated to the analytics data warehouse for population health analytics, the replication pipeline did not carry classification labels , the warehouse ingested data without classification metadata. The analytics team, needing to classify the warehouse data, applied their own classification scheme , which used different tier names and was calibrated differently than the governance platform's scheme. Clinical records in the warehouse were classified under the analytics scheme as 'research data' , a category with significantly broader access than Tier 1 in the governance platform's scheme. The same patient clinical records had two different classifications in two different systems, and the classification that governed access in the analytics environment was the weaker one. Every analytics platform user who could access 'research data' could access clinical records that the source system's Tier 1 classification restricted to a handful of authorized clinicians.
What is the Data Tagging Inconsistency Problem, Really?
Data tagging is the application of classification labels , sensitivity tiers, data categories, handling requirements , to data elements across the systems where they exist. Consistent tagging ensures that a data element carries the same classification and corresponding controls regardless of which system it resides in. Inconsistent tagging produces a governance posture where the same data element has different classification labels, different access controls, and different handling requirements in different systems , with the most permissive classification determining what is practically possible with the data.
The label propagation problem is the structural root of tagging inconsistency. When data moves from one system to another , through replication, ETL processes, API feeds, analytics pipelines, or export workflows , the classification metadata that governs the data in the source system may not travel with the data. The receiving system may lack the technical capability to ingest classification metadata, may use a different classification scheme, or may simply not receive the metadata as part of the data transfer. The result is data that exists in the new system without its classification context , requiring either that classification be applied independently in the new system, or that data is processed without any classification at all.
The multi-scheme problem compounds the propagation issue. Organizations with complex data environments frequently have multiple classification schemes operating in parallel , the corporate governance platform's scheme, the analytics team's internal scheme, the cloud storage platform's default classification, and the SaaS application's own data categories. Each scheme was designed for its specific context and makes sense within that context. Across contexts, they produce inconsistent classification of the same data, with no canonical classification that governs the data across all systems. The most permissive scheme that any system applies to a data element is the effective classification for that element from an access control perspective.
- Classification not propagated through ETL and replication , data moved between systems losing classification metadata
- Multi-scheme inconsistency , different classification schemes in different systems producing different classifications of the same data
- No canonical classification authority , absence of a single authoritative classification that governs across all systems
- Reclassification at weaker tier , receiving systems applying independent classification that is less restrictive than source classification
- Classification coverage gaps in secondary systems , analytics platforms, data warehouses, and collaboration tools operating without classification
Why this matters
Data tagging inconsistency matters for TPRM because it means that data protection controls applied to the data's primary system do not necessarily apply to the same data in secondary systems , and secondary systems may have significantly broader access populations. A healthcare vendor whose Tier 1 clinical records are accessible to all analytics platform users because the analytics platform applies a less restrictive classification has provided clinical record access to a population far broader than the Tier 1 controls contemplate , not through a breach, but through a classification inconsistency that was built into the replication pipeline.
The regulatory implication is direct. GDPR's data protection principles apply to personal data regardless of which system it is in and regardless of what classification label a specific system has applied to it. A regulator who asks why clinical records were accessible to analytics platform users with 'research data' access will not accept 'the analytics platform classified them differently' as a compliance defense. The data is what it is. The classification that governs it must reflect that regardless of which system holds it.
Where most teams get this wrong
The most consistent failure is assessing classification in primary systems and assuming consistency across secondary systems. A vendor whose primary data management platform has excellent classification may have analytics platforms, data warehouses, and collaboration tools with entirely different classification schemes or no classification at all. The primary system assessment confirms classification capability. The secondary system gap remains unexamined.
- Assessing classification in primary systems without secondary system coverage
- No label propagation assessment , whether classification metadata travels with data through pipelines
- Multi-scheme inconsistency not assessed , different schemes in different systems
- Most permissive classification effect not evaluated , weakest scheme determining actual access
- No canonical classification authority , absence of cross-system classification governance
What good looks like
Mature classification programs maintain a canonical classification scheme that governs across all systems, implement technical mechanisms for propagating classification labels through data pipelines, and validate classification consistency across systems through periodic scanning.
- Canonical cross-system classification scheme , single authoritative tier structure governing all systems
- Label propagation through pipelines , ETL and replication processes that carry classification metadata
- Consistent classification validation , periodic scanning confirming that data classification matches across source and secondary systems
- Secondary system classification coverage , analytics platforms and data warehouses included in classification program
- Access control alignment to canonical classification , controls in each system calibrated to canonical tier regardless of system-specific naming
Tooling
Data Catalog with Classification , Collibra, Alation, Microsoft Purview
Data catalog platforms provide canonical classification management across multiple systems , maintaining a single classification definition that applies to data regardless of which system holds it. Microsoft Purview specifically provides classification propagation through data pipelines, carrying sensitivity labels from source systems to connected data stores. For TPRM practitioners, asking whether the vendor uses a data catalog with cross-system classification propagation surfaces whether canonical classification governance exists.
Data Quality and Lineage , Monte Carlo, dbt, Great Expectations
Data lineage platforms track how data flows between systems , enabling identification of classification inconsistencies by comparing classification in source systems against classification in downstream systems. For TPRM practitioners, asking whether classification consistency is validated through data lineage tooling provides a specific cross-system governance question.
Governance challenges
The governance challenge with classification consistency is the organizational ownership problem. The primary data governance team owns the classification scheme. The analytics team owns the data warehouse. The infrastructure team owns the replication pipelines. No single team owns the question of whether classification is consistent across all three. Closing the gap requires cross-functional ownership of canonical classification governance that spans all data domains.
- Ask about cross-system classification consistency , whether the same data element has consistent classification across primary and secondary systems
- Ask about label propagation , whether classification metadata travels with data through replication and ETL processes
- Ask about secondary system classification , analytics warehouses and platforms specifically
- Require canonical classification documentation , single authoritative scheme governing all systems
If you are a small team
Ask your data-intensive vendors one question that surfaces classification inconsistency: does the classification tier you apply to sensitive data in your primary system propagate automatically to your analytics platforms and data warehouses when that data is replicated , or does each system classify independently? The answer immediately surfaces whether canonical classification governance exists or whether each system operates on its own classification scheme.
- Ask whether classification labels propagate through replication and ETL pipelines
- Ask whether analytics platforms use the same classification scheme as primary systems
- Ask whether cross-system classification consistency is validated through scanning
What to require
Ask directly:
"Does the classification tier applied to sensitive data in your primary systems propagate automatically when that data is replicated to analytics platforms or data warehouses , or does each system classify independently?"
"What is the canonical classification scheme that governs data across all your systems , and is the same tier structure applied consistently in your primary database, your analytics warehouse, and your collaboration platforms?"
Expect as evidence
- Label propagation mechanism description or gap acknowledgment
- Canonical classification scheme documentation covering all systems
- Cross-system classification consistency validation evidence
A vendor who responds with 'we have a comprehensive classification scheme' should be asked specifically whether that scheme applies consistently in the analytics data warehouse and whether clinical or sensitive records classified at the highest tier in the primary system carry the same classification in the warehouse. The scheme is the specification. The consistency is the governance.
How to evidence it
- Classification consistency assessment records for primary and secondary systems
- Label propagation documentation
- Cross-system validation evidence
- Canonical classification scheme documentation
Key Takeaway
The classification that governs data must follow the data across every system it inhabits. When it does not , when the label changes as the data moves, when each system applies its own scheme independently, when the analytics warehouse reclassifies Tier 1 records as research data because nobody configured the pipeline to carry the label , the classification program that was designed to protect the data protects it only in the systems that apply the intended label. The most permissive classification in any system is the effective classification for all access purposes. The label that changes when the data moves is not a label. It is a suggestion.
Speak to It™
The term you nodded along to, explained in ninety seconds, so you can speak to it professionally. It is how most readers find these articles.
Join the Association