Vendor Data Transformation Risks
You Sent Customer Records. The Vendor Made Something New From Them.
6 min read · 19 June 2026 · Privacy
A retail company shared customer purchase history with a loyalty analytics vendor to enable personalization recommendations. The vendor's analytics platform ingested the purchase history, normalized it, calculated derived attributes , purchase frequency scores, category affinity indexes, churn propensity metrics , and joined it with their reference product taxonomy. The resulting dataset had dozens of calculated fields that did not exist in the original data. The retail company's DPA specified the lawful basis and handling requirements for customer purchase history. It said nothing about the calculated attributes the vendor derived from that history , which had different sensitivity characteristics, different retention implications, and arguably different regulatory status. When a customer submitted a DSAR requesting all data the retailer held about them, the vendor provided the original purchase history fields plus twenty-three calculated attributes. The retail company's privacy notice described the original purchase history processing. It did not describe the derivation of churn propensity scores or category affinity indexes. The DSAR response included data the customer had not been informed about. The ICO inquiry that followed found inadequate transparency around data derived from customer profiles.
What is the Vendor Data Transformation Risk Problem, Really?
Data transformation is the process by which input data is modified, enriched, derived, or calculated into output data with different characteristics , normalization that standardizes formats, derivation that calculates new attributes from existing ones, enrichment that joins external data with input records, and aggregation that produces summary attributes from individual records. The risk arises when the transformed output has regulatory, sensitivity, or rights implications that differ from the input data, and when the DPA and governance program address the input but not the output.
The derived attribute problem is the most common transformation risk for analytics vendors. When a vendor's platform calculates a churn propensity score from purchase history, that score is derived personal data , it is personal data because it is attributed to an individual, and it derives from the original personal data the vendor was authorized to process. The lawful basis for the original purchase history processing may or may not extend to deriving predictive scores from that history , the derivation may require a different basis, particularly if the scores are used for automated decision-making. The data subject rights obligations for the derived data may differ from those for the source data , a data subject has rights over automated profiling outputs that may be broader than their rights over raw purchase records.
The regulatory classification shift is a specific transformation risk for vendors who apply normalization or categorization to input data. A vendor whose platform normalizes free-text medical history into standardized diagnostic codes has transformed unstructured text into structured health data , potentially shifting the regulatory classification from general personal data to special category health data, with different processing basis requirements, different retention obligations, and different cross-border transfer restrictions. The input was an unstructured text field. The output is a health code. The processing basis that covered the text field may not cover the health code.
- Derived attribute creation without governance , calculated fields derived from customer data without DPA provisions addressing derived attributes
- Regulatory classification shift through transformation , normalization producing data with different regulatory status than the input
- Data subject rights scope for derived data , DSAR and erasure obligations for derived attributes not addressed in governance
- Transparency obligation for derivation , privacy notices that describe input processing but not output derivation
- Derived data retention without governance , calculated attributes subject to different retention requirements than source data but governed by the same schedule
Why this matters
Vendor data transformation matters for TPRM because transformation creates new data , data that did not exist before the vendor processed the input , with characteristics that may differ materially from the input data in regulatory status, sensitivity, and rights obligations. A governance program that covers the input but not the output has a systematic gap for every transformed dataset the vendor creates from customer input.
The DSAR implication is immediately practical. When a customer submits a DSAR, the vendor must provide all personal data they hold about the subject , which includes derived attributes as well as source data. A controller who did not know the vendor was deriving calculated attributes from their customers' records cannot include those attributes in the DSAR response , creating an incomplete response and a transparency failure that regulators will identify.
Where most teams get this wrong
The most consistent failure is treating the data sent to a vendor as equivalent to the data the vendor holds. Analytics vendors by definition transform data , that is the value proposition. What the vendor holds after processing is not what was sent. It is what was sent plus everything the processing derived from it. Governance that covers only what was sent leaves everything derived from it unaddressed.
- Treating sent data as equivalent to held data for governance purposes
- Derived attributes not addressed in DPA
- DSAR scope not extended to derived data
- Privacy notice transparency not covering derivation
- Derived attribute retention not separately governed
What good looks like
Mature data transformation governance programs address derived data as a distinct data category , specifying the lawful basis for derivation, the retention schedule for derived attributes, the data subject rights obligations that apply to derived data, and the transparency obligations that require disclosure of significant derivation in privacy notices.
- Derivation disclosure in DPA , types of derived attributes the vendor creates, lawful basis for derivation
- Derived data inventory , what calculated attributes are created from customer input data
- Data subject rights scope for derivations , DSAR and erasure obligations extending to derived attributes
- Privacy notice transparency , disclosure of significant derivation that creates profileing or automated decision-making
- Derived data retention schedule , separate retention governance for derived attributes
Tooling
Data Lineage , Collibra, Alation, dbt
Data lineage platforms track how data transforms between systems , documenting derived attributes and their source data lineage. For TPRM practitioners, asking whether the vendor maintains lineage documentation for derived attributes created from customer data provides a specific derivation governance question.
Privacy Management , OneTrust, BigID
Privacy management platforms support extension of data inventories to include derived data categories , enabling governance of derived attributes alongside source data. For TPRM practitioners, asking whether the vendor's ROPA includes derived attribute categories provides a specific regulatory documentation question.
Governance challenges
The governance challenge with data transformation is the design-time governance problem. Transformation specifications are designed by data engineers optimizing for analytical value. The governance implications of derived attributes , regulatory classification, rights obligations, retention requirements , require governance team input at design time, not after the derivation pipeline has been running for two years. Connecting transformation design to governance review before derivation begins is the organizational intervention that closes the gap.
- Add derivation disclosure to analytics vendor DPA , types of attributes derived from customer input
- Extend DSAR scope to derived attributes
- Update privacy notices for significant derivation , profiling and automated decision-making transparency
- Require derived data inventory from analytics vendors
- Ask what the vendor creates from your data , not just what they receive
If you are a small team
Ask your analytics vendors one question that surfaces transformation governance immediately: beyond the data fields we send you, what calculated or derived attributes does your platform create from our customer data , and does our DPA address those derived attributes as a distinct data category? The answer tells you what the vendor holds that you did not send and whether your governance covers it.
- Ask what derived attributes the vendor's platform creates from customer input
- Ask whether the DPA addresses derived attributes as a distinct data category
- Ask what data the vendor would provide in a DSAR response , source data plus derived attributes
What to require
Ask directly:
"What calculated or derived attributes does your platform create from the customer data we provide , and are those derived attributes addressed in our DPA as a distinct data category with its own lawful basis and retention schedule?"
"If we received a subject access request for one of our customers whose data you hold, what data would you provide , specifically, would your response include calculated attributes and scores derived from their source data?"
Expect as evidence
- Derived attribute inventory , what is created from customer input data
- DPA derivation provisions , lawful basis and retention for derived attributes
- DSAR scope confirmation including derived attributes
- Privacy notice transparency for significant derivation
A vendor who responds to the derivation question with 'we analyze the data you provide' has confirmed analytics occurs without describing what it produces. Ask specifically what calculated fields exist in their system attributed to your customers that were not in the original data you sent. The original data is what you governed. The derived data is what they made. Both require governance.
How to evidence it
- Derived attribute inventory documentation
- DPA derivation provisions
- DSAR scope extension to derived attributes
- Privacy notice transparency for significant derivation
Key Takeaway
You sent purchase history. The vendor created churn propensity scores, category affinity indexes, and twenty-three other calculated attributes from it. Those attributes are personal data. They were derived from personal data. They are attributed to your customers. They exist in the vendor's system right now, and your DPA says nothing about them because the DPA was written about what you sent, not about what the vendor made from it. The data you sent and the data the vendor holds are related but not the same. Governance for what you sent is governance for part of what the vendor holds. The rest needs its own governance, its own lawful basis, its own retention schedule, and its own transparency disclosure. Ask what they made. Then govern it.
Speak to It™
The term you nodded along to, explained in ninety seconds, so you can speak to it professionally. It is how most readers find these articles.
Join the Association