Customer data platform (CDP)
A customer data platform is software that collects customer data from many systems, matches it to one person and makes that profile available to marketing, analytics and service tools.
A customer data platform (CDP) is software that creates and maintains one persistent, unified record per customer and makes it available to other systems. David Raab, who named the category in 2013, defined it first as packaged software. The CDP Institute's 2026 definition also accepts CDPs composed on top of a data warehouse, which is the main architecture debate today.
- Origin
- David Raab (named the category in a blog post), 2013
- Level
- 401 · Expert
- Fits
- Scale-up, Enterprise
- Time to apply
- 2 to 3 weeks to scope one use case and its identity rules; implementation depends on the architecture chosen
- What you need
- one customer decision the profile must improve, such as which message goes to whom · an inventory of systems that hold customer data, with the identifiers each one uses · someone from data or engineering and someone who owns consent and privacy
A customer data platform (CDP) is software that gathers customer data from many systems, matches it to individual people and keeps one persistent profile per person that other tools can read. Marketing consultant David Raab named the category in an April 2013 blog post, according to the CDP Institute’s own history. He saw new marketing applications that built a unified customer database from several sources, and Raab Associates published its first report on the group in autumn 2013, profiling eleven systems in three groups: B2B data enhancement, campaign systems and audience management. The CDP Institute followed in 2016 and became part of Customer Data Alliance in 2025.
What does a CDP actually do?
A CDP ingests data from every source, links records that belong to the same person and shares the profile with other systems. It sits between the systems that collect data and the systems that act on it.

Two definitions are worth knowing. Raab’s original was packaged software that builds a persistent, unified customer database accessible to other systems, a wording MarTech quotes. Gartner, as quoted by Salesforce when it announced a Leader placement in the first Gartner Magic Quadrant for CDPs (report dated 14 February 2024), describes software that supports marketing and customer experience use cases by unifying a company’s customer data. The Institute’s version is wider. Its 2026 text says the CDP takes primary responsibility for defining customer identity and record structure over time, and it sorts products into four types: Data CDPs, Analytics CDPs, Campaign CDPs and Delivery CDPs, each including the functions of the one before. A Data CDP, the minimum, gathers data, links it to identities and stores it for external systems; a Delivery CDP also sends the messages by email, web, mobile app or ads.
A CRM stores interactions for people to work in. A tag manager or integration tool passes data along, and the Institute says it lacks the permanent database a CDP needs. A warehouse stores anything but has no customer identity unless someone models it.
How does identity resolution work?
Identity resolution links the identifiers that belong to one person into a single profile. The identifiers are cookie IDs, device IDs, emails, phone numbers and customer IDs. It is the hardest part of a CDP: a wrong link merges two people, a missed link splits one.

There are two approaches. Deterministic matching merges records that share an exact identifier, such as the same email. Probabilistic matching scores how likely two records are the same person from fuzzy signals such as name and address. The statistical idea is old: Ivan Fellegi and Alan Sunter’s 1969 paper framed each comparison as a link, a non-link or a possible link where the evidence is insufficient to decide. Products expose both. Salesforce’s training material describes match rules, such as fuzzy name with a normalised address, and reconciliation rules that pick which value wins by frequency, recency or source. Segment’s documentation lists ID rules and a merge protection feature for non-unique anonymous IDs.
Anonymous traffic is the weak spot. WebKit’s Intelligent Tracking Prevention in Safari deletes script-written storage, including cookies, IndexedDB and LocalStorage, after 7 days without interaction, and caps cookies set on landing pages reached through decorated links at 24 hours. The cookie that links a visit to a later login can disappear. mParticle’s documentation shows a design choice in the other direction: in its profile conversion strategy, a login adds an identifier to the existing profile, so the anonymous history stays with the person.
Hightouch’s product post describes pausing a merge when a profile collects too many emails, a guard against a shared device fusing a household. Identity can only join what was captured with a usable identifier, so start from an event tracking plan.
How does consent fit in?
Consent has to be stored on the profile and obeyed by every destination. Under Article 4(11) of the GDPR, consent is a freely given, specific, informed and unambiguous indication of the person’s wishes, and Article 7(3) lets the person withdraw it at any time. Recital 32 says silence and pre-ticked boxes are not consent, and the Court of Justice of the EU applied that in Planet49 (case C-673/17) to a pre-ticked checkbox. The European Data Protection Board’s Guidelines 05/2020, version 1.1 of 4 May 2020, add granularity: separate purposes need separate consent.
For a CDP that means a consent record per purpose, with time and source, and a withdrawal that propagates. Google’s consent mode does not store choices for you, so your own system must persist and reapply them. The server-side tracking and consent page covers the collection side.
Article 18, part 5 of Russia’s Federal Law 152-FZ bars recording, storing and extracting Russian citizens’ personal data using databases outside Russia, except in the cases in items 2, 3, 4 and 8 of Article 6, part 1. For a team serving Russian users, the cloud region of a CDP or warehouse is a first-order choice, and a lawyer should read the exceptions against the case.
Packaged or warehouse-native: what is the debate?
The debate is whether the profile should live in a vendor’s database or in your own cloud data warehouse, with tools on top. Warehouse-native vendors such as Hightouch argue that packaged CDPs see only part of the data, impose a fixed profile model and lock identity logic inside the vendor. RudderStack describes itself as a warehouse-native CDP that resolves identities in the customer’s warehouse. Snowflake positions a customer 360 on its platform with identity partners. These are vendor claims, and the Hightouch post offers no benchmark. Inside the camp the views differ too. In Kim Davis’s August 2024 MarTech piece, Tasso Argyros of ActionIQ notes that even zero-copy designs can leave data in two places once it is edited, while Barry Padgett of Amperity argues that only reading warehouse data limits what a CDP can create.
David Raab argues that a composable CDP is a component, not a CDP, and prefers the term warehouse-native. In late 2024 he argued that suites win: integration costs are high, a weak component drags down the rest, and warehouses struggle with some real-time work. He expects a hybrid where most data stays in the warehouse and real-time response runs in a linked CDP. He cites an Acxiom survey in which 17% of marketers planned to move to a suite within 12 months and 3% to a composable strategy, while about 80% already ran best-of-breed stacks, and a MessageGears survey of data-quality professionals in which 56% preferred composable architecture and at most 11% a packaged system. The samples differ, so the surveys point in different directions. The Institute’s 2026 definition settles the wording by saying the services can be native or composed from warehouses.
| Packaged CDP | Warehouse-native (composable) | |
|---|---|---|
| Where the profile lives | Vendor’s database | Your warehouse |
| Strength | Fast profile updates, usable by marketers | Reuses existing models, full access to data |
| Main risk | Lock-in, partial data | Engineering effort, batch latency |
Once the profile exists, it feeds decisions such as next best action, and it gives customer journey analytics and unified measurement a person-level spine to join to. A Growth Lab plan starts from the decision the profile must change, then works back to the data.
How to apply Customer data platform (CDP), step by step
- Start from one decision. Write down the first decision the profile must support: suppress ads to people who just bought, trigger a lapsed-customer email, or route a high-value customer to a senior agent. Result: one use case with an owner and a metric.
- List sources and identifiers. For each system (website, app, CRM, billing, support), note which identifiers it holds and which expire; Safari deletes script-written storage after 7 days without interaction. Result: a table of sources by identifier that shows where joins will come from.
- Write the matching rules. Decide which identifiers merge two records, which need a second signal, and which never merge (shared devices, generic emails), and who wins when values conflict. Result: a written rule set tested on a sample first.
- Attach consent to the profile. Store consent per purpose with time, source and wording shown, and make every outbound feed read it, so a GDPR Article 7(3) withdrawal reaches every tool. Result: consent held as data on the profile.
- Choose where the profile lives. Compare packaged, warehouse-native and hybrid setups on update speed, who maintains the pipelines and where the law requires storage. Result: a written architecture decision.
- Connect one activation and measure it. Send the profile to one destination, keep a holdout group and compare. Result: a measured first use case that justifies or ends the investment.
Examples
A payments app removes wasted retargeting
Illustrative. A fintech app sees web visits under a cookie, app sessions under a device ID and accounts under a customer ID, so a user who has just opened an account still sees acquisition ads. After the cookie is matched to the customer ID at login, an 'account opened' attribute excludes that person from the ad audience. The first measurable result is the share of ad spend aimed at existing customers.
A clinic network joins booking and follow-up
Illustrative. A clinic network keeps bookings in one system and email subscribers in another, keyed by email or phone. A linked profile stops a patient who has just visited from receiving a 'book your first visit' offer. Because it joins contact details to visit history, the team limits access, keeps consent per purpose and has a lawyer review the rules.
A retailer with a warehouse already in place
Illustrative. A retailer already models orders and loyalty data in a cloud data warehouse. It writes identity rules as SQL models there and syncs audiences to ad and email tools with a reverse ETL product, and keeps a separate real-time tool for on-site personalisation because the warehouse refreshes in batches. This is the hybrid Raab predicts.
When to use it
Use a CDP when the same customers appear in several systems under different identifiers, several teams need one profile, and a decision (suppression, triggers, personalisation) suffers from a missing match.
When not to use it
Skip it while customer data sits in one system, when the CRM can serve as the profile, or when nobody owns data quality, because a CDP copies the mess from its sources. If the need is reporting rather than activation, a clean warehouse may be enough.
Common mistakes
- Buying a tool before choosing a use case.
- Merging on weak identifiers. A shared tablet or family email can fuse two people, and Hightouch says un-merging in packaged tools can be hard.
- Treating consent as a banner. If the choice never reaches the profile, a withdrawal changes nothing downstream.
- Calling any customer database a CDP. The CDP Institute separates systems that build a persistent profile from integration tools and tag managers that only pass data along.
- Ignoring where data is stored. Article 18 of Russia's 152-FZ restricts databases outside Russia for Russian citizens' personal data.
FAQ
What is a customer data platform?
Software that ingests customer data from many sources, matches records to individuals and keeps a persistent profile other systems can read. The CDP Institute adds that it owns customer identity and record structure over time.
What is the difference between a CDP and a CRM?
A CRM stores sales and service interactions for people to work in. A CDP ingests behaviour, transactions and CRM data, resolves identities across them and serves the result to many tools. Many companies run both, with the CRM as one source.
What is a composable CDP?
A CDP assembled from components on a cloud data warehouse instead of a vendor database. Hightouch and RudderStack use the term or 'warehouse-native'. David Raab calls a composable CDP a component, not a CDP, and prefers warehouse-native.
Do I need a CDP if I already have a data warehouse?
Not necessarily. If your team can model identity in the warehouse and sync audiences to tools, warehouse-native covers many use cases. A packaged or hybrid CDP helps with real-time decisions or when marketers must run it without engineers.
How does consent work in a CDP?
Stored on the profile per purpose, with timestamp and source, and checked by every outbound feed. GDPR consent must be specific and unambiguous and can be withdrawn at any time, so withdrawal must reach every connected tool.
Sources
- CDP Institute, What Is a Customer Data Platform (CDP)? (definition as updated in 2026)
- CDP Institute, The CDP Institute Backstory
- Raab Associates, Report: New Systems Help Marketers Build Better Customer Databases, 9 October 2013
- Salesforce (quoting Gartner), Salesforce Named a Leader in the 2024 Gartner Magic Quadrant for Customer Data Platforms, 21 February 2024. Vendor page
- MarTech, Kim Davis, What the Composability Revolution Means for CDPs, 7 August 2024
- CustomerThink, David Raab, Composable CDP Is Dead, 5 November 2024
- Twilio Segment Documentation, Identity Resolution Overview. Vendor documentation
- mParticle Documentation, IDSync Profile Conversion Strategy. Vendor documentation
- Salesforce Trailhead, Unify Your Data in Data Cloud Core Functionality. Vendor training
- Hightouch Blog, Announcing Identity Resolution. Vendor blog
- Hightouch Blog, Andrew Jesien, Identity Resolution: Why CDPs Fall Short, 15 August 2023. Vendor blog
- RudderStack Documentation, RudderStack Cloud. Vendor documentation
- Snowflake Marketing Solutions, Customer 360. Vendor page
- Ivan Fellegi and Alan Sunter, A Theory for Record Linkage, Journal of the American Statistical Association 64(328), December 1969
- Ahmed Elmagarmid, Panagiotis Ipeirotis and Vassilios Verykios, Duplicate Record Detection: A Survey, IEEE Transactions on Knowledge and Data Engineering 19(1), 2007 (NYU copy)
- WebKit, Tracking Prevention in WebKit
- EUR-Lex, Regulation (EU) 2016/679, General Data Protection Regulation
- European Data Protection Board, Guidelines 05/2020 on Consent under Regulation 2016/679, version 1.1, 4 May 2020
- Court of Justice of the European Union, Planet49 GmbH, Case C-673/17, judgment
- Google for Developers, Manage Consent Settings in Consent Mode
- Consultant.ru, Federal Law No. 152-FZ On Personal Data, Article 18, Duties of the Operator
Last updated Oct 9, 2026


