Buyer's guideGrowth, Marketing & Sales
Customer data platform buyer's guide: packaged, composable or neither
A customer data platform collects customer data from many sources, resolves it into profiles and sends audiences to the tools that act on them. Whether to buy a packaged platform, assemble one on your data warehouse, or skip it entirely depends on your use cases, your data team and how you handle consent. This guide gives the criteria, the questions to ask vendors and a way to test on your own data.
On this page
- Use cases that justify a CDP, and the cheaper alternatives
- The layers of a composable, warehouse-native CDP
- Packaged CDP, composable CDP or no CDP: how they compare
- Identity resolution: deterministic, probabilistic and who decides
- Consent and privacy: carrying permissions to every channel
- Activation latency, AI personalization inputs and total cost
- RFP questions and a proof of concept on your own data
- Matching the architecture to the company profile
- Questions and answers
- Sources
Use cases that justify a CDP, and the cheaper alternatives
A CDP earns its cost when marketing needs to build audiences from several sources, such as web and app behavior, purchases, support history and CRM fields, and push them to many channels without a data engineer in the loop each time. Typical examples are suppressing existing customers from acquisition ads, triggering lifecycle messages from product events and personalizing a site from recent behavior.
Many teams do not need one. If most customer data already lives in the CRM and the main need is reporting, a data warehouse with a BI tool answers the question. If activation is limited to one email platform fed by the CRM, the existing integration may be enough. A CDP bought for a vague single customer view tends to become another silo.
The layers of a composable, warehouse-native CDP
A packaged CDP bundles all of these layers in one product. A composable approach assembles them around the warehouse your data team already runs.
- Activation channels
Email, ads, site personalization, sales tools and support desks that receive audiences and attributes.
- Reverse ETL and sync
Moves modeled audiences and traits from the warehouse to each channel on a schedule or in near real time.
- Audience builder
A marketer-facing interface for defining segments over the warehouse tables.
- Identity and data models
Rules that merge identifiers into profiles, plus modeled tables for accounts, people and events.
- Cloud data warehouse
The single store for customer data, with access controls, retention and audit.
- Collection and connectors
Event tracking from web and app plus connectors from CRM, billing and support systems.
Packaged CDP, composable CDP or no CDP: how they compare
| Criterion | Packaged CDP | Composable CDP | CRM plus warehouse, no CDP |
|---|---|---|---|
| Where customer data lives | A copy inside the vendor's platform | Your own warehouse | CRM and warehouse |
| Time to first audience | Fast, with prebuilt connectors | Depends on warehouse maturity | Fast for simple CRM-based lists |
| Skills needed | Marketing operations, light data support | Data engineering and analytics engineering | CRM administration and analysts |
| Real-time behavior | Usually strong for web and app events | Varies; streaming adds cost and complexity | Weak |
| Identity resolution | Vendor's rules, sometimes a black box | Your rules, fully inspectable | Manual matching in the CRM |
| Lock-in | High; profiles and logic sit with the vendor | Lower; components can be swapped | Low |
| Cost drivers | Profiles, events or tracked users | Tool licenses plus warehouse compute | Existing licenses and analyst time |
Typical characteristics, not vendor rankings. Several packaged vendors now offer warehouse connections, so check each product against each row.
Identity resolution: deterministic, probabilistic and who decides
Deterministic matching joins records that share an exact identifier, such as a hashed email, a customer ID or a login. It is explainable and rarely wrong, but misses people who browse anonymously or use several addresses. Probabilistic matching infers that records belong to one person from signals such as device, location and behavior. It finds more matches, at the cost of false merges that are hard to spot and harder to undo.
Treat identity rules as governed configuration. Document which identifiers merge profiles, how conflicts are resolved and how a wrong merge is split, and keep B2B account matching separate from person matching. Ask any vendor to show the merge history for a single profile; if it cannot, you will not be able to answer a customer's access request with confidence.
Consent and privacy: carrying permissions to every channel
Under the GDPR, personal data must be processed for specified purposes and limited to what is necessary for them, and consent, where it is the lawful basis, must be as easy to withdraw as to give.1 A CDP therefore needs to store the purpose and basis alongside each attribute and check them at the moment of activation, not just at collection.
In California, the CCPA as amended gives consumers the right to opt out of the sale or sharing of their personal information, and the state's guidance says businesses must treat opt-out preference signals such as Global Privacy Control as valid requests.2 Practical tests for any platform: does a withdrawn consent or opt-out reach every downstream audience, how quickly, and can you prove it? Retention limits and deletion requests should flow the same way.
Activation latency, AI personalization inputs and total cost
Be specific about latency. Abandoned-cart and on-site personalization use cases need event-to-action times measured in seconds or minutes; most lifecycle and advertising audiences work fine with hourly or daily syncs. Paying for streaming everywhere is a common overspend.
Personalization models need consented, fresh, well-defined features, such as recency of key actions, product affinity and predicted value, rather than raw events. Ask whether the platform can serve model outputs back to channels and log which version produced each decision.
Total cost includes far more than the license: warehouse compute for composable setups, implementation, connector maintenance and the people who own the data models. Model at least three years, including the cost of leaving.
RFP questions and a proof of concept on your own data
Run the proof of concept on a real extract with known edge cases, not on vendor demo data.
Matching the architecture to the company profile
Hypothetical profiles to show how the criteria combine; your situation may mix several.
- If
A consumer brand with a small data team, heavy web and app traffic and many marketing channels.
ThenShortlist packaged CDPs with strong real-time web and app collection.
Speed to activation and marketer self-service matter more than control over the identity logic.
- If
A company with a mature warehouse and analytics engineering team already modeling customer data.
ThenGo composable: add an audience builder and reverse ETL on top of the warehouse.
Duplicating clean warehouse data into a second platform adds cost and a second version of the truth.
- If
A B2B company whose customer data lives mostly in the CRM, with a handful of channels.
ThenSkip the CDP for now; fix the CRM data model and connect the warehouse for reporting.
The problem is usually definitions and data quality, which a CDP will not solve.
- If
A regulated firm with strict residency rules and sensitive attributes.
ThenFavor architectures that keep data in your own environment and enforce purpose checks at activation.
Every extra copy of personal data widens the compliance surface.
Questions and answers
What is the difference between a CDP and a CRM?
A CRM is the working system for people who sell to and serve customers: accounts, contacts, opportunities and cases, edited by hand. A CDP collects behavioral and transactional data from many systems, resolves it into profiles and feeds audiences to channels automatically. Many companies need both, with the CRM as the system of record for relationships and the CDP or warehouse holding behavioral data.
Is reverse ETL the same as a composable CDP?
Reverse ETL is one layer of a composable CDP: it syncs modeled data from the warehouse to operational tools. A full composable setup also needs event collection, identity resolution, data models and usually an audience builder for marketers. Some products cover several of those layers, so compare capabilities rather than labels.
Can a customer data platform fix poor data quality?
Only partly. A CDP can standardize formats, deduplicate on rules you define and flag gaps, but it cannot correct records that were entered wrongly or define stages your teams have not agreed on. If the same customer is described differently in three systems because of process problems, fix the process and the source systems first, then let the platform consolidate clean inputs.
Do we need real-time data for personalization?
Only for some use cases. On-site personalization, cart recovery and in-session offers benefit from data that arrives within seconds or minutes. Most email programs, ad audiences and sales alerts work well with hourly or daily updates. List each use case with the latency it truly needs before paying for streaming infrastructure across the board.
Sources
- Regulation (EU) 2016/679 (General Data Protection Regulation) — EUR-Lex · checked 10 October 2026
- California Consumer Privacy Act (CCPA) — State of California Department of Justice · checked 10 October 2026