BUSINESS SCENARIO LIBRARY

A collection of representative B2B lead discovery scenarios, showing how AI identifies qualified sales opportunities from real-world business conversations.

SCENARIO 125Marketing tech & creator services

Every Business Unit Wants a Unified Customer View: Which Data Sources Should a CDP Start With?

Around the CDP implementation data integration scenario, this piece walks a marketing technology lead through defining a minimum viable data model when data sources are fragmented and uneven, building an identity resolution strategy, and phasing the integration scope.

Business stage
CDP selection and implementation
Lead quality
★★★★★
Typical buyer
Marketing technology lead
Estimated intent
Very high · implementation window
Illustrative scenario

This is an illustrative scenario designed to explain the product’s judgement logic. It is not a real customer case, testimonial, contract, revenue result, or conversion claim.

HOW TO READ THIS SCENARIO

01Situation

02Signal judgement

03Confidence vs priority

04Human next step

Signals considered

  • each business unit defines customer-data terms differently
  • source-system update cadences vary by an order of magnitude
  • identity keys are inconsistent across systems
  • an existing data team has voiced concern about integration scope

Illustrative scenario. This article explains business-signal judgement and human verification. It does not represent a real customer, conversation, contract, revenue result or conversion claim.

The Situation

You are the marketing technology lead. Three months ago, the CMO declared in a strategy meeting that “we need a unified customer view.” The team began evaluating customer data platforms. On the surface, the products look similar — real-time event collection, identity stitching, audience segmentation, activation to downstream channels. But when you started mapping the internal landscape, the real problem surfaced: the existing data sources are not lined up neatly waiting for you. They are scattered across the CRM, website tracking, transaction databases, customer-service systems, email platforms, and several bespoke operations backends built by individual business units.

The tougher part: each data source has a different owner in a different team. CRM belongs to sales operations. Website tracking belongs to the growth team. Transaction data belongs to finance and ecommerce. Customer-service data belongs to the experience team. And every team defines “customer” differently — CRM uses phone numbers as the primary key, the website uses cookies and device IDs, the transaction system ties order IDs to membership IDs. No single person can say “standardize all sources first, then integrate” — because that means at least two quarters with no visible CDP output.

Your situation: you need a start plan that produces visible results within weeks, without building a data foundation you will have to tear down later.

Why “Ingest All Sources at Once” Is the Most Common Failure Mode

CDP projects rarely fail because the product was wrong. The overwhelming majority fail for one reason: the integration scope was not constrained at the start.

Every new data source introduces semantic conflict, not just data volume. When the CRM says “customer status — active” and the transaction system says “purchased in the last 90 days,” someone must decide which definition is authoritative once both land in the CDP. The number of these conflicts does not grow linearly with each new source — it grows combinatorially, because each new source must be reconciled against every existing source.

The cost of fixing integration mistakes rises with scope, not linearly. If you integrate three sources and discover the identity resolution rule is wrong, the fix is adjusting three mapping rules. If you integrate eight sources before discovering the same problem, the fix is not five more rules — it is tracing how each rule change ripples through every already-merged customer profile.

Data-owner patience is a finite resource. Business-unit data owners are usually cooperative at project kickoff — willing to discuss requirements and share data dictionaries. But if three months pass and they still cannot see what value their data is producing inside the CDP, cooperation decays quickly. Planning a dozen-source integration from day one means the source at the end of the queue waits at least three quarters — by which time its owner may have changed roles.

Evidence to Verify Before You Define the Scope

Before locking the integration scope, complete these six verification steps.

  1. Data-source inventory with accountable owners. List every candidate data source and tag each with the actual managing team and approver. Do not assume “IT can help.” Data-source ingestion requires at least two kinds of authorization: technical access — API credentials and data dictionaries — and business permission — field definitions and usage approval. Without both, the source is a paper candidate and should not enter the initial scope.

  2. Data-quality snapshot. For each candidate source, spot-check the last month of records on three measures: fill rate of key fields (unique identifier, timestamp, core attributes), cross-system value consistency for the same record, and the prevalence of null or default-filled values. You do not need a full data-quality audit, but you need to know roughly where each source sits on the garbage-in-garbage-out spectrum.

  3. Identity-resolution feasibility. Take one week of sample data and attempt the simplest identity-matching rules — phone number, email, membership ID — across systems. Calculate the match rate and catalogue the typical mismatch cases. If the match rate between core systems is very low, you have a strategic question to answer before the project starts: push for unified identifier standards across systems first, or accept that the CDP can only build a unified view for a subset of customers.

  4. Real-time versus batch boundary. Sort every business use case into two buckets: data that must flow in real time — like current-session website behavior for triggering real-time personalization — and data that can arrive in batch — like last month’s transaction summary for monthly RFM segmentation. A real-time pipeline is an order of magnitude more complex and expensive than batch. Do not default to “everything must be real-time.” Most early CDP use cases are adequately served by batch.

  5. Privacy compliance baseline. For each data source, identify whether the data category includes personal sensitive information, whether user consent records exist, and whether cross-border data transfer triggers local regulatory restrictions. Privacy compliance is not a post-project patch — it determines whether certain sources can legally be integrated at all, and if so, what additional user notification and opt-out mechanisms are required.

  6. Internal operations team readiness. After the CDP goes live, who maintains data quality? Who responds to business teams requesting new audience segments? Who monitors whether data pipelines have stalled? If the answer is “we will figure it out later,” the CDP will degrade into an expensive zombie system once the vendor team exits. Before the project starts, you need at minimum a named owner for operational handoff and a committed weekly time budget.

The Human Next Step

With the verification complete, proceed in three stages.

First, define a minimum viable data model. From the candidate sources, select three to five core ones — prioritizing those with acceptable data quality, committed owners, and coverage of the customer-identity-behavior-transaction loop. Use these to define a minimum model: unified customer ID, core event types (page view, add to cart, order, payment), and key customer attributes (registration date, membership tier, last active date). This model does not need to be perfect, but it must go live within your stated timeline and produce usable output.

Second, sign a “data contract” with every source owner. This is not a legal document — it is a mutually confirmed working document with four items: the field list the source will provide with definitions, the update frequency and expected latency, the escalation and fix process for data-quality issues, and the data owner’s acknowledgment of how the data will be used. No source enters the CDP without both signatures on this contract. It cannot prevent every problem, but it prevents the most common one: when data goes wrong, nobody knows where to start looking.

Third, set a hard expansion gate. After the core data model goes live, it must run stably through at least one full business cycle — a month that includes a complete marketing campaign rhythm — before the next batch of sources is evaluated for ingestion. The gate’s purpose: validate the foundation’s data quality and business adoption before adding complexity. If the core model produces visible business problems in that first month — identity merge errors sending mismatched messages to customers — those problems must be fixed before expansion.

What Community Messages Cannot Prove

An informal recommendation such as “this CDP product connects really fast,” “this vendor has a strong implementation team,” or “we had similar data sources and finished in two weeks” — these describe partial experience and subjective impression, not verifiable CDP implementation planning evidence. Informal recommendations cannot confirm:

  • Whether the recommended product’s data model matches your source structure and business logic
  • The vendor’s actual identity-resolution capability — especially given the complexity of your identifier landscape
  • Whether the implementation timeline estimate already accounts for your data governance and compliance approval cycles
  • Whether the recommender’s source count, quality, and complexity are comparable to yours
  • The actual workload and maintenance cost the operations team will face post-launch

Every item above must come from your own source verification and vendor capability validation.


This is an illustrative business scenario demonstrating typical verification and decision sequencing in CDP implementation data-source integration. It references no specific customer, CDP product name, contract value, project data, or outcome claim. Actual decisions should follow internal data architecture, vendor contracts, and applicable regulations.

Frequently asked questions

Which data sources should the CDP ingest first?

Start with three core sources: the CRM system for customer identity and interaction records, the website or mini-program behavioral event stream, and the transaction system for order and payment data. These three cover the minimum closed loop of who, what they did, and what they bought. Other sources — customer-service tickets, email opens, ad impressions — should wait until the core data model has run stably through at least one full business cycle. If you layer complexity onto an unstable foundation, every diagnosis starts with 'which source is wrong this time?'

How do you know data quality is good enough to integrate?

You do not need perfect data to begin. You need two minimum conditions. First, key fields — unique customer identifier, event timestamp, event type — must have sufficient fill rates in the source system to support identity resolution. Second, the data owner must commit to a timeline and named owner for future quality improvement. Without these two, the integrated data will create noise instead of insight, and business teams will lose trust in the CDP before it has a chance to deliver value.