Does Your CDP Query the Warehouse Directly or Copy Your Data? A Vendor-by-Vendor Answer (2026)

  • When a CDP vendor says “warehouse-native,” it can mean anything from genuine in-place querying to a full proprietary copy with a Snowflake logo on the sales deck.
  • The testable difference is whether your warehouse compute bill moves when the CDP runs a segment. If it does, the query ran in your warehouse. If it does not, the CDP already pulled a copy.
  • Data replication creates three separate problems: egress fees, a second governance surface you have to control, and a sync lag that makes your “real-time” segments stale by the time they activate.
  • Hightouch, Census, and dbt-native approaches (including GrowthLoop) sit at one end of the spectrum. Segment, Salesforce Data Cloud, and Adobe Real-Time CDP sit at the other. Most vendors occupy messy middle ground.
  • A four-step verification test at the end of this article lets you confirm a vendor’s actual architecture before you sign a contract.

When a customer data platform queries your Snowflake or BigQuery warehouse directly, compute charges appear in your cloud bill and your data never leaves the environment your security team controls. When a CDP copies your data into its own cloud first, you pay egress, accept a second compliance surface, and live with whatever latency the replication pipeline introduces. Most vendors claim the former. The architecture verdict table below shows which ones actually deliver it.


Why “Zero Copy CDP” Does Not Mean the Same Thing Twice

The term zero copy entered the CDP vocabulary around the same time Snowflake’s Data Sharing feature made it technically possible to grant another application read access to your tables without moving bytes across a network. Vendors immediately appropriated the phrase. The problem is that “zero copy” now covers at least three meaningfully different architectures, and a sales team can use the same words to describe all of them.

The first architecture is genuine query pushdown: the CDP sends SQL (or a compiled query plan) to your warehouse, the warehouse runs the compute, and results return to the CDP only when needed for activation. Your data stays in your cloud account. Your IAM policies apply. Your query log shows every execution. This is what “zero copy” originally meant in a data engineering context, borrowed from operating-system design where zero-copy I/O avoids redundant memory writes.

The second architecture is copy-on-connect: the vendor replicates your tables into their managed cloud at setup, then queries their copy. They may still call this zero copy because they are not copying on every query, only once. The sync runs on a schedule, typically every few hours, meaning any segment you build is already working from yesterday’s data.

The third is hybrid pushdown: the CDP pushes some queries to your warehouse for profiling and audience building but pulls a subset of records into its own compute layer for real-time decisioning, enrichment, or identity resolution. This is architecturally honest, but vendors rarely disclose which operations run where.

Distinguishing between these three requires more than reading a product page. It requires understanding what query pushdown actually means technically: the CDP’s query planner translates a segment definition into SQL or a Spark execution plan, ships that plan to the warehouse engine, and the warehouse processes it using your credits, your network, your region. The CDP receives only the result set, not the underlying rows.


What Data Replication Actually Costs You (Beyond the Obvious)

Egress fees are the visible cost. On Google Cloud, data leaving a region costs money per gigabyte, and a large customer table copied nightly adds up faster than most procurement teams model. On AWS, inter-region data transfer runs at standard data transfer rates. But egress is the cheapest part of the replication problem.

The deeper cost is governance duplication. When a CDP copies your warehouse tables into its proprietary store, you now have two authoritative copies of customer PII. Any deletion request under GDPR or CCPA must reach both systems. Any schema change in your warehouse must be propagated to the CDP’s copy before analysts notice a broken sync. Any data retention policy your security team sets must be negotiated with the CDP vendor’s storage layer separately. This is not a theoretical concern; it becomes concrete the first time a customer submits a right-to-erasure request and your legal team asks you to confirm the data is gone.

The third cost is latency-as-trust-erosion. A CDP that syncs your warehouse every four hours will build segments on four-hour-old behavior. For ecommerce cart abandonment flows triggered within minutes of a session, that lag is fatal to relevance. For B2B lead scoring where the buying signal is a product-qualified action, a four-hour delta means your sales team is calling on yesterday’s intent. Understanding how latency compounds across your stack is part of the modern marketing data stack that every RevOps team should audit before adding another layer.


Vendor-by-Vendor Architecture Verdict

The table below reflects each vendor’s publicly documented architecture as of their product documentation, press releases, and technical blog posts. Where a vendor’s architecture is ambiguous or occupies middle ground, that is noted explicitly. This is not a ranking and no vendor paid for placement.

VendorArchitecture ClassWhere Compute RunsDoes Customer Data Leave Your Cloud?Replication Required?Notable Detail
HightouchQuery Pushdown / Reverse ETLYour warehouseOnly activation payloads to destinationsNoSQL runs in your warehouse; Hightouch reads only the rows needed for sync
CensusQuery Pushdown / Reverse ETLYour warehouseOnly activation payloads to destinationsNoSyncs query results directly from warehouse to destinations without intermediate storage
GrowthLoopWarehouse-NativeYour warehouse (BigQuery-native)No (activation payloads only)NoBuilt inside BigQuery; audience computation runs as BigQuery jobs in your project
CohesionWarehouse-NativeYour warehouseNoNoOperates as a Snowflake Native App; no data egress by design
Segment (Twilio)Cloud-Native with Warehouse SyncSegment’s cloud (proprietary)Yes , Segment stores its own profile copyYes (bidirectional)Unify product syncs to Snowflake/BigQuery but also maintains proprietary profile store; data lives in both places
Salesforce Data CloudProprietary Cloud + Federated QuerySalesforce infrastructure (with optional Zero Copy via Data Cloud)Yes by default; Zero Copy Partner Network reduces but does not eliminate replicationYes for most connectorsZero Copy in Salesforce refers to Snowflake Data Sharing or Iceberg reads , compute still runs on Salesforce infrastructure for most operations
Adobe Real-Time CDPProprietary CloudAdobe Experience Platform infrastructureYesYesData ingested into Adobe’s proprietary Profile Store and Data Lake; warehouse integration available but data moves into AEP
RudderStackWarehouse-First / Open-Source OptionYour warehouse (when self-hosted or using Cloud Extract)Depends on deployment; cloud plan routes through RudderStack’s infraPartial , new events flow through RudderStack’s pipeline before landing in your warehouse; this is event routing, not replication of existing warehouse dataRudderStack collects and forwards event streams to your warehouse destination; it does not copy tables that already exist in your warehouse
dbt (with semantic layer)In-Warehouse ComputeYour warehouse entirelyNoNoFunctions as the transformation and audience-definition layer in warehouse-native stacks; surfaces metrics to BI and activation tools. Not a CDP, but included here because it is frequently positioned alongside CDP tooling in warehouse-native architectures.
ActionIQHybrid PushdownMix: pushes query to warehouse, but maintains its own compute for orchestration and identityPartialPartialQuery pushdown for audience segmentation, but identity resolution and some ML features operate in ActionIQ’s cloud
LyticsHybrid (Warehouse Sync + Proprietary)Lytics cloud + your warehouseYesYesCan sync from BigQuery/Snowflake but builds behavioral models in its own layer

The most important column is “Does Customer Data Leave Your Cloud?” , that is the one your information security team will ask about first, and the one that determines whether you need a separate data processing agreement with the CDP vendor covering their storage practices. For teams already doing sophisticated audience work, the shortlist of warehouse-native CDPs worth evaluating covers these vendors with more operational depth.


What Is Query Pushdown in a CDP, Specifically?

Query pushdown means the CDP translates a marketer’s segment definition into a SQL statement (or equivalent), ships that SQL to your warehouse endpoint, and the warehouse engine executes it. The CDP receives back only the output: a list of user IDs, or a summary count, or a sample for preview. The underlying customer rows never transit the CDP’s network.

The practical implication is that your warehouse query log records every segment computation the CDP performs. You can open Snowflake’s Query History, filter by the service account the CDP uses, and see exactly what ran and when. That is a complete audit trail, which is meaningfully different from trusting a vendor’s dashboard to tell you what happened to your data.

Contrast this with a copy-first architecture, where the CDP ingests your tables and runs queries against its own replica. Your warehouse logs show an export job, but nothing about what queries ran afterward. The CDP’s internal query infrastructure is a black box from your security team’s perspective.

A hybrid pushdown vendor like ActionIQ pushes the predicate filtering to your warehouse (so the rough audience computation runs on your credits) but then fetches the matching rows into its own environment for enrichment, identity stitching, or ML scoring. This is a meaningful middle ground: your warehouse handles the heavy compute, but PII does still leave your cloud for certain operations. Honest vendors document exactly which operations trigger data movement. Ask for that documentation before signing.


Does Segment Store a Copy of Your Customer Data?

Yes. Segment’s Unify product, which powers its CDP functionality, maintains a proprietary profile store separate from your warehouse. When you connect Segment to Snowflake or BigQuery, Segment writes event data into your warehouse (via Segment Connections) and can read from your warehouse via Reverse ETL, but the canonical user profile Segment uses for identity resolution and real-time activation lives in Segment’s managed infrastructure.

Segment’s documentation describes this as a “Profile Store” that merges identity across events using deterministic and probabilistic matching. The benefit is that Segment can serve real-time profile lookups without waiting for a warehouse query. The cost is that you now have customer PII in Twilio’s infrastructure, subject to Twilio’s data processing terms, retained on Twilio’s retention schedule, and outside your own deletion controls unless you explicitly configure that through Segment’s Privacy Portal.

For teams evaluating Segment who want warehouse-first architecture instead, the Segment alternatives overview covers tools that handle event collection and identity without the proprietary profile store.


How to Tell If a CDP Is Truly Warehouse-Native: The AboutMartech Architecture Verification Test

Vendor documentation alone is not sufficient verification. This four-step test, which you can run during a proof-of-concept, gives you observable evidence rather than marketing claims. We call it the AboutMartech Architecture Verification Test (AVT) , four checks any marketing or data engineering team can execute before signing a contract.

Step 1: Check Your Warehouse Query Log During a Segment Build

Ask the CDP vendor to walk you through building a test audience during the POC. Before they do, open your warehouse query history in a second window: Snowflake’s Query History tab, BigQuery’s Job History, or Databricks’ Query History. The moment the vendor runs the segment, check whether a new query appears in your log, executed by the CDP’s service account. If a query appears with a recognizable SQL pattern (SELECT user_id FROM events WHERE…), the computation ran in your warehouse. If nothing appears, the vendor is querying a replica.

Step 2: Monitor Your Egress Metrics During Onboarding

Your cloud provider’s billing console shows data transfer volume in near real-time. Note your baseline egress for a normal 24-hour period before the CDP connects. Then connect the CDP and monitor egress for the following 48 hours. A genuine query-pushdown CDP should not materially change your egress figures , only activation payloads (lists of IDs going to an ad platform or email tool) should appear. A replication-first CDP will show an egress spike proportional to the size of the tables it copies. This is observable, not self-reported.

Step 3: Request the Service Account Permission Scope

Ask the vendor’s implementation team to provide the full list of permissions required for their service account in your warehouse. A query-pushdown CDP needs SELECT access on the relevant schemas and the ability to create views or temporary tables for query optimization. A copy-first CDP needs EXPORT or similar bulk data extraction permissions. The permission list tells you the architecture, because you cannot copy data with only read permissions on individual tables.

Step 4: Run a GDPR Deletion Test Before You Go Live

Create a synthetic test user in your warehouse and let the CDP index them. Then delete that user from your warehouse tables and submit a deletion request through the CDP’s privacy controls. After the CDP confirms deletion, query your warehouse to verify the synthetic user is gone there too. Then ask the vendor to provide written confirmation that the user’s profile has been purged from their systems. A vendor with a copy-first architecture will need to purge two separate data stores. A warehouse-native vendor with genuine query pushdown has nothing to purge on their end because they never stored your rows.

If a vendor hesitates on Step 4, or cannot provide written confirmation of deletion from their own systems within a defined SLA, that is the answer to your architecture question. For teams building out their full data infrastructure, understanding how CDPs fit into the difference between ETL, reverse ETL, and CDP architectures makes this test easier to interpret.


Why Does Data Replication Increase Your Warehouse Bill?

The counterintuitive part: replication does not always increase your warehouse compute bill. It often decreases it. The CDP is doing its compute on its own infrastructure, so your warehouse sits idle more of the time. The cost it increases is your egress bill and your total-cost-of-ownership when you factor in the CDP’s own compute pricing.

The warehouse bill increases when you use a hybrid pushdown architecture. The CDP pushes queries to your warehouse (which run on your credits) and then also charges you for its own compute layer on top. You pay twice for some operations. This is worth modeling before a proof-of-concept: ask the vendor to run their standard audience queries during the trial and then check your warehouse credit consumption against the previous equivalent period.

For teams where the warehouse bill is a real constraint, reverse ETL tools are architecturally closer to zero copy than most CDPs. Reverse ETL tools like Hightouch and Census push SQL to your warehouse and ship only the resulting rows to destinations, making your warehouse the system of record without adding a proprietary profile store in the middle.


Warehouse-Native vs Cloud-Native: What the Terms Actually Mean

Cloud-native CDP refers to a platform built to run in a cloud environment, typically the vendor’s own managed cloud. It is a statement about the vendor’s infrastructure approach, not about where your data lives. Segment, Adobe Real-Time CDP, and Salesforce Data Cloud are cloud-native CDPs: they were built on modern cloud infrastructure, they scale elastically, but they run in the vendor’s accounts and manage their own data stores.

Warehouse-native CDP means the application is built to run inside your warehouse, or at minimum to use your warehouse as the primary compute and storage layer. GrowthLoop running as a BigQuery application is warehouse-native. Cohesion operating as a Snowflake Native App is warehouse-native. The critical distinction is where the query optimizer runs: the vendor’s infrastructure or yours.

Some vendors use “warehouse-native” to mean they have a strong warehouse integration. That is not the same thing. A strong integration means they can read from and write to your warehouse. Warehouse-native means the core product cannot function without your warehouse as its compute backbone. You can verify this by asking one simple question during a demo: “What happens to segment computation if you lose access to our warehouse credentials?” A genuinely warehouse-native vendor will say segments cannot be built. A vendor with warehouse integration will say they fall back to their own store.


Single Source of Truth: Which Architectures Actually Preserve It

The “single source of truth” argument for warehouse-native CDPs is real but often stated too simply. The claim is that if the CDP queries your warehouse directly, your warehouse remains the authoritative record for customer data. That is true for storage. It is less true for identity resolution.

Most CDPs, even warehouse-native ones, build their own identity graph: a mapping between device IDs, email addresses, anonymous visitor IDs, and known user IDs. That identity graph may live in the CDP’s infrastructure or in your warehouse depending on the vendor. When it lives in the vendor’s infrastructure, you effectively have two sources of truth: raw events and profiles in your warehouse, and resolved identities in the CDP. Any query that needs to know “which anonymous sessions belong to this known customer” must go through the CDP.

Hightouch and Census handle this by letting you define identity resolution logic yourself in SQL and store the resulting identity graph in your own warehouse tables. GrowthLoop similarly writes its computed audiences back to BigQuery as native tables you own. These approaches preserve the warehouse as the genuine single source of truth, including for identity. Compare this to Segment’s Unify, where identity resolution is a proprietary process producing a profile that lives in Segment’s infrastructure and is not directly queryable from your warehouse without additional integration work.


Frequently Asked Questions

What does “zero copy” mean in the context of a customer data platform?

In CDP architecture, zero copy means the platform queries your data where it already lives, typically a cloud data warehouse like Snowflake or BigQuery, without replicating that data into the vendor’s own storage. The vendor’s application sends queries to your warehouse, your warehouse executes them using your compute credits, and only the resulting output (not the underlying rows) travels back. A true zero copy CDP leaves no persistent copy of your customer data in the vendor’s infrastructure.

How do I tell if a CDP is copying my data or querying it in place?

Run the AboutMartech Architecture Verification Test: monitor your warehouse query log during a segment build to see if the vendor’s service account appears there; check your cloud egress metrics for an unusual spike during onboarding; review the permission scope the vendor’s service account requires (export permissions indicate replication); and run a GDPR deletion test to see whether the vendor can confirm purging their own copy. These four steps give you observable evidence, not vendor self-reporting.

Does Segment copy my data into its own cloud?

Yes. Segment’s Unify product maintains a proprietary Profile Store in Twilio’s infrastructure that is separate from any warehouse you connect. Segment writes events to your warehouse and can read from it via reverse ETL, but the identity-resolved customer profile used for real-time activation lives in Segment’s managed environment. This means customer PII is subject to Twilio’s data processing terms and stored outside your own deletion controls unless you configure Segment’s Privacy Portal explicitly.

What is query pushdown in a CDP?

Query pushdown means the CDP translates a segment or audience definition into SQL (or a compatible query plan), sends that SQL to your warehouse, and your warehouse engine executes the computation. The CDP receives only the result set. Your warehouse query log records every execution. The opposite of query pushdown is the vendor running queries against a replica of your data in their own infrastructure, which leaves no trace in your warehouse logs and creates a second governance surface outside your control.

Why does data replication create compliance risk?

When a CDP copies your warehouse tables into its own storage, you now have two systems holding customer PII. Any deletion request under GDPR, CCPA, or similar regulations must reach both systems. Any schema change must propagate to the replica. Any data retention policy must be negotiated with the vendor’s storage layer independently of your own warehouse policies. If the vendor is breached, your customer data is exposed even if your own warehouse is secure. A single data processing agreement may not cover all the vendor’s storage locations.

What is the difference between warehouse-native and cloud-native CDPs?

Cloud-native describes the vendor’s infrastructure approach: the platform runs in a modern cloud environment, typically the vendor’s own managed cloud. Warehouse-native describes where your data lives and where computation runs: the CDP uses your cloud data warehouse as its primary compute and storage layer. A cloud-native CDP (Segment, Adobe Real-Time CDP) runs in the vendor’s infrastructure. A warehouse-native CDP (GrowthLoop, Cohesion) runs inside your warehouse environment. The terms are not interchangeable.

Which CDPs do not require data replication?

Hightouch and Census operate as reverse ETL tools that query your warehouse without maintaining a proprietary copy. GrowthLoop runs natively inside BigQuery. Cohesion operates as a Snowflake Native App. These are the architectures with the clearest zero-copy credentials based on publicly documented technical approaches. ActionIQ offers partial query pushdown but still performs some operations in its own infrastructure for identity resolution and ML features. Any vendor claiming zero copy should be verified using the four-step test described in this article.

How does data replication affect real-time segmentation?

A replication-first CDP syncs your warehouse tables on a schedule, commonly every one to several hours depending on the vendor and plan tier. Any segment built in the CDP reflects the state of that replica at the last sync, not the current state of your warehouse. For use cases where recency matters, such as cart abandonment flows, intent scoring on recent product page views, or time-sensitive B2B triggers, this lag means the CDP is working from stale data. A query-pushdown CDP that runs directly against your warehouse returns results as current as your most recent event ingestion into the warehouse.


What the Architecture Question Is Really About

Most buyers frame this as a cost question. Egress fees, warehouse credits, CDP licensing , all real, all worth modeling. But the architecture question is fundamentally a control question. A CDP that stores your data stores your customer relationships in its own system. The vendor’s pricing decisions, their data breach exposure, their acquisition by a larger company: all of these now directly affect your ability to access and govern your own customer data.

A warehouse-native or genuine query-pushdown architecture returns that control to your team. Your warehouse IAM policies govern access. Your retention schedules apply. Your query logs provide the audit trail. When the CDP vendor raises prices, you can switch tools without a data migration, because the data never left your warehouse. That optionality has real dollar value, even if it does not appear on a vendor’s pricing comparison table.

The distinction between querying in place and replicating is not subtle once you know what to look for. The four-step test in this article gives you the verification mechanism. Run it during every CDP proof-of-concept, not just when a vendor makes an explicit zero-copy claim. The vendors with genuinely warehouse-native architectures will pass it easily and welcome the question. The ones with replication-first architectures dressed in warehouse-native language will hesitate on Step 3 and stall on Step 4. That hesitation is your answer.

Benjamin Parker
Benjamin Parker

Benjamin Parker covers martech stacks at About Martech, with a focus on AI search tooling, analytics platforms, and the data layer underneath both. He writes the category breakdowns and vendor comparisons marketers reach for when they are actively choosing a tool. His work spans GEO and AEO software, product analytics, reporting dashboards, warehouse-native CDPs, and RevOps routing.

Articles: 26