Consented, connected, omnitraffic data does not become cheaper as AI improves — it becomes more valuable, because AI raises the marginal utility of having every shopper signal on the same consumer identity.
The data AI can’t manufacture
Frontier AI models are commoditizing. GPT, Claude, Gemini, and Llama now perform within points of each other on most public benchmarks; DeepSeek showed a near-frontier model could be built for roughly six million dollars3; Apple chose to license Gemini for about a billion a year rather than build its own.2 In a world where the model is rented, the durable advantage moves down the stack to data the model does not possess.
For brands, agencies, and retailers trying to understand and influence shopper behavior, that means one thing in particular: consented, connected, omnitraffic data — every shopper signal stitched, with permission, to a single consumer identity across every channel, surface, and touchpoint. Location, app usage, web browsing, receipts, demographics, attitudes, survey responses, and increasingly conversations with AI assistants.
The strategic acquisitions of the past five years prove the point — buyers are paying premium multiples specifically to assemble what a small number of first-party panel platforms already hold natively. This paper lays out the argument and the evidence.
Connected signals beat isolated datasets
One dataset tells you what happened. Connected datasets tell you why and what will happen next. A receipt tells you what a shopper bought. A receipt joined to app usage, web browsing, location, AI-assistant conversation, demographics, and stated opinion on the same consumer tells you what they considered, what they rejected, where they were when they decided, what they told ChatGPT they were looking for, and what they will probably buy next month. That difference — from observation to intent, motivation, substitution, and prediction — is the entire game for brands and retailers trying to grow share.
It is also exactly what AI agents need in order to act on a shopper's behalf, and exactly what generative models cannot manufacture. Synthetic data does not substitute for real behavioral linkage; recent academic work (Shumailov et al., Nature 2024)1 has settled the model-collapse question for any model trained recursively on synthetic shopper behavior. Connected omnitraffic data is the input AI cannot synthesize, and the input that makes AI useful at the consumer layer.
Building one panel or one dataset is hard. Building consent-permissioned linkages across panels — every signal joined on the same shopper identity — is exponentially harder. Each new signal added to the identity record multiplies the value of every prior signal, because every joined pair adds a new measurable behavior. This is the textbook data network effect, and it is what makes the assets in this paper structurally unreplicable rather than just expensive to copy.
The distinction that decides every downstream number
Probabilistic matching — fuzzy joins, device-graph inference, lookalike modeling — is what most ad-tech identity solutions use, and what frontier AI is now getting better at automating. It is good enough for some ad targeting and bad for almost everything else. Real-world probabilistic identity resolution lands at roughly 85–95% accuracy4 under ideal conditions, which means the 5–15% of records matched wrong propagate into every downstream measurement, attribution model, and decision. For high-stakes shopper questions — incremental sales lift, true cross-channel attribution, share-of-wallet shifts, post-AI-assistant conversion — probabilistic accuracy is not good enough.
Deterministic data is different. The same identifiable consumer, with explicit consent, contributing receipts, location, app usage, browsing, opinions, and demographic context to a single platform that joins everything to one record. No inference, no probability, no decay as cookies die. It is the only way to deterministically see the full journey — from a ChatGPT question about "best running shoes for flat feet" through a brand site visit, a competitor app open, a physical store walk-in, and a receipt at checkout. That is the use case the acquisitions below are circling, and the one probabilistic methods cannot reliably serve.
“But won’t AI just link the data soon?”
The sharpest objection to this thesis is technical: as models get better at entity resolution and on-the-fly retrieval, joining datasets gets cheaper, and the premium for pre-joined data should compress. It is the right question. The honest answer is that AI does compress part of the moat — but not the parts that matter for shopper behavior. Three constraints are unaffected by AI capability.
Access is contractual, not computational
A retailer does not have ad-exposure data; an ad platform does not have purchase data; a payment network does not have AI-assistant conversation data. No model — however capable — gives Amazon access to Walmart's loyalty file or any third party access to Vizio's ACR feed. Walmart paid $2.3B for Vizio in 2024 not because the join was hard, but because it had no legal right to the data until it owned the company. AI closes the join gap; it does not close the access gap.
Consent is a legal artifact AI cannot synthesize
GDPR, CCPA and CPRA, Apple's ATT framework, and the spreading US state-by-state privacy patchwork increasingly require consented identity linkage, not inferred linkage. Probabilistic matching without permission produces regulatory exposure, not a moat. Consent is paperwork the model cannot generate, signed by a human the model cannot reach.
You cannot link a signal that was never collected
Probabilistic matching gets 85–95% accuracy on records that exist; it gets zero on records never captured. If a shopper never permissioned their location, no AI infers it. If they never told ChatGPT what they were considering, no LLM hallucinates a defensible conversation. The data-capture step is upstream of AI — at the consumer-panel, sensor, or contributory-network layer. That is precisely where the durable moat sits: at the point of collection, with the consumer, under the consent agreement.
The market is voting
If "AI will link it later" were the correct read, sophisticated post-GPT-4 acquirers would not be paying premium multiples in 2024–2026 for already-linked shopper data. They are — Publicis, Walmart, TransUnion, Omnicom, Adobe, and more. These are not naive buyers. They have the same access to AI tooling as everyone else. They are paying the premium because they understand AI compresses the join layer while leaving access, consent, and collection untouched — and those layers are where the value is.
Ten acquisitions that prove the shopper-data thesis
Selected from a larger universe of ~35 data-asset acquisitions. Chosen because each deal is, at its core, a matching deal — the acquirer already had one shopper signal and paid a premium to join it, with consent, to another. The pattern repeats across agency holdcos, retailers, identity providers, retail-media networks, and ad-tech measurement platforms.
| Acquirer / Target | Date | Price | Signal being joined |
|---|---|---|---|
| Publicis / LiveRamp | Nov 2025 | ~$2.5B | Identity graph × agency data stack |
| Walmart / Vizio | Dec 2024 | $2.3B | TV ad exposure (ACR) × purchase |
| Circana / NCS + Nielsen MMM | Q4 2024 | Undisclosed | CPG purchase × ad exposure |
| TransUnion / Neustar | Dec 2021 | $3.1B | Authoritative identity resolution |
| Omnicom / Flywheel | Jan 2024 | $835M | Retail-media performance data |
| Adobe / Semrush | Apr 2026 | $1.9B | AI-assistant brand visibility (GEO/AEO) |
| Experian / Audigent | Dec 2024 | ~$200–250M | Publisher first-party data × identity |
| Mediaocean / Innovid | Feb 2025 | ~$500M | Independent CTV ad-serving / measurement |
| Mastercard / Dynamic Yield | Apr 2022 | Undisclosed | Personalization × transaction network |
| Google / Fitbit | Jan 2021 | $2.1B | Consumer-behavior graph (regulatory case) |
Publicis is buying LiveRamp's identity-resolution and data-collaboration infrastructure: interoperability across 25,000+ publisher sites, hundreds of data and tech partners, and the RampID clean-room layer. CEO Arthur Sadoun framed it as “data co-creation for smarter agents.” The matching value layers LiveRamp's identity graph onto Publicis's existing Epsilon and Profitero stack — an integrated agency-side shopper graph the holdco could not build organically. Replicating 25,000 publisher consent agreements is not a five-year project; it is not feasible.
The canonical screen-to-wallet acquisition. Walmart bought a low-margin TV hardware business primarily for SmartCast OS and its automatic content recognition data across ~18M active accounts. ACR sees what is on the screen across HDMI inputs and streaming apps regardless of source device. Combined with Walmart's first-party purchase data, it gives Walmart Connect closed-loop view-to-purchase attribution at scale. Walmart now owns one of the rarest assets in shopper measurement: deterministically matched ad exposure and purchase on the same household.
The closest analog in this paper to the consumer-panel architecture. Circana acquired the matched CPG purchase-to-exposure dataset NCS spent a decade building by linking Catalina loyalty-card purchase data to ad exposure logs, plus Nielsen's MMM business for cross-channel measurement. The matched panel is the asset — independent CPG measurement linking sales to media without walled-garden dependency is structurally hard to recreate. Direct read-through for any brand that needs deterministic shopper-journey measurement outside Amazon and Meta walls.
The direct precedent for Publicis–LiveRamp. TransUnion paid 5.4x revenue for Neustar's OneID identity-resolution platform — explicitly the identity graph, not the security business (excluded). TransUnion rebuilt TruAudience around OneID; integration reportedly increased marketable phone numbers 25% and marketable IP addresses 54%. The moat is decades of authoritative identifiers — phone provisioning data from Neustar's old NPAC role plus offline-to-online linkages — that you cannot crawl together at any price.
Omnicom bought Flywheel from Ascential for the near-real-time retail-media performance dataset across Amazon, Walmart, Alibaba, and adjacent marketplaces. Flywheel manages tens of billions in product sales and billions in ad spend, giving Omnicom its own first-party view of marketplace pricing, share-of-search, and competitor ad performance. John Wren named all three rationales — data, talent, technology — in that order. The data position is what makes the services business defensible long-term.
Adobe acquired the proprietary discoverability index — a multi-year crawl of how brands appear across search engines and, increasingly, generative engines (GEO and AEO). As consumers shift from googling to asking ChatGPT, Gemini, Perplexity, and Claude for recommendations, the Semrush index is the only systematic measurement of how AI assistants surface a brand. Joined to Adobe's Experience Cloud and Brand Concierge, it lets marketers measure and steer AI-assistant visibility — a category that did not exist 24 months ago, and a critical missing piece in modern shopper journey measurement.
Audigent curates first-party data from a publisher network (SmartPMP, ContextualPMP) plus its Hadron identity layer. Experian cited identity and activation explicitly — adding sell-side first-party data and curated PMPs to its existing demand-side identity graph. The publisher-data relationships and curated audience packages are the asset; replicating publisher consent agreements takes years. A cleaner-side mirror of the LiveRamp acquisition logic.
Combined with Mediaocean's existing Flashtalking, this consolidation produced the only independent omnichannel ad serving + CTV measurement dataset at scale outside the walled gardens. The real asset is the combined ad-serving impression dataset — Innovid serves a huge share of US CTV impressions, and log-level data is what makes independent measurement and outcome attribution possible. With holdco and CTV consolidation accelerating, a non-walled-garden measurement dataset is one of the scarcest assets in shopper marketing.
Included as the instructive negative case. McDonald's bought Dynamic Yield in 2019 for $300M to personalize the drive-thru; divested three years later. Mastercard's twist: fuse the personalization engine with its own transaction-data network. The resulting product, Element, uses “Mastercard's propensity models ... over 112 billion transactions ... more than 15 petabytes of proprietary data” to power merchant personalization. The data asset was not Dynamic Yield's. It was Mastercard's preexisting spend data, newly distributed through a personalization layer. The lesson is the entire thesis of this paper: a personalization tool sitting on top of a brand's data alone is structurally less valuable than the same tool sitting on top of a horizontal, multi-brand shopper-data network. Joining is everything.
Included as the cleanest regulatory recognition that consumer-data graphs are themselves competition-relevant assets — not just incidental to a hardware deal. The European Commission cleared Fitbit only after Google made legally binding commitments — for 10 years, extendable to 20 — not to use Fitbit health, fitness, or location data for advertising in the EEA, to keep Fitbit data siloed, to require explicit user consent for Google apps to see Fitbit data, and to maintain open APIs for third-party wearables. Read between the lines: regulators believe combining a consented consumer-behavior dataset with an ad-targeting graph is itself the competitive concern. Validates the premise that consented, connected omnitraffic data is, in fact, the strategic asset.
Adjacent evidence: AI labs are paying real money for data
AI labs writing recurring checks for data access functionally reprice the underlying asset upward without changing ownership: Reddit at ~$60M/year from Google plus a reported ~$70M from OpenAI ($200M+ total AI-licensing revenue in 2024); News Corp at up to $250M over five years with OpenAI; Stack Overflow, Photobucket, Shutterstock, Taylor & Francis, and Wiley have all signed comparable deals.15 These establish a real, observable market price for proprietary data where none existed in 2022. The corollary for shopper-data holders is immediate: assets carried at near-zero AI-marginal value two years ago are being repriced today.
Implications for brands, retailers, and agencies
1 The connected shopper graph is the scarcest asset class in the market.
Every clean-fit deal in this paper buys or builds a dataset that matches two otherwise-separate shopper worlds — ad exposure to purchase, screen to wallet, retail-media spend to commerce outcomes, AI-assistant query to brand visibility. Matched shopper data is harder to fake than raw data, harder to replicate than scraped data, and far more useful to AI agents than isolated data.
2 The asset has been silently repriced.
Datasets carried at near-zero AI-marginal value in 2022 now have observable market prices set by AI-lab licensing deals and acquisition multiples. Connected omnitraffic shopper assets price at a premium to single-signal assets in every comparable transaction examined. Any brand, retailer, or agency that depends on shopper data should be running a real valuation exercise on what they own, license, and would need to pay to assemble organically.
3 The work is at the collection layer, not the inference layer.
AI compresses the join layer but cannot manufacture access, consent, or collection. The strategic moves that compound — first-party panel growth, consent refresh, new signal types (AI-assistant conversations, in particular), contributory-network expansion — all happen upstream of AI. Waiting for AI to substitute for connected shopper data is a strategy of expecting the model to do what the model cannot do.
4 What to look for in a shopper-data platform now.
The criteria implied by the deals above and by the AI counter-argument:
- Directly-consented consumers rather than inferred panels — every signal contributed knowingly.
- Multiple signal streams collected natively on the same identity, not stitched probabilistically from third-party sources — receipts, location, app usage, web browsing, opinions, demographics, ideally AI-assistant conversations.
- Sufficient panel scale to support brand-level and category-level reads at statistical significance.
- Deterministic identity resolution rather than probabilistic graphs.
- AI-readable architecture so the data is queryable in natural language, not only via SQL or pre-built dashboards.
- Clean read on the full shopper journey — from intent (AI-assistant query, search, social) through consideration (browsing, app sessions, in-store visits) through outcome (purchase, repeat, switch).
Sources & references
Figures and dates in this paper are drawn from public company announcements, regulatory decisions, investor disclosures, and press reports published between 2021 and 2026, plus the peer-reviewed source cited below. Acquisition values reflect reported or estimated figures at announcement; undisclosed deals are noted as such. The acquisition universe analyzed comprises approximately 35 data-asset transactions, of which the ten above were selected as the clearest "matching deal" proof points.
- Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R. & Gal, Y. AI models collapse when trained on recursively generated data. Nature, vol. 631 (2024).
- Apple–Google agreement to license Gemini for Apple Intelligence / Siri (reported ~$1B/year). Press reports, 2025–2026.
- DeepSeek-V3 reported training cost (~$6M pre-training compute; $5.576M). DeepSeek-V3 Technical Report, 2024.
- Probabilistic identity-resolution accuracy (~85–95% under ideal conditions). Industry identity-resolution benchmarks and ad-tech vendor documentation.
- Publicis Groupe to acquire LiveRamp (~$2.5B), announced Nov 2025. LiveRamp investor announcement (SEC Form 8-K).
- Walmart to acquire Vizio ($2.3B), announced Feb 2024, closed Dec 2024. Walmart corporate press release.
- Circana acquisition of NCSolutions and Nielsen's Marketing Mix Modeling business, 2024. Circana announcement.
- TransUnion acquisition of Neustar ($3.1B), closed Dec 2021. TransUnion press release.
- Omnicom acquisition of Flywheel Digital from Ascential ($835M), announced Oct 2023, closed Jan 2024. Omnicom announcement.
- Adobe acquisition of Semrush ($1.9B all-cash), announced Nov 2025, closed Apr 2026. Adobe press release.
- Experian acquisition of Audigent (reported ~$200–250M), Dec 2024. Experian announcement.
- Mediaocean acquisition of Innovid (~$500M enterprise value), announced Nov 2024, closed Feb 2025. Mediaocean announcement.
- Mastercard acquisition of Dynamic Yield (closed Apr 2022) and the Element personalization product built on Mastercard's transaction network; McDonald's 2019 purchase ($300M) and 2022 divestiture ($271M gain on sale). Press reports, 2021–2022.
- Google acquisition of Fitbit ($2.1B), closed Jan 2021, with European Commission commitments. European Commission decision, Case M.9660 (2020).
- AI data-licensing agreements: Reddit–Google and Reddit–OpenAI; News Corp–OpenAI (up to $250M / 5 years); Stack Overflow, Photobucket, Shutterstock, Taylor & Francis, and Wiley. Company filings and press reports, 2023–2024.


