
In August, Google offered $10 million for de-identified internal business data from Spirit Airlines, including employee communications, spreadsheets, calendars, and operating records. Three months earlier, Publicis agreed to acquire LiveRamp at a $2.2 billion enterprise value, describing data co-creation as a way to build more intelligent AI agents.
Those deals put a price on a question marketers are now confronting inside their own organizations: As access to capable foundation models expands, what will make one company’s AI meaningfully better than another’s?
At Marketecture Live Chicago, Abby Roulston, SVP of Business Marketing at Attain, sat down with Crissi Cupak, Head of Product at PMG, to discuss why the answer increasingly depends on the information feeding the model and the systems used to check its conclusions.
Editor’s note: This conversation has been edited and condensed for clarity.
Abby Roulston: Crissi, you sit across brands and technology every day. Are you seeing the conversation move from “Which model should we use?” to “What are we actually feeding it?”
Crissi Cupak: Model choice still matters, and the models will keep improving. It’s becoming less likely to create a durable advantage because competitors can access many of the same foundation models.
Differentiation increasingly comes from the context you provide, the proprietary data you can use, and the workflows built around the model. The question becomes: What does my model know that yours doesn’t?
At PMG, that’s part of the thinking behind Alli. Its usefulness grows because it works across media, audience, creative, planning, business, and performance data. PMG is building Alli Intel to draw from UDA, Business Insights, Creative Insights, and Audience Planner, producing recommendations alongside answers. The model is one part of the system. Its value depends on the information and operating context around it.
Abby Roulston: Owning proprietary data doesn’t automatically create an advantage. What makes a dataset defensible?
Crissi Cupak: A company could have 20 years of unique PDFs stored somewhere and technically own proprietary data with no real moat in it.
The value depends on whether the data is hard to reproduce, permissioned for the intended use, close to real behavior, longitudinal, current, and connectable to other information. Fit matters far more than raw volume. Billions of observations that can’t answer the question in front of you won’t create an advantage.
Longitudinal data is especially valuable because history can’t be manufactured retroactively. If you’ve built a consistent record over time, competitors can’t recreate that asset by buying a new tool next quarter.
Abby Roulston: Are clients still focused on accumulating more data?
Crissi Cupak: The first instinct is often “more.” The better question is whether the data can answer something useful. I’m seeing more clients make that shift. Scale has value when it fits the decision you’re trying to make.
Abby Roulston: As synthetic content proliferates, does real human behavior become more valuable?
Crissi Cupak: Synthetic data is useful, and its importance will keep growing. AI still needs an anchor back to independently observed reality, especially as more online material has been generated by other systems.
That gives actual human behavior a distinct role as validation. A purchase provides different evidence from a search, a click, or a survey response about what someone plans to buy. Each signal can contribute to understanding a consumer. The purchase tells you whether the predicted behavior occurred.
Abby Roulston: The industry already has a blunt term for content that drifts too far from reality: AI slop. Research on model collapse offers a technical version of the same warning. Systems trained recursively on generated output can degrade in ways that become difficult to detect.
For marketers, the practical question is whether an AI system making decisions about customers has a check against real behavior. If you can’t identify the data validating its output, you may be several generations removed from the event you’re trying to understand.
Abby Roulston: How much of what marketers believe they know about a consumer is inferred from indirect signals?
Crissi Cupak: Marketing has always worked with incomplete information because that was what the industry had available. AI is very good at finding relationships among those signals, but better reasoning over a proxy doesn’t turn the proxy into proof.
The risk is false precision. The analysis becomes more sophisticated, so confidence rises even when the underlying information hasn’t improved.
Consider a consumer who reads three articles about running shoes, searches for the best options, and tells a survey they plan to buy a pair. Those are useful signals. They help build a picture of interest and intent. They still don’t confirm that a purchase happened.
Abby Roulston: What should marketers examine before trusting an AI recommendation?
Crissi Cupak: It comes down to validation. Can you see the underlying data? Can you reproduce the result, compare the recommendation with what happened, measure the error, and understand where the model performs poorly?
Generative AI is optimized to produce plausible answers, and plausibility is different from truth. The answer may sound complete even when its support is weak. AI systems should increasingly be judged on concrete results, measured error, and known limitations.
We may have plenty of data in aggregate. What’s scarce is a high-quality record of real human behavior that can show whether a forecast was correct.
Abby Roulston: How important is it to build a loop where an AI prediction eventually meets the real outcome?
Crissi Cupak: Marketing measurement has historically looked backward: Did the campaign work? Who purchased? What was the return on ad spend?
The same information can support a continuous learning system. You observe what’s happening, form a hypothesis, act, compare the result with the expectation, and adjust. Measurement then becomes an input for the next decision, with learning continuing after the campaign ends.
Purchase data is especially useful because it helps close the distance between inferred intent and actual behavior. At PMG, we’re already bringing point-of-sale sales data alongside media performance data into Alli for client use cases. The goal is to make that ingestion increasingly automated and secure so learning can flow back into planning and optimization.
Abby Roulston: Better models and stronger validation can expand what consumer intelligence can do. The system also needs to show its work. Before scaling an AI-driven decision across an organization, leaders should ask: When the model is wrong, how will we know? If nobody has an answer, the organization is relying on a confident guess.
Abby Roulston: In two years, what will be the key learnings we wish we knew now? Any advice for marketers in this room?
Crissi Cupak: Don’t view data as a firehose; focus on the strategic decision-making upfront as to what data makes sense for your specific objectives. Start by inventorying what your company uniquely knows. Look beyond the data warehouse and identify information that competitors can’t easily recreate.
Then verify your rights. Proprietary data won’t help an AI strategy if the company doesn’t have permission to use it for that purpose.
Connect the information that already exists across the business. AI becomes more useful when media, customer, commerce, creative, and business signals can be understood together.
Finally, build systems where assumptions meet actual results. We’ve spent the last few years asking how intelligent models can become. The next question is how good the information will be that we give them to learn from.
Get the latest from The Outcome—expert takes on the industry, emerging trends and exclusive insights into how consumers are actually spending.