Every day, millions of consumers open ChatGPT, Claude, or Gemini and do something they used to do on a store shelf, in a search bar, or with a friend: they ask what to buy. They describe a symptom and ask what will fix it. Sometimes they name your brand or a competitor. Other times they do not name it at all and the model recommends it first. How would you know?
Those conversations are a new source of AI brand intelligence, and most brands have no idea what is in theirs. Unlike a survey, nobody handed these consumers a question or a list of answers. They spoke first, unprompted, in their own words. And unlike traditional social listening, there is a second voice in the room, the AI's, actively recommending, comparing, and qualifying brands to the people asking.
The good news: there is a rigorous, repeatable way to pull real signal out of this. The method moves from the broadest view to the most granular, category to brand to competitor to intent, and each stage builds on the one before it. Here is how it works, with real examples from studies we have run.
1. Define the universe of conversations that matters
Resist the urge to go straight to your brand. The first move is to step all the way back, scope the entire population of conversations relevant to your business, then segment them by category: the broadest and most revealing cut you can make.
Say you are a consumer-health company. Isolate the spaces you play in (anti-itch, body pain, allergy, sleep, digestive health) and listen to how people describe each one in their own words. You learn two things at once: what consumers are actually seeking (a recommendation, a comparison, plain reassurance), and what the AI hands back to them. Cluster the keywords and topics inside each category and a second picture appears: what is rising, what is fading, and where the category is quietly moving.
This is where it stops being theory. In one hydration study, the category view rewrote the strategy immediately. The workout, the occasion the entire category is built around, ranked only ninth. The biggest moment was the sick day, at 56%. You cannot see that on a shelf. You can see it in the conversation.
Hydration study / ChatGPT conversation data
2. Move to the brand: who brings you up first?
It is not enough to count how often your brand appears. The sharper question is who introduces it, the consumer or the AI.
Calculate the split. What percentage of the time is your brand surfaced first by the person, versus recommended first by the AI? That is a direct read on your AI presence: whether the model is proactively advocating for you, or merely responding to people who already know your name.
Then split every conversation into user voice and AI voice and run sentiment on each separately. This two-sided view is the whole game, because the AI can be a strong advocate for your brand while quietly attaching a caveat that reshapes how consumers perceive it.
Take Cortizone-10, from a separate OTC study of 33,588 health consumers across 339,000 conversations. Its AI voice runs 68.4% cautionary, far more hesitant than how consumers themselves talk about it. And yet the brand posted a +94.2% net AI lift, among the highest in the study: only 4.6% of consumers asked about it unprompted, but ChatGPT recommended it in 98.8% of anti-itch conversations. Its dominance is almost entirely AI-built, a dependency you would never see by counting mentions alone.
In practice: the Gatorade case
Gatorade was the strongest brand in the category: the most-mentioned, still growing, organically known. On the study's AI-to-user scale, where a ratio under 1.5 signals strong organic pull and a ratio above 4 means a brand is essentially AI-manufactured, Gatorade sat at 1.40. Consumers bring it up themselves, and the AI reinforces rather than manufactures its presence. By the old scoreboard, Gatorade had the recommendation locked up.
In 56% of the times ChatGPT named Gatorade, it attached a “high in sugar” caveat, in the very same breath as the recommendation.
MFour hydration study / AI voice analysis
So the brand had the recommendation, but the model was reliably planting a hesitation about a specific product attribute alongside it. Worse for the whole category: ChatGPT led with “drink more water” in 41% of hydration answers, recommending free and DIY options before it named any brand at all.
The takeaway for the beverage maker was not “are we recommended?” They were. It was “on what terms are we recommended?” That is a direct signal to rethink product and packaging, and it only exists in the conversation data.
3. Layer in the competition
Now run the same favorability and sentiment indexes against competitors, at both the brand and the product level. Compare:
- Total mention volume, yours versus theirs
- The consumer-initiated versus AI-recommended split for each
- Relative favorability across both the user voice and the AI voice
This shows where you actually stand inside AI conversations, which can look nothing like your traditional share of voice. In the hydration set, AI-to-user ratios ranged from genuine organic pull (BodyArmor at 1.23, the strongest in the category) to brands that barely exist in consumers' minds and survive almost entirely on the model's recommendation (Nuun and DripDrop at 26 to 28). Two brands with similar mention counts can have completely different relationships with the consumer: one earned, one AI-manufactured.
4. Go deep on intent
The richest insight appears when your brand shows up in the same conversation thread as a competitor. Those co-occurrences demand intent analysis: why are the two being compared?
- Are consumers comparing ingredients?
- Are they comparing price?
- Are they hunting for a lookalike, same quality at a lower cost?
Mapping these comparative intents tells you exactly where you are vulnerable and where you win. It is the difference between knowing you were compared and knowing what the consumer was actually deciding between.
DANI / prompt guide
Now write the prompts
Every cut in this post starts as a prompt. The guide walks the two-step process, the three building blocks, and a template you can copy straight into MFour Studio.
5. The craft that makes it work: keyword discipline
Everything above depends on one underrated skill: deciding which conversations count. This is where most analyses quietly go wrong, and it deserves real attention.
Remember what makes this data different. In a survey, you write the questions and supply the answers. Here, consumers have already spoken, unprompted, with no dropdown and no controlled vocabulary. The only way to find the conversations that matter is to anticipate every word and phrase a real person might use, and to rule out the look-alikes that mean something else entirely.
Good keyword work is really three distinct jobs:
Include terms
The words that qualify a conversation: the brand name plus every variant and sub-brand (Pepsi, Diet Pepsi, Pepsi Zero Sugar, Pepsi Wild Cherry). A variant you forget to list is a mention you never count.
Exclude terms
The words that disqualify a conversation even when an include term is present. A Pepsi analysis strips out “Pepsi stock,” “PepsiCo earnings,” and “PepsiCo job interview”: financial and career chatter, not the drink.
Disambiguation
For terms only sometimes about your topic, require a second qualifying word in the same message. “Soda” versus “baking soda.” Target the retailer versus “target audience.” “Delivery” of groceries, packages, or a baby.
This is not theoretical. In the hydration study, raw keyword matching pulled a far larger pool, and disciplined cleaning removed 38% of it as false positives: “Propel your career,” a video game called Ultima, and the like. The clean universe of 52,011 is the one you can trust; the raw pull would have inflated every number that followed.
The standard to aim for is clean, not perfect. A few hundred stray conversations in a universe of 100,000+ is acceptable; a few thousand is not. Precision is earned in passes: run the analysis, have the system flag suspected false positives, tighten the logic, rerun.
One more point on rigor: not every analysis is a semantic search. Matching by meaning has its place when you are mapping conversations to a broad topic like travel, provided the source has been chunked appropriately before embedding. Go one level down into intent and brand mentions, though, and keywords are the right tool, because generating vector embeddings for every individual word in a broader conversation is untenable. Quantitative analysis at this scale comes from AI translating your natural-language question into a SQL query that can run hundreds of keywords at once with the necessary inclusion and exclusion logic. Getting that logic right is what separates real insight from noise.
The takeaway
Consumer conversations with AI are a new frontier of brand intelligence. The brands that win will not be the ones that simply ask whether AI recommends them. They will be the ones that systematically measure who brings them up first, in what tone, against which competitors, and with what caveats attached, and then act on it.
The conversations are already happening. The only question is whether you are listening.



