Designing a prompt set that actually measures brand visibility in AI answers
A prompt set design only works when it mirrors how buyers actually navigate their decision journey.

Most brand visibility frameworks for AI answers are built on a bad assumption: that you can take your existing SEO keyword list, run it through ChatGPT a few times, and call what comes out a measurement system. You cannot. The prompt is a search query in name only, and treating it like one produces data that feels rigorous while telling you almost nothing useful.
The actual problem is architectural. When a user asks an AI a question, the model does not retrieve a ranked list of URLs. It synthesizes an answer from internalized training patterns and, in some architectures, retrieval-augmented sources. Your brand either appears in that synthesis or it does not, and whether it appears depends less on keyword density than on how the model has learned to frame a given category, use case, or decision context. That framing is what your prompt set needs to surface. Most prompt sets fall far short.
Why Most Prompt Sets Fail Before You Even Run Them
The instinct is to measure brand visibility by asking the obvious questions: "What are the best project management tools?" or "Who are the leading cybersecurity vendors?" These queries are not useless, but they are incomplete in a way that flatters brands more than it informs them.
Why exactly does this happen? Category-level queries are the easiest context for any established brand to surface in. They measure name recognition, not presence across the decision journey. A brand can dominate "top five" lists and still be completely invisible when a buyer asks, "What should I look for when evaluating endpoint detection for a mid-market company?" That second question is closer to real purchase intent. It also represents the moment where AI-mediated influence is quietly compounding, one synthesized answer at a time.
It is also worth considering that AI outputs are not static artifacts. The same model, the same prompt, and the same day can produce different brand mentions across sessions. If your methodology does not account for this stochasticity, you are building a spreadsheet around noise and calling it research — not running a measurement system.
Building the Prompt Architecture
A defensible prompt set is designed around decision stages, not keyword themes. Think of it as the buyer journey expressed in natural language — because that is essentially what it is, a road map drawn in the words real buyers use, and your brand either appears on that map or it doesn't.
Stage One: Problem Articulation
These are queries where the user does not yet have a category in mind. They are describing a symptom. "Our sales team keeps losing deals in the final stage and we don't know why." "We're spending too much on cloud infrastructure and can't figure out where it's going." What you are measuring here is whether the model associates your category, let alone your brand, with this class of problem. If it does not, you have an awareness deficit that no volume of category-level mentions will compensate for.
Stage Two: Category Education
Here the user understands the category but is learning its contours. "What's the difference between a CDP and a CRM?" "How does zero-trust architecture actually work?" Your brand does not necessarily need to appear by name at this stage. But the concepts, terminology, and evaluation criteria your brand champions should. If a competitor's vocabulary is doing the educating, that competitor is constructing the mental model through which buyers will eventually evaluate everyone, including you. That is a structural disadvantage, and it accumulates — like interest on a debt you didn't know you were carrying.
Stage Three: Evaluation and Comparison
This is where most teams pour their energy, and it is the least differentiated place to compete. "Compare Salesforce and HubSpot for enterprise sales teams." Direct comparison prompts matter, but the model's behavior at this stage is substantially shaped by what happened in stages one and two. By the time a user reaches an evaluation query, the AI has already been calibrating their framework. Your stage-three visibility is, in meaningful ways, downstream of your stage-one and stage-two presence. Most teams measuring only here are reading the final chapter and thinking they understand the plot.
Stage Four: Objection and Risk Framing
This stage is routinely overlooked, which is its own commentary on how shallow most brand monitoring exercises are. Buyers ask AI questions like, "What are the risks of moving to a composable architecture?" or "What do companies get wrong when implementing a new ERP?" If your brand surfaces in these queries as a cautionary example, you have a problem your category-level rankings will rarely reveal. If your brand appears as a recommended path through the risk, you have an asset you are almost certainly not tracking.
The Prompt Variables You Actually Need to Control
Once the stage architecture is in place, each prompt needs to vary across at least three dimensions.
Persona specificity. A CFO asking about spend management software and a procurement manager asking about the same topic will receive meaningfully different answers from a well-calibrated model. "What should I consider as a CFO evaluating spend management platforms?" surfaces different brand associations than its generic equivalent. Your prompts should reflect the range of people actually making or influencing the purchase.
Stakes and context. "What's a good tool for social media scheduling?" and "What tool should a Series B company use for social media if they're managing multiple brands across regions?" are structurally different questions. The second introduces constraints that sharpen the model's response and will promote or demote your brand relative to the unconstrained version. Unconstrained prompts produce generic answers; generic answers produce misleading baselines.
Phrasing variation. The same semantic intent, expressed differently, can yield meaningfully different brand mentions. Run multiple phrasings of the same underlying question. If your brand appears consistently across phrasings, that is a robust signal. If it appears only under narrow formulations, that is fragility wearing the costume of presence.
What You Are Actually Measuring
After running a well-constructed prompt set across multiple sessions and models, you are looking at several distinct signals. Brand mention frequency is the obvious one; it is also the least interesting in isolation.
The more diagnostic signals are: at which decision stages does the brand appear, and in what framing? A brand mentioned frequently at the category stage but rarely at the evaluation stage has a recognition problem. A brand surfacing at the evaluation stage in a risk or cautionary context has a reputation problem. A brand that appears only when explicitly named in the prompt has a generative presence problem, meaning the model does not volunteer it unprompted. Each failure mode requires a different response, and a flat mention-count metric will not distinguish between them.
That raises an important question: how do you actually code and analyze the outputs? Manual review is slow but, at the current state of the tooling, still more reliable than automated tagging for nuanced framing signals. The pragmatic approach is automated detection for mention frequency and position, with human review layered in for framing and sentiment on a sampled basis. Tools like Profound, Semrush's AI visibility features, and platforms like Scrunch AI are beginning to build structured measurement workflows around exactly this problem. They are worth evaluating seriously rather than defaulting to ad-hoc prompt testing in a chat window, which is, unfortunately, where every team I've audited currently lives.
The Validity Problem You Cannot Wish Away
But what if the whole enterprise is compromised from the start? One serious objection runs like this: measuring AI brand visibility is inherently unreliable because models update continuously, outputs are probabilistic, and the training data is auditable by no one. This is not a frivolous concern.
It does not, however, make measurement futile. It makes methodological rigor more consequential, not less. The answer is to build a longitudinal tracking system, run the same prompt set on a defined cadence, and measure directional change rather than absolute position — rather than pretending the measurement is stable. A brand gaining mention frequency and positive framing across decision stages over successive measurement periods is doing something right, even if the exact mechanism remains opaque. The signal is relative and temporal. That is a real constraint. It is not a disqualifying one, and treating it as such is a convenient excuse for not doing the work.
The Discomfort Is the Point
Most companies are not building this infrastructure because discovering that a model you cannot directly optimize has formed a clear and possibly unflattering view of your brand is genuinely uncomfortable. Easier to run a few category-level queries, screenshot the mention, and move on.
But what if the discomfort is exactly the information? The brands investing in prompt architecture now, before AI-mediated discovery becomes the default entry point for most buying journeys, are building something that compounds over time. The brands waiting for the trend to become obvious will find themselves measuring catch-up while calling it strategy.
The prompt set is not a reporting artifact. It is a diagnostic. Build it carefully enough, run it honestly enough, and it will tell you what the model thinks about your brand across the entire decision journey, which is increasingly indistinguishable from what prospective customers will encounter when they ask the machine where to start. That is the most important brand audit you run this year, and it will be the most uncomfortable one.


