Key Takeaways
- AI visibility scores are modeled samples, not ground truth. Most platforms lack access to complete buyer query data, and AI responses differ substantially from run to run.
- Build your own measurement baseline from first-party data. Sales calls, support tickets, win-and-loss debriefs, and community threads provide authentic buyer questions that can anchor a repeatable query panel.
- If you own the question panel and understand how observations are collected, a platform becomes an instrument you can audit rather than a score you must blindly trust.
Nearly every founder who has purchased an AI visibility platform describes a strikingly similar initial meeting. Three claims arrive in a predictable sequence: your buyers are asking these questions, your brand appears here in the response, and your competitor ranks above you.
The pitch is compelling — I have sat through numerous variations of it. The first time I pressed for details on where the question set originated, the conversation grew notably less specific. That was the moment I began investigating further.
What I uncovered points to a step founders can take before committing to another dashboard: construct the question set from data your business already possesses.
No one has the complete query stream
No major AI discovery platform currently provides a comprehensive query stream comparable to traditional search-query reporting. As a result, the prompt list in your visibility report reflects a model — not a faithful recording of what buyers are actually asking.
Some vendors are candid about this. Otterly documents how it draws on Search Console data, keyword research, and generated brainstorming. Ahrefs publishes its own methodology for expanding related questions. That level of transparency matters because it allows a buyer to evaluate the methodology rather than simply admiring the interface.
The Interactive Advertising Bureau made the broader measurement challenge explicit in August 2026. Its AI visibility guidance reveals that more than twenty companies employ different methodologies, each capable of producing different results for the same brand.
The IAB also draws a critical distinction between directional data and decision-grade data, classifying any measurement program containing fewer than fifty queries as exploratory rather than directional. This is a valuable frame for founders: understand what kind of evidence you are examining before acting on it.
This does not render modeled prompt panels useless. It means they should be priced, governed, and reported as modeled demand — not as a direct feed of buyer behavior.
The question set is only the first source of uncertainty. Even if you resolved it perfectly, the answer itself would still shift from run to run.
Even a flawless prompt list cannot produce a stable rank
The output continues to move. In a 2026 crowdsourced study, 600 volunteers ran identical brand-recommendation prompts through major AI systems close to 3,000 times. The same list of brands surfaced in fewer than one out of every hundred repeated runs.
Separate research covering 693,509 repeat answers found that two responses to the same ChatGPT prompt shared only 21.2% of their cited domains.
A single-run ranking is not dependable measurement under these conditions. Repeated observations across a fixed question set can reveal direction, but a screenshot of one answer cannot confirm whether the result is durable.
This is why I value repeatability, source patterns, and disclosed methodology more than the most impressive-looking score in a demonstration.
The data no vendor can sell you
The most valuable question set in your category may already exist within your own organization. It lives in sales calls, support tickets, win-and-loss debriefs, and community threads. These are genuine buyer questions, articulated in the language buyers actually use — and competitors have no access to your first-party context.
There is an important limitation: first-party questions do not represent the entire market. They reflect the buyers who reached you, not everyone researching the category broadly. Treat them as a protected starting point, then augment them with public category questions, and keep the panel consistent long enough to compare results over time.
The strongest first-party panel is not simply a list of frequently asked questions. It should represent the range of decisions a buyer is attempting to make. Incorporate discovery questions about the category, comparison questions about alternatives, risk questions about implementation or switching, proof questions about results, and commercial questions about cost or timing.
That mix matters because a brand can appear highly visible at the top of the funnel and vanish entirely once the buyer enters evaluation mode. If you only test the questions marketing prefers to answer, you risk building a flattering baseline that misses the moments where revenue is genuinely won or lost. The objective is not more prompts — it is a panel that mirrors the buying journey closely enough to reveal where your evidence grows thin.
Here are four steps to convert that material into a practical baseline.
- Pull real questions: For a quick internal pilot, start with ten recurring buyer questions. If you want a directional category read, expand to at least fifty unique queries covering multiple intent types, since the IAB treats smaller programs as exploratory. Use the buyer’s language, not the wording from your positioning deck.
- Run each question repeatedly across a defined engine panel: Five runs per prompt per engine is the Bullzeye repetition floor for an exploratory pass, as it exposes run-to-run variance without making manual testing unmanageable. Treat this as a methodology choice, not an industry rule. Use a panel you can defend, and report each platform separately.
- Record the reference list, not just the answer: Log every cited or referenced source into a spreadsheet. The answer tells you what appeared in that particular run. The source inventory tells you which evidence environment you can examine — and in some cases, influence.
- Read every source against three tests: Who does it say you serve? What problem does it say you solve? What category does it place you in? Record the differences rather than collapsing them into a simple pass or fail. The reference inventory serves as the bridge between measurement and action.
This is also the point where AI visibility becomes useful beyond marketing. Suppose your company is consistently mentioned for a broad category question but disappears when a buyer asks who is best for a regulated use case, implementation support, or a specific integration. That is not automatically an SEO problem. It may be a proof problem, a positioning problem, a product-marketing problem, or a third-party credibility problem.
The source inventory helps separate these possibilities. Instead of telling leadership that visibility dropped six points, you can identify which buyer question exposed the gap, which sources shaped the answer, and which evidence is missing. That is a far more productive management conversation.
Why the inconsistency log matters
Consistent information across independent sources provides a retrieval system with a clearer evidence environment to draw from. When your website, reviews, press coverage, and leadership profiles each describe different versions of your company, you have an evidence-governance problem long before you have an AI problem.
Most companies I audit are carrying some form of positioning drift within that stack. Address what you can control first: website copy, review profiles, professional profiles, and sales collateral. Then tackle the sources you influence over a longer cycle — customer stories, analyst coverage, and earned media.
The objective is not to replace every platform with a spreadsheet. Automation, historical tracking, and competitive monitoring can still justify their cost. The objective is to stop outsourcing the definition of buyer intent. If you own the question panel and understand how observations are collected, a platform becomes an instrument you can audit rather than a score you must trust.
Owning the panel also transforms the vendor conversation. You can ask a platform to measure against your fixed questions, disclose what changed when a model or methodology is updated, and preserve a baseline you can compare over time. If a provider cannot deliver on that, you know what you are buying: useful monitoring, perhaps — but not a decision system worthy of being treated as ground truth.
That distinction protects both budget and credibility. Founders and marketing leaders do not need perfect certainty from an unstable channel. They need sufficient methodological discipline to recognize when a pattern is emerging, when it is still noise, and what action the evidence genuinely supports.
An AI visibility score is not a ranking. It is a sample. The inconsistency log is one of the evidence conditions you can actively manage — and it often produces a more useful roadmap than chasing a position that may shift on the next run.


