Skip to content
8 min read

Can you measure brand visibility in AI search?

Measure brand visibility in AI search with 20 prompts, a repeatable scoring sheet, competitor comparisons, and a plan for fixing gaps.

Can you measure brand visibility in AI search?
Table of Contents

Brand visibility in AI search can be measured, but only if you stop treating one flattering answer as evidence. A useful audit tests whether answer engines recall your brand, recommend it for the right jobs, describe it accurately, and cite sources that a buyer can inspect. Those are separate outcomes. Combining them hides the reason a competitor keeps appearing while you do not.

I use a fixed set of 20 prompts, record every result in a scoring sheet, and rerun the test under controlled conditions. The exercise takes a few focused hours for one market. It will not produce an official share-of-voice number because no answer engine exposes the full population of prompts or users. It will show where your brand enters the buying conversation, where it drops out, and which gap deserves work first.

AI visibility is four different tests

A brand is visible only when an answer helps the user place it in a market and make a decision. A bare mention in a long list is weaker than a recommendation with a correct reason, and a recommendation without supporting evidence is less defensible than one with a citation. I score four observable layers: recall, fit, accuracy, and evidence.

Recall asks whether the system names the brand without being prompted. Fit asks whether it connects the brand to the problem, buyer, or use case you actually serve. Accuracy checks whether material claims such as category, audience, geography, pricing model, or capabilities agree with your current public information. Evidence asks whether the answer cites your site or another credible source that supports the claim.

This distinction catches a common reporting error. A team sees its name in six answers and declares success. Reading the answers shows that four mentions came only after the prompt named the company, two used an outdated category, and none appeared in a recommendation. The brand has recognition, but it has little influence on an unprompted shortlist.

Classic search rank and AI answer visibility also differ. A page can rank for a query without the answer naming its publisher. An answer can mention a company after gathering evidence through several rewritten or related searches. OpenAI's ChatGPT Search documentation says the product may rewrite a user's question into one or more targeted queries. Google Search Central describes a similar query fan-out process for AI Overviews and AI Mode. That is why checking your rank for the literal audit prompt does not explain the answer by itself.

Do not turn the four layers into one grand metric too early. A total is useful for comparing runs, but the layer scores tell you what to repair. Low recall points toward weak category association. Good recall with poor accuracy points toward inconsistent or stale public facts. Recommendations without citations call for stronger evidence, not more repetitions of the brand name.

Set the test before you see the answers

A credible audit fixes its scope in advance, because changing prompts after seeing weak results turns measurement into theater. Write down the market, buyer, geography, language, product category, three to five competitors, answer engines, and test date. Use the same scope for every row.

Choose competitors that buyers truly compare. Include the obvious category leader, one company near your size, and any substitute that solves the job differently. A startup selling incident response software may compete with another platform, an observability suite, and an internal process built around existing tools. If you only enter companies that resemble your homepage, the audit misses the alternatives an answer engine may reasonably suggest.

Run the prompts on at least two answer engines that your buyers are likely to use. Keep product names and versions in the run notes because interfaces and retrieval behavior change. Use a fresh conversation for every prompt so earlier answers do not seed later ones. Stay logged in or logged out consistently, keep location settings stable, and do not regenerate an answer until you get one you prefer. The first completed answer is the observation.

Personalization is a source of variance, not a reason to abandon the test. If the engine uses memory, account history, or precise location, either disable those features or label the run as personalized. For a regional company, repeat the same set from each material geography using legitimate local testers. Do not pretend a VPN alone reproduces language, account history, local inventory, and user intent.

Record the full answer, visible citations, and any follow-up query the interface reveals. Screenshots help with later disputes, but text is easier to compare. Save enough context to tell whether a citation supports the sentence beside it. A source attached to a paragraph may establish one fact while leaving the recommendation unsupported.

You are building a repeatable sample, not impersonating a statistically representative panel. The sample becomes useful when it stays fixed across competitors and across time. Its limits should remain visible in the report.

Give every run an ID and keep an untouched copy of its inputs. Record the prompt template version, substitutions, engine, mode, account state, operator, local time, and geography. If an engine shows a source panel or search activity, capture that too. This audit log stops a later team from comparing two numbers produced under different conditions and calling the difference progress.

Decide how you will handle unavailable facts before scoring. If a private-company price is not published, an answer that says the price is unknown may be accurate; an invented number is not. If your own site contradicts itself, mark the answer against the current approved fact and log the public conflict as a remediation item. The engine cannot reliably resolve a disagreement that the company created.

Keep raw observations separate from analyst comments. The answer, citation text, and retrieval state belong in locked source columns. Scores and notes belong in review columns. That separation lets another reviewer revisit a judgment without reconstructing the session, and it prevents a polished summary from replacing inconvenient evidence.

Define exclusions as carefully as inclusions. Skip prompts that require private account data, regulated advice beyond the engine intended use, or a geography where the product is unavailable. Document the exclusion once. Quietly dropping a question after a poor answer damages the benchmark more than reporting the miss.

These 20 prompts expose the buying journey

The prompt set should move from category discovery to a named-brand challenge. Replace the bracketed terms once, then preserve the exact wording for every run. Do not stuff your company into the unprompted questions. Its absence is part of the result.

  1. What are the best [category] options for [buyer type]?
  2. Which [category] companies should a [company size] business consider?
  3. What tools can solve [specific job] for a team with [constraint]?
  4. What are good alternatives to [category leader]?
  5. Compare the leading approaches to [specific job].
  6. Which [category] option is best for [industry] and why?
  7. What should I shortlist for [use case] in [geography]?
  8. Which vendors support [must-have requirement]?
  9. What is the safest choice for [high-risk use case]?
  10. Which option is easiest for a small team to adopt?
  11. Which [category] products work well as a company grows?
  12. What are the tradeoffs between [competitor A] and [competitor B]?
  13. Is [your brand] a good fit for [buyer type]?
  14. What does [your brand] do?
  15. What are the strengths and weaknesses of [your brand]?
  16. How does [your brand] compare with [competitor A]?
  17. What does [your brand] cost and who is it for?
  18. Is [your brand] credible for [specific job]?
  19. What are the main complaints about [your brand]?
  20. Should I choose [your brand] or [competitor B] for [constraint]?

The first 12 prompts test discovery without giving your brand a free pass. Prompts 13 through 20 test understanding after the user supplies the name. Keep both groups. If you test only named prompts, a system can summarize your homepage and look competent while never recalling you during category discovery. If you test only unprompted prompts, you cannot see whether the engine misunderstands the company once a buyer asks directly.

Adapt the nouns, not the intent. A local services firm might replace "products" with "firms" and add a city. A developer tool may make a required integration the constraint. Preserve the sequence of discovery, comparison, risk, objections, and decision. These stages reveal more than 20 variations of "best vendor."

Some prompts will not trigger web retrieval or citations. Keep those answers. Retrieval behavior is part of the observation, and forcing search on every query can produce a test that does not resemble how buyers use the product. Note whether search occurred, then compare like with like on the next run.

The scoring sheet keeps opinions out of the score

Score each prompt against explicit rules, then store the details needed to challenge the score. A spreadsheet with one row per prompt and one block per answer engine is enough. Use this header as the working artifact:

<table><thead><tr><th>Field</th><th>Allowed value</th><th>What to record</th></tr></thead><tbody><tr><td>Prompt ID</td><td>1 to 20</td><td>Fixed prompt number</td></tr><tr><td>Brand mentioned</td><td>0 or 1</td><td>Unprompted or prompted</td></tr><tr><td>Position</td><td>0 to 3</td><td>3 first, 2 top three, 1 later, 0 absent</td></tr><tr><td>Fit</td><td>0 to 2</td><td>2 correct fit, 1 partial, 0 wrong or absent</td></tr><tr><td>Accuracy</td><td>0 to 2</td><td>2 material facts correct, 1 mixed, 0 materially wrong</td></tr><tr><td>Recommendation</td><td>0 to 2</td><td>2 recommended, 1 considered, 0 dismissed or absent</td></tr><tr><td>Evidence</td><td>0 to 2</td><td>2 supporting citation, 1 citation with weak support, 0 none</td></tr><tr><td>Sentiment</td><td>-1 to 1</td><td>Negative, neutral, or positive</td></tr><tr><td>Cited domain</td><td>Text</td><td>Domain beside the relevant claim</td></tr><tr><td>Competitors named</td><td>Text</td><td>Every competing brand in order</td></tr><tr><td>Notes</td><td>Text</td><td>Exact error, qualification, or reason</td></tr></tbody></table>

For unprompted prompts, count a mention only if the answer text names the brand. Do not count a hidden search result, autocomplete suggestion, or citation that never connects the source to the company. For named prompts, the mention field will usually be 1 by construction, so position and recommendation carry more information. Label the mention as prompted to prevent it from inflating discovery recall.

Position needs a rule because prose answers do not always create ranked lists. "First" means the first company discussed as a candidate, not the first occurrence in a disclaimer. "Top three" applies when the answer clearly presents choices in order or gives three early candidates comparable space. If the answer says there is no universal winner and then discusses companies alphabetically, record the text order and explain the absence of a ranking signal in notes.

Accuracy should focus on facts that could change a decision. A minor wording difference does not deserve a penalty. A wrong target customer, unavailable capability, old price, false compliance claim, or incorrect geography does. Keep a small reference sheet of approved current facts so two reviewers judge against the same source.

Evidence receives 2 only when the cited page supports the nearby material claim. A link to your homepage beside a detailed pricing statement may earn 0 if the homepage contains no price. This stricter rule prevents citation decoration from looking like substantiation. It also points directly to missing pages or weak third-party coverage.

Calculate separate rates for the first 12 and last eight prompts. Discovery mention rate equals unprompted prompts with a mention divided by 12. Named-brand accuracy equals the sum of accuracy points on prompts 13 through 20 divided by 16. Keep recommendation and evidence rates separate. You may add a total for trend reporting, but never use it without the component table.

Run the audit without contaminating it

Price the repair plan honestly
A five-business-day review identifies savings before you commit budget to every visibility gap.

Consistency matters more than squeezing every possible answer out of an engine. Run all competitors and engines within a short window, use the same operator instructions, and log exceptions as they happen. If one interface fails, mark the row missing rather than replacing it with a different mode silently.

Start with the unprompted set. Copy the answer exactly, record citations, score it, and close the conversation. Finish prompts 1 through 12 before moving to the named set. This order keeps the operator from unconsciously revising category terms after learning how the engine describes the brand.

For comparisons, do not replace your brand in all 20 prompts with each competitor. The unprompted prompts already expose every company the system volunteers. Run the eight named prompts for each selected competitor, changing only the company name and paired competitor where required. That produces comparable accuracy and recommendation profiles without multiplying meaningless category questions.

Use a second reviewer for disputed rows, especially fit, recommendation, and evidence. The reviewer should see the answer and scoring rules, not the first reviewer's score. Resolve disagreements in notes and tighten the rule for future runs. A scoring system that relies on one person's mood will drift faster than the engines do.

Do not ask the model to score its own answer. It may accept your rubric, but it can also rationalize unsupported claims, overlook a weak citation, or interpret polite wording as a recommendation. Human review is slower and auditable. Automate storage and arithmetic if the volume grows, while keeping the decision fields reviewable.

Repeat a small stability sample before presenting the findings. Choose four prompts covering discovery, comparison, accuracy, and objections, then rerun them in fresh sessions. If the answers swing widely, report the range and avoid claiming that a one-point lead matters. Variance is information: it tells you the association is weak or the result depends heavily on retrieval.

Read gaps by pattern, not by ego

The useful comparison is not "Did we beat everyone?" It is "At which decision stage does each company enter, and what evidence supports it?" Build a matrix with brands as rows and prompt themes as columns. Place the component scores in each cell. Patterns appear quickly.

Low discovery recall with strong named-brand accuracy means the engine can understand your company when asked, but does not associate it strongly enough with the category or use case. Look at which competitors appear and which cited domains introduce them. You may have clear product pages but little independent category evidence, or you may use internal language that buyers never use.

High recall with weak fit is a positioning problem. The brand is known, but public sources connect it to the wrong buyer or an older product. Inspect your homepage title, company descriptions, profile pages, press mentions, comparison pages, and old documentation. Fix contradictions before publishing more content. More pages repeating different descriptions will make the entity harder to resolve.

Good recommendations with low evidence are fragile. The answer may rely on model memory, a search snippet, or a source that does not substantiate the claim. Publish the missing proof in plain text: eligibility rules, current pricing logic, supported use cases, limitations, customer evidence you have permission to share, and clear ownership of claims. Do not manufacture reviews or consensus.

Strong evidence with weak recommendations is not necessarily a content failure. The source may correctly show that a competitor fits the prompt better. Read the constraint. If the product does not serve that segment, an accurate exclusion saves a bad sales conversation. Mark the result as correct positioning rather than forcing a visibility campaign around an audience you should not win.

When every brand scores poorly, examine the prompt and engine before celebrating a tie. The category may be ambiguous, too new, too local, or poorly documented. The answer may also avoid recommendations for a high-risk request. A weak market-wide result calls for better public explanation and careful testing, not a claim that you lead because nobody appeared.

Competitor citations are a research queue, not a copying plan. Identify why a source was usable: direct comparison, concrete definition, named author, current product facts, original data, or clear limitations. Create better evidence for your own claims. Copying headings and phrasing can erase the distinctive information an answer system needs to tell companies apart.

Crawler checks prevent false content diagnoses

Build a smaller execution team
I have run production with two AI-augmented engineers while preserving output and uptime.

Before rewriting the site, confirm that answer engines and their search suppliers can fetch the pages you expect them to use. A blocked crawler, authentication wall, bot challenge, noindex, or snippet restriction can make good content unavailable to retrieval. The correction belongs in infrastructure, not in another blog post.

OpenAI's publisher guidance says sites should allow OAI-SearchBot if they want content included in ChatGPT summaries and snippets. Google Search Central says pages must be indexed and eligible for a snippet to appear as supporting links in AI Overviews or AI Mode. Google also says there is no special AI schema or machine-readable file required. I would not buy an "AI visibility" package whose main deliverable is an invented markup vocabulary.

Check the public response and robots rules from outside your authenticated application. These commands do not prove indexing, but they catch redirects, accidental blocks, and headers that deserve investigation:

curl -I https://example.com/important-page
curl -s https://example.com/robots.txt
curl -s https://example.com/important-page | grep -iE 'noindex|nosnippet|max-snippet'

A healthy first response usually shows a successful status, a crawlable canonical destination, and no unintended X-Robots-Tag restriction. The HTML check should match what the server returns to ordinary visitors. If a CDN or web application firewall challenges automated clients, inspect its logs rather than assuming robots.txt is the only gate. OpenAI's documentation specifically warns that host or content delivery network controls can block its published search crawler traffic.

Crawler access creates eligibility, not inclusion. Google says meeting technical requirements does not guarantee crawling, indexing, or serving. OpenAI likewise says there is no way to guarantee top placement. Anyone selling guaranteed citations is claiming control they do not have.

Use the platform's own webmaster tools where available. Google Search Console can confirm indexing and show the rendered HTML through URL Inspection, but Google reports AI feature traffic inside the broader Web search type. That reporting limit is another reason to pair referral and conversion data with the controlled prompt audit instead of pretending either source is complete.

Fix the source that caused the gap

Stop funding an oversized queue
The audit maps work and cost so visibility repairs do not require another ten-person team.

A good remediation plan maps each failed score to a source, owner, and observable correction. Do not respond to every absence by publishing another generic article. The engine may need a clearer product page, a crawl fix, consistent organization details, a comparison that admits tradeoffs, or credible third-party evidence.

For accuracy errors, create one canonical public statement for each decision-grade fact and remove contradictions. State who the product is for, what it does, where it is available, how pricing works, and which limitations matter. Put essential facts in visible text. Structured data should agree with that text, not introduce claims users cannot see. Google Search Central makes the same point in its AI feature guidance.

For weak fit, write around the buyer's actual problem and constraints. A useful page explains when the product fits, when it does not, prerequisites, implementation effort, and the alternative a buyer should choose in the excluded case. This language gives retrieval systems discriminating facts. Generic claims about being faster or smarter give them little reason to choose one brand over another.

For an evidence gap, decide what proof you can publish honestly. Original research needs a described method. Customer examples need permission and enough context to connect action with outcome. Documentation needs stable pages and explicit scope. A founder's opinion can help with category framing, but it should not masquerade as independent validation.

Assign technical fixes to engineering, factual ownership to the product or operations owner, and evidence development to whoever can verify the claim. Marketing can coordinate, but it cannot validate every compliance, pricing, or product statement alone. Put a review date beside facts that expire.

Prioritize by decision impact and repairability. A false claim about eligibility or price deserves faster work than absence from a broad brainstorming prompt. A crawler block on all product pages has wider impact than weak wording on one comparison. The scores create the queue; business risk sets its order.

Measure change without promising attribution

Rerun the fixed prompt set monthly at first and after material site, product, or crawler changes. Preserve the original wording, engine, account state, geography, and scoring rules. Add prompts only as a new version of the benchmark, because silently replacing weak questions makes the trend meaningless.

Compare component rates, competitor patterns, citation domains, and factual errors. Track whether corrected claims disappear, whether your own authoritative pages begin supporting answers, and whether unprompted recall spreads to the use cases you chose. Keep example answers beside the numbers. Executives need to see whether a five-point change reflects one new mention or a genuine shift across the journey.

Then connect the audit to business signals without claiming that one caused the other. Tag referral traffic from answer engines where analytics expose it. Watch branded search, qualified lead notes, sales calls that mention an answer tool, and conversions on cited pages. OpenAI says publishers that allow OAI-SearchBot can track ChatGPT referral traffic in analytics. Google includes AI feature activity in Search Console's Web reporting, so its interface does not provide a clean standalone AI visibility total.

Set a decision threshold before the rerun. For example, investigate any new material error immediately, open a remediation item when the same discovery theme misses twice, and treat small score movement inside your stability range as noise. Thresholds keep teams from celebrating random positive answers or panicking over one omission.

The score is a diagnostic, not a revenue forecast. It samples how systems answer a controlled set of questions. It cannot measure every buyer prompt, private model behavior, or the persuasive effect of a mention. Use it to choose work and test whether public understanding changed. Use pipeline data to judge commercial value.

If the audit exposes a long repair queue, fund the highest-impact corrections instead of creating an "AI search" side project with no owner. In a Team & AI Audit, I look for where a smaller AI-augmented team can reclaim engineering capacity and ship the work that already matters; the same discipline applies here. Fix access, facts, fit, and evidence in that order, rerun the unchanged prompts, and let the observed answers decide whether the gap closed.

Frequently Asked Questions

What is brand visibility in AI search?

It is the degree to which an answer engine recalls your brand, connects it to the right need, describes it accurately, and supports claims with evidence. A mention alone measures only recall, so a useful audit keeps those outcomes separate.

How many prompts do I need for an AI visibility audit?

Twenty well-chosen prompts are enough for a practical baseline in one market. Keep a mix of unprompted discovery questions and named-brand questions, then preserve their wording across runs.

Which AI search engines should I test?

Test at least two answer engines your buyers plausibly use. Record the product, mode, login state, location setting, and date because retrieval and personalization can change the result.

How often should I rerun the audit?

A monthly run works well while you are correcting gaps. Rerun after material crawler, website, positioning, or product changes, but keep the benchmark prompts unchanged.

Can a high AI visibility score predict revenue?

No. The score diagnoses public recall, fit, accuracy, recommendations, and evidence for a controlled prompt set. Use referral, lead, sales, and conversion data to assess commercial value.

Why does a competitor appear when my site ranks higher?

Answer engines may rewrite a prompt and retrieve sources across related subtopics, so the literal search ranking does not determine the final answer. The competitor may also have clearer category associations or stronger supporting sources.

Does schema markup improve AI search visibility?

Accurate structured data can help search systems understand eligible content, but Google says there is no special schema required for its AI features. Markup must match visible facts, and it cannot compensate for blocked crawling or weak evidence.

Should I block AI crawlers from my website?

That is a policy decision, but it has a visibility consequence. If you want content included in ChatGPT summaries and snippets, OpenAI says not to block OAI-SearchBot; restrictions may reduce what a system can retrieve or show.

What should I fix first after the audit?

Fix material false claims and broad crawler blocks first, then address weak category fit and missing evidence. Business risk should set the order, not whichever score looks easiest to raise.

Can I automate AI brand visibility monitoring?

You can automate prompt execution, storage, and arithmetic when platform terms permit it. Keep human review for fit, recommendation strength, factual accuracy, and whether a citation truly supports the claim.

Related Posts