Skip to content
8 min read

AEO tools and the numbers they actually measure

Compare what AEO tools track, how current plans are priced, where their visibility scores fail, and when a spreadsheet gives you better evidence.

AEO tools and the numbers they actually measure
Table of Contents

AEO tools do not measure whether an AI system "knows" your company. They run a chosen set of prompts, collect the answers they receive, classify mentions and citations, then compress that sample into charts. The useful output is evidence about those sampled answers. The dangerous output is a tidy score that encourages executives to mistake the sample for the whole market.

I have bought enough analytics software to distrust a metric until I can explain its denominator. A brand can gain ten points of "AI visibility" because the tracked prompt set changed, a vendor added a model, or one assistant started returning longer answers. None of those changes proves that more buyers discovered the brand.

That does not make this software useless. It makes tool selection a measurement design problem. Choose the question first, define the sample, preserve the raw answers, and only then decide whether automation is worth paying for.

An AI visibility score is a sample statistic

Every visibility score begins with a finite prompt set. The platform sends those prompts to one or more answer engines, often with a selected country, language, device context, or model variant. It records whether the target brand appears, where it appears, which sources the answer cites, and sometimes the tone of the text. The score summarizes those observations.

A simple mention rate is easy to understand:

mention_rate = answers_that_mention_brand / valid_answers

Share of voice usually compares your mentions with the mentions of named competitors. Average position may count the order in which brands appear. Citation share may count cited URLs, cited domains, answers with at least one citation, or all three. Vendors can use the same label for different formulas, so the label alone tells you very little.

The denominator matters more than the dashboard design. Suppose you track 40 prompts in ChatGPT and Perplexity once a day. Your theoretical monthly sample is 2,400 answers before retries, blocked requests, empty responses, and parsing failures. If 30 prompts describe early research while 10 name a product category, the blended score mostly reports early research visibility. It does not represent an even view of the buying journey.

Ask a vendor to expose four things: the exact prompts, the complete answers, the execution timestamp, and the formula behind every aggregate. If you cannot export those records, you cannot investigate a spike or reproduce a board report. You are renting a conclusion without keeping the evidence.

This distinction is the one teams routinely blur: observed answer presence is not customer demand. A prompt tracker tells you what appeared when the tracker asked. It cannot tell you how many people asked that question, saw that answer, trusted it, visited later through another channel, or bought. Search volume estimates and referral analytics can add context, but they do not turn sampled answers into audience measurement.

Four tool categories answer different questions

The AEO market is easier to compare when you separate four categories that vendors often bundle together.

Prompt monitors repeatedly run questions you choose. They are best for a controlled watchlist such as "best payroll software for a 50 person company" or "Brand A versus Brand B." Peec AI, OtterlyAI, and the custom tracking parts of Scrunch, Semrush, and Ahrefs fit here. Their core unit is usually a prompt, an answer, or a check.

Discovery indexes collect a much larger vendor chosen set of prompts and let you search it for brands, topics, and cited domains. Ahrefs Brand Radar and Semrush Visibility Overview lean into this category. Discovery is useful when you do not yet know which prompts matter or when you want a broad competitor baseline. The blind spot is sampling control: a huge index can be irrelevant to your actual market if its questions, countries, or intent mix do not match your buyers.

Technical audits inspect whether crawlers can reach and parse a site. They check items such as robots directives, server responses, rendered content, structured data, and page clarity. Scrunch, Semrush, and OtterlyAI include audit features at different limits. An audit can find a blocked crawler or a page that returns an empty application shell. It cannot prove that fixing the page will make an assistant cite it. Access is a prerequisite, not a ranking formula.

Attribution products connect AI referrals, server logs, web analytics, or conversion data to visibility work. This category gets closest to commercial value, but it sees only part of the journey. A buyer may read an AI answer and later arrive through branded search or direct navigation. Conversely, a referral from an assistant does not prove that a monitored prompt caused the visit.

Content recommendation layers sit across these categories. They turn gaps into suggested pages, edits, or outreach targets. Treat those suggestions as hypotheses. A system can observe that competitors get cited from comparison pages, but it cannot know whether publishing another generic comparison page will earn a citation or help a buyer. The recommendation becomes useful only when it names the observed answer, cited source, missing fact, and expected business effect.

The current plans are not priced on one unit

The table below uses public vendor pages checked on August 8, 2026. Prices change, annual discounts obscure monthly cash cost, and some pages localize currency. Recheck the live plan before approving a purchase.

PlatformPublished entry pointWhat the entry plan measuresMaterial limit or blind spot
OtterlyAI$29 per month15 prompts across ChatGPT, Google AI Overviews, Perplexity, and Microsoft Copilot, with daily trackingGoogle AI Mode and Gemini are paid additions; a small prompt set can produce a fragile trend
Peec AI$95 per month50 prompts, three selected models, daily tracking, one project, and unlimited usersPrice grows with prompts and added models; a project is not the same as a market or product line
Semrush AI Visibility$99 per month per domain when billed annuallyDatabase research plus 25 custom prompts across ChatGPT, Google AI, Gemini, and PerplexityOne domain, limited prompt tracking, and separate charges for extra domains or user access
Ahrefs custom prompts$50 per month2,500 checks that can be allocated across platforms, locations, prompts, and frequencyCheck consumption requires planning; it is depth tracking rather than the full discovery index
Ahrefs Brand Radar$199 per month for one index or $699 for all platformsSearch across a large vendor built prompt index; all platform access includes 2,500 custom checksThe index reflects Ahrefs' modeled prompts, while full platform coverage costs far more than custom prompts alone
Scrunch$250 per month billed annually350 custom prompts, 1,000 industry prompts, three users, five page audits, personas, and agent traffic monitoringBroader workflow costs more; enterprise security and the data API sit above the entry plan

These prices are not comparable until you normalize the unit. OtterlyAI prices a small library of prompts over an included group of engines. Peec prices a prompt allowance and a selected model count. Ahrefs custom tracking prices checks, so one prompt run daily in two locations on three platforms consumes six checks per day. Semrush combines a discovery database, a per domain workspace, and a smaller custom tracker. Scrunch combines monitoring with audits and traffic features.

Calculate the monthly observation volume you need:

observations = prompts x platforms x locations x personas x runs_per_month

Then add seats, projects, exports, API access, retained history, and onboarding. A $29 plan that cannot represent your two countries may cost more in analyst time than a $250 plan. A $699 index is wasteful when the team cares about 30 purchase prompts. Headline price is the wrong comparison column.

The vendor documentation also reveals the intended use. Ahrefs says its index is built from hundreds of millions of search backed prompts modeled from its keyword database. That breadth is useful for discovery, but it is still a modeled corpus. Peec defines an AI answer as one prompt result for one model, which is a clean operational unit. Semrush separates its prompt database from daily tracking of prompts you select. Those are three different datasets, even when each screen calls the output visibility.

Your prompt library determines most of the result

A prompt library needs to represent buyer decisions, not the phrases your marketing team wants to win. If the set contains only flattering category questions, the score rewards brand mentions in an invented market. If it contains only prompts that name your brand, it measures whether the model can repeat your positioning. Neither set tells you whether an undecided buyer encounters you.

Build the library from real evidence: sales call notes, site search, support questions, paid search terms, comparison pages that convert, and objections heard during procurement. Keep branded diagnostic prompts separate from unbranded discovery and purchase prompts. A blended score hides a painful but common state where assistants describe your brand accurately when asked by name yet never recommend it to a buyer who starts with a problem.

Location and language belong in the prompt definition. "Best invoicing software" asked in English from the United States is not the same test as its literal translation run in Germany. Regulations, available products, currency, and local sources change the answer. Translating a US prompt list creates neat cross market charts with little decision value.

Personas deserve restraint. Adding "I am a CFO" can materially change a prompt. Adding five decorative persona descriptions multiplies cost without proving that actual users speak that way. Start with intent and constraints that affect the answer, such as company size, regulated data, existing stack, or buying authority. Drop any modifier that does not change the decision.

Before a trial, write a measurement contract in one page. It should name the business question, prompt source, intent groups, competitors, engines, locales, run frequency, success metric, and exclusions. Include the rule for changing prompts. If a team can silently replace weak prompts, its trend line becomes a record of editorial changes rather than market movement.

A useful contract might say: "Track 60 fixed English prompts used by US founders evaluating engineering services. Report mention rate by intent group, citations to our domain, competitor co-mentions, and valid answer count. Review prompts quarterly; never rewrite historical results. Treat a five point change as investigatory until raw answers show a consistent cause." That is more defensible than asking a platform to "track AEO."

Mentions, citations, position, and sentiment disagree for good reasons

Match measurement to execution
Ongoing fractional CTO support starts at $5,000-10,000 monthly for practical AI team transformation.

A mention means the answer contains a recognized brand name or alias. It may be a recommendation, a warning, a comparison, or a passing reference. Alias handling creates obvious errors: short brand names can match ordinary words, while product names, parent companies, misspellings, and localized names can be missed. Inspect the matcher and keep a false positive list.

A citation is a source attached to an answer, but platforms expose sources differently. Some answers show inline links, some provide a source panel, and some mention a domain without linking it. A tool may count citation instances, unique cited pages, or answers that cite the brand domain. Those numbers answer different questions. A page cited twice in one answer should not automatically count as two successful prompts.

Position sounds familiar to anyone who has tracked search rankings, yet generated answers do not have ten stable blue links. One response may put a brand in a table, another may mention it in prose, and another may list it under "not suitable for." Converting those layouts into position 1, 2, or 3 creates a convenient number with weak semantic consistency.

Sentiment is even more delicate. Product advice often mixes praise and qualification in the same paragraph: suitable for enterprises, expensive for small teams, strong integrations, difficult migration. A positive or negative label throws away the condition that makes the text useful. Store the passage and the rationale behind the label. Use sentiment to find answers for human review, not as a board target.

Citations also expose a problem that mention rate hides. Your brand can appear often while every supporting source belongs to review sites, affiliates, or competitors. That leaves your product facts dependent on intermediaries. The response is not to publish hundreds of pages. Find the recurring factual gap, put a clear answer on the authoritative page, and check whether assistants can retrieve it. Pricing, limits, supported regions, security statements, and comparison criteria need explicit prose that a reader can verify.

Repeated runs reveal variance, not ground truth

Generated answers vary. Models sample tokens, search indexes change, retrieval systems choose different sources, and providers update their products. The same prompt can return different brands minutes apart. One daily answer gives you a time series, but it does not tell you the within day variance.

Run a small stability test before trusting a score. Pick ten prompts from different intent groups and run each several times on the same engine, locale, and day. Compare brand inclusion, citations, and recommendation order. If a target appears in two of five runs, recording a single run as a binary win or loss exaggerates certainty. The platform should let you see retries or at least disclose how it handles them.

Failures need their own denominator. Timeouts, refusals, empty answers, login walls, and parser errors are not negative brand observations. If a tool quietly counts them as absence, an engine outage can look like a visibility collapse. If it drops them without reporting valid answer count, a rising score may simply reflect a smaller sample.

Here is a failure I have seen in adjacent monitoring systems. A weekly report showed a competitor jumping from 18 percent to 31 percent share of voice. The team prepared a content response. Raw data revealed that the vendor had added recommendation prompts to the tracked set, and those prompts disproportionately favored the competitor's category. Nothing in the market had moved. The sample had moved.

Version the measurement inputs. Store a prompt set version, engine label, locale, location, run schedule, and parser version beside every observation. When a vendor changes an engine connector or classification rule, mark the chart. Historical continuity is a product feature, not a cosmetic reporting choice.

Do not average every engine into one executive score. ChatGPT, Google AI Overviews, Gemini, Perplexity, and Copilot have different interfaces and retrieval behavior. Buyers also use them differently. Report each engine first, then create a weighted total only if you have defensible usage or conversion weights. Equal weighting is an assumption, not neutrality.

Crawler readiness cannot promise answer inclusion

Shrink the team around AI
The operating target is one or two AI-augmented engineers shipping about three times faster.

A technical AEO audit can prove that a crawler was allowed to request a page and that the returned document contained usable content. It cannot prove that an answer engine indexed the page, selected it for retrieval, trusted its claims, or will cite it for a particular prompt. Tools that collapse this chain into one readiness grade make a diagnostic check look like an outcome forecast.

RFC 9309 defines how crawlers interpret robots.txt rules. That standard gives teams a precise way to test declared access, but it says nothing about citation eligibility. A site can welcome every crawler and still offer thin, contradictory, or stale product facts. Another site can block a named training crawler while remaining visible through search indexes or other retrieval paths. Treat each user agent and access path separately.

Rendering checks matter when the initial HTML contains little more than an application shell. A human browser may execute scripts and show a complete pricing table, while a basic fetch receives headings with no prices. Server logs can confirm that a declared crawler requested the URL and the response status, but they do not reveal whether the engine understood or retained the content. A crawler visit is evidence of access, not evidence of use.

Structured data has the same boundary. Valid markup can disambiguate an organization, product, author, price, or FAQ. It reduces parsing work and helps machines connect facts. It does not override a weak page, manufacture authority, or force a citation. Reject any sales pitch that turns schema validation into a ranking guarantee.

Audit output becomes actionable when it preserves the chain of evidence. For each finding, record the affected URL class, requesting user agent, robots rule, response code, returned content, render difference, and the business facts missing from the response. Then assign the fix to engineering, content, or legal. A generic score of 72 gives none of those owners a testable task.

The most useful retest is narrow. After fixing a blocked or empty page, verify that the relevant crawler can fetch the intended text, then watch the prompt group that depends on those facts. A later citation supports the hypothesis that access mattered. No change means you keep investigating retrieval, source authority, factual clarity, or prompt relevance instead of endlessly polishing technical checks.

This is where combined platforms can justify a higher price. If an audit finding links directly to affected prompts, raw answers, cited competitors, and later observations, the tool shortens an investigation. If the audit and visibility chart merely share a navigation menu, the bundle has not created evidence. It has put two products on one invoice.

A spreadsheet wins while the questions are still changing

Audit the bigger cost first
A $5,000 Team & AI Audit must identify $50,000 in yearly savings or it is free.

A spreadsheet beats a subscription when the team is still learning what to measure, has fewer than roughly 50 important prompts, checks one or two markets, and can tolerate weekly rather than daily collection. The manual work forces useful decisions about aliases, intent, citations, and failures. Software purchased before those decisions only automates confusion.

Use one row per observed answer with these columns:

FieldExampleWhy it exists
run_at2026-08-08T14:00:00ZSeparates observations and exposes gaps
prompt_idBUY-014Keeps wording changes out of the identifier
prompt_version3Preserves comparability after an approved edit
enginePerplexityPrevents meaningless cross engine averaging
localeen-USCaptures market context
valid_answerTRUEProtects the denominator from collection failures
target_mentionedTRUESupports a transparent mention rate
competitorsBrand B; Brand CMakes share calculations auditable
cited_domainsPublisher; review sitePreserves evidence behind citation totals
answer_textFull response textLets a person investigate any score
notesQualified recommendationRetains nuance the classifier misses

Keep a second sheet for prompts with the exact text, intent group, source, owner, added date, retired date, and exclusion reason. Never overwrite a prompt. Retire one version and add another. A third sheet can summarize rates with pivot tables, but the observation sheet remains the source of truth.

The manual process has limits. Personal accounts may produce results influenced by chat history or account settings. Copying answers is slow, and interface changes break routines. Automated use may face provider terms or rate limits. Do not build an unauthorized scraper because a subscription feels expensive. Use supported interfaces, document the collection context, and accept that a manual baseline is a research sample.

The spreadsheet stops winning when collection consumes the review time. If an analyst spends six hours copying answers and ten minutes studying them, automation has a clear job. It also loses when agencies need strict client separation, several locales multiply runs, executives need scheduled reports, or an API must feed a warehouse. The threshold is operational, not ideological.

A subscription earns its cost through evidence and workflow

Buy a prompt monitor when you have a stable library and need repeated collection. Buy a discovery index when you need to find unknown prompts, sources, or competitors at scale. Buy an audit layer when crawler access and page rendering are active problems. Buy the combined suite only if the same team will use those parts in one operating rhythm.

During a trial, ignore the polished overview for the first hour. Export raw answers. Recalculate mention rate for one prompt group. Check five citations by hand. Trigger or locate a failed run. Change an alias and see whether history changes. Compare the same prompt across two engines. If the platform cannot support that inspection, its score is unsuitable for an executive decision.

Then test workflow questions. Can an analyst annotate a false positive? Can you freeze and version a prompt set? Do exports include full answer text and timestamps? Can you separate countries, products, and customer segments without buying a project for each? Does the API expose the fields shown in the user interface? How long does the platform retain raw responses? A lower price does not compensate for evidence trapped in screenshots.

Connect visibility to behavior carefully. Tag assistant referrals in analytics, inspect server logs for declared AI crawlers, ask sales how buyers researched the shortlist, and watch branded search. These signals can support a causal story, but none proves it alone. A visibility increase with no citations, traffic, pipeline influence, or improved answer accuracy is an interesting observation, not a return calculation.

The unpopular recommendation is to delay buying for two weeks and run the spreadsheet baseline. Vendors dislike it, and teams prefer a dashboard because procurement looks like progress. The baseline gives you the prompt library, aliases, variance estimate, and required export fields that make a trial decisive. Without those, every demo looks comprehensive.

A Team & AI Audit from oleg.is can examine this measurement work alongside the engineering and AI workflow that produces it. The fixed engagement costs $5,000, takes five business days, and is free if it does not identify at least $50,000 a year in savings.

Do not ask which AEO platform has the best visibility score. Ask which dataset can support the next decision, whether you can inspect its evidence, and what it costs per valid observation. If the vendor cannot answer those questions, keep the spreadsheet.

Frequently Asked Questions

What do AEO tools actually measure?

They measure a sample of generated answers produced from a defined set of prompts, engines, locations, and run times. Most platforms classify brand mentions, citations, position, competitors, and sometimes sentiment, then aggregate those observations into a score.

Is an AI visibility score the same as market share?

No. It is a share within the vendor's observed prompt sample, not a count of real users or purchases. Treat it as a monitoring statistic whose meaning depends on the prompt set and denominator.

How much do AEO tools cost?

Public entry plans in this comparison run from $29 a month for a small prompt allowance to hundreds of dollars for broader indexes and workflows. Normalize prompts, platforms, locations, frequency, seats, exports, and API access before comparing prices.

Which AEO metric should a startup track first?

Start with valid answer count and mention rate for a small set of purchase and comparison prompts. Keep citations and full answer text beside the rate, because a mention without context can be negative or irrelevant.

Can AEO software measure revenue from AI search?

It can connect some assistant referrals and conversions, but it cannot see every research path. Buyers often return through branded search or direct navigation, so combine referral data with sales research notes and other demand signals.

How many prompts should I track?

Track enough prompts to cover distinct buyer intents without filling the library with minor wording variants. For many startups, 30 to 60 well sourced prompts teach more than hundreds of synthetic questions.

Why do AI visibility results change between runs?

Generated answers vary because model sampling, retrieval, indexes, and product versions change. Repeated runs on a small test set reveal how much of a movement is ordinary variance.

When is a spreadsheet enough for AEO tracking?

A spreadsheet is enough while prompts are changing, the watchlist is small, and weekly checks can support the decision. It stops being economical when collection across engines and markets consumes more time than analysis.

Should I track citations or brand mentions?

Track both because they answer different questions. Mentions show whether the brand appears, while citations show which pages and domains supply evidence for the answer.

What should I test during an AEO tool trial?

Export raw answers, recalculate one metric, inspect citations, find failed runs, test aliases, and compare engines. Also verify prompt versioning, history retention, project limits, and whether the API exposes the evidence shown in the dashboard.

Related Posts