Skip to content
8 min read

Share of model belongs in your board report

Learn how to define share of model, measure it consistently across ChatGPT, Perplexity and Gemini, and turn the result into a credible board report.

Share of model belongs in your board report
Table of Contents

Share of model should appear in a board report when AI answers influence how buyers discover, compare, and shortlist companies in your category. It should not appear as a shiny percentage with no definition, sample, or connection to revenue. A board can act on the first version. The second produces a discussion about the number and nothing else.

I use the term for a brand's share of relevant AI-generated answers under a controlled set of buyer questions. That definition sounds simple. The work lies in controlling the questions, platform settings, geography, time, scoring, and uncertainty well enough that a later result means the same thing. If any of those inputs drift silently, a chart can rise while the brand's position has not changed at all.

This is an emerging metric, not an audited accounting standard. Treat it as a repeatable market signal alongside search share, qualified pipeline, win-loss evidence, and direct customer research. It can reveal a weak position before revenue reports do, but it cannot prove that an AI answer caused a sale.

The business question comes first: when a prospective buyer asks an AI service for help in your category, does the answer represent your company accurately and give it fair consideration? A trend line is useful only insofar as it sharpens that question. This framing also stops the metric from becoming a contest to mention the brand everywhere. A company should not want a recommendation for a customer it cannot serve, and it should not count an irrelevant mention as progress.

Set the measurement boundary in a one-paragraph charter before collecting anything. Name the buyer population, markets, products, competitors, AI surfaces, decision stages, reporting cadence, and executive decision the metric will support. State what it will not measure, including total consumer awareness and sales attribution. Ask the metric owner and finance lead to approve that charter. This small document prevents six months of debate about a denominator nobody agreed to.

Share of model measures answer presence, not awareness

Share of model is the percentage of eligible AI answers in which a company receives a defined kind of presence, relative to named competitors or to all answers in a category. The denominator matters more than most published examples admit. You can measure share among brands mentioned, share of recommendations, share of citations, or share of weighted visibility. Those are different metrics and should never share one label.

A basic unweighted version for brand b is:

mention_share(b) = answers_that_mention(b) / all_eligible_answers

If an answer names three competitors, each can receive one mention. Shares across companies can therefore total more than 100 percent. If the board expects a market-share chart that always sums to 100, use a normalized competitive share instead:

normalized_share(b) = mentions_of(b) / mentions_of_all_tracked_brands

That normalization answers a narrower question: how much of the tracked competitive attention did each brand receive? It hides answers that mention no tracked brand. Keep a separate "no brand mentioned" rate or the report will make a weak category presence look healthy.

Share of voice usually counts exposure in media, advertising, social posts, or search results. Those surfaces have observable units such as impressions, placements, or posts. AI answers are synthesized. One response may recommend a company, criticize it, cite its documentation, or mention it only as a comparison. Counting all four as equal presence destroys the meaning of the metric.

Three distinctions keep the language honest. Presence means the brand appears. Prominence means where and how strongly it appears. Preference means the answer recommends it for the user's need. Citation means the system points to the brand or its content as evidence. Report each separately before considering a composite score.

A governed question set is the real measurement asset

A stable, documented set of buyer questions determines whether the measurement reflects your market or merely your prompt writer's instincts. Build it from actual decision moments: discovery, problem diagnosis, alternative comparison, risk review, migration, pricing, and final selection. Product marketing, sales, support, search data, and win-loss interviews should all contribute, but no single team should own the set alone.

Start with the customer and buying situation, not the desired brand mention. "What tools can replace manual month-end reconciliation for a 50-person finance team?" is useful. "Why is Acme the best reconciliation tool?" measures prompt compliance. Questions should sound like something a buyer would type when the company is not in the room.

Give every question a permanent ID and attach a small amount of metadata:

  • Journey stage and business problem
  • Customer segment and geography
  • Commercial importance, expressed as a declared weight
  • Expected answer type, such as explanation, comparison, or shortlist
  • Inclusion date and reason for any later change

Do not create hundreds of paraphrases and pretend they are hundreds of independent buyer needs. Near-duplicate prompts can make a result look precise while multiplying one editorial assumption. I prefer a smaller core set with deliberate variants for wording, locale, and customer context. Separate the stable benchmark set from an exploratory set that can change as the market changes.

A quarterly review can add or retire questions, but preserve the prior version and report the effect of the change. Run the old and new sets in parallel once. That overlap tells the board whether a jump came from the market or from the instrument. Quietly replacing weak prompts is the AI visibility equivalent of moving a sales target after the quarter closes.

The popular recommendation to scrape every "people also ask" query and send the whole pile to each model is wrong. It is popular because volume feels objective and automation feels cheap. The resulting sample overweights search behavior, repeated phrasing, and informational questions, then underweights the commercial questions a board cares about. Coverage comes from a reasoned sampling frame, not prompt volume.

Weights need the same discipline. A weight of five for an enterprise security question and one for a general definition says the first scenario matters five times as much to management. It does not say five times as many buyers ask it. Keep both weighted and unweighted results so a director can see how much the declared priorities affect the story. If one heavily weighted prompt can swing the total, show that sensitivity rather than hiding it inside the aggregate.

Before freezing the set, ask sales and support to classify each question without seeing who wrote it. Disagreement often reveals a segment that uses different language or a question that combines two buying stages. Rewrite those items, then pilot them on all platforms. A good pilot checks whether the services understand the scenario, whether the answer can be scored under the rubric, and whether the question accidentally names the category in a way real buyers do not.

Comparable tests require a written run protocol

You can compare ChatGPT, Perplexity, and Gemini only after you define exactly what a run means on each service. Record the product surface, model or mode label shown to the user, account state, browsing or research mode, locale, location setting when available, date, and whether the conversation started clean. Do not combine answers from a consumer chat, an API, and a research mode under one platform name.

The minimum protocol is plain:

  1. Start a new conversation for every prompt so earlier answers cannot contaminate the next one.
  2. Use the same prompt text, language, buyer facts, and requested output format on all three services.
  3. Save the full answer, visible citations, run timestamp, and the settings a reviewer can see.
  4. Repeat the prompt on multiple scheduled runs instead of retrying until an attractive answer appears.
  5. Store refusals, errors, and empty answers rather than silently replacing them.

AI products change. Model labels, default behaviors, citation displays, and access conditions can differ by plan and region. A credible report identifies the observed surface instead of claiming to measure an entire company. Write "Gemini consumer chat, standard mode" or the equivalent label captured during the run, not simply "Gemini." This is less elegant on a slide and much more useful when the result changes.

Keep prompt formatting neutral. Asking for "the five best vendors in a table with links" can force mentions and citations that a natural answer would not contain. If your buyer usually asks for a shortlist, request one. If the buyer asks for help understanding a problem, let the answer remain explanatory. The instrument should imitate the decision, not manufacture a convenient scoring surface.

Manual collection is acceptable for an initial benchmark and often exposes protocol gaps faster than automation. Automation becomes useful once the questions and scoring rules have survived a few rounds. Platform terms, available interfaces, and data-handling rules must constrain collection. Do not build the board's metric on brittle browser automation that you cannot explain or maintain.

Score presence, prominence, preference, and citation separately

A useful scoring system preserves the differences that a single visibility score erases. For each answer, record four fields: whether the brand appears, how prominent it is, whether the answer recommends it for the stated need, and whether a citation supports the claim. Add sentiment or factual accuracy only if reviewers can apply a written rule consistently.

I use a simple prominence scale that a second reviewer can reproduce:

0 = absent
1 = incidental mention or undifferentiated long list
2 = substantive description or direct comparison
3 = top recommendation tied to the buyer's stated need

Recommendation needs its own Boolean field because a detailed warning can earn prominence without preference. Citation also needs its own field. A cited article written by the brand is evidence of content influence, while a citation to an unrelated publisher that happens to name the brand is different. If owned and third-party citations matter strategically, split them.

Do not let an opaque weighted score lead the report. A formula such as 40 percent mentions, 30 percent rank, 20 percent sentiment, and 10 percent citations expresses management preferences, not a natural law. It can help an operating team sort work, but the board should see the components and the weights. Change a weight only prospectively, then show the old calculation once for continuity.

Accuracy deserves a defect log rather than a bonus inside visibility. A confident recommendation based on an obsolete feature, wrong market, or invented price creates exposure and risk at the same time. Record the claim, the answer excerpt needed to identify it, a severity rating, the likely source when visible, and the owner who will investigate. Never celebrate a rising share while material errors rise with it.

Use two reviewers on a sample before scaling. Have them score independently, compare disagreements, and rewrite the rubric where reasonable people interpreted it differently. The goal is not to train reviewers to agree with the metric's owner. The goal is to remove decisions that exist only in one person's head.

Set an adjudication rule for the disagreements that remain. The second reviewer should not simply defer to the first or to the more senior employee. Both should point to the exact answer text and the relevant rubric clause, then an adjudicator records the decision and whether the clause needs a prospective change. Keep original scores. Otherwise the cleaned data will conceal where the instrument was hard to apply.

Names create another quiet source of error. A product name may also be a common word, a company may appear under a former name, and an answer may cite a parent company while recommending a subsidiary. Build an alias table with explicit include and exclude examples. Run automated matching only to propose candidates; a reviewer should resolve ambiguous cases. False matches tend to cluster, so one bad alias can move a chart far more than one variable answer.

A reproducible worksheet beats a mystery dashboard

Audit the team behind AI visibility
The five-day Team & AI Audit identifies where your current engineering setup wastes money and attention.

The first working system can be a table that preserves raw evidence and calculates transparent aggregates. One row should represent one answer-brand pair, not one prompt. That shape handles answers with several competitors without stuffing a list into a cell.

A practical CSV header looks like this:

run_id,run_date,platform,surface,question_id,segment,question_weight,brand,mentioned,prominence,recommended,cited,accuracy_issue,reviewer
2026Q1-0042,2026-02-12,ExampleAI,consumer_chat,Q017,midmarket,2,Acme,1,2,0,1,0,RS

The example uses a fictional platform and brand. Your evidence store should also keep the prompt and complete answer under run_id; the board table does not need that text, but an analyst must be able to trace every cell back to it. Remove or restrict any personal or confidential information before collection. Real customer prompts often contain details that do not belong in a third-party chat or a broadly shared worksheet.

Calculate a weighted mention rate without hiding the denominator:

weighted_mentions = sum(question_weight where mentioned = 1)
eligible_weight = sum(question_weight for eligible answers)
weighted_mention_rate = weighted_mentions / eligible_weight

Define "eligible" before the run. A platform error may be excluded from the visibility denominator but included in a reliability appendix. A refusal can be eligible if the question itself was valid, because the refusal is part of the buyer experience. Whichever rule you choose, publish error, refusal, and exclusion counts beside sample size.

For normalized competitive share, calculate the weighted mentions for every tracked brand, then divide each brand's value by the total. For preference share, repeat the calculation with recommended = 1. Do not compare your weighted rate with a competitor's unweighted rate or mix an answer-level denominator with a mention-level denominator. Many surprising dashboard movements are denominator mistakes wearing a strategic story.

Version the prompt set, scoring rubric, brand dictionary, and calculation code or formulas. Brand dictionaries need aliases, product names, and rules for ambiguous terms. Keep exclusions visible. If a reviewer cannot reconstruct last quarter's number after an employee leaves, you do not have a metric yet.

Sampling controls keep the trend believable

A share of model result is an estimate drawn from variable outputs, so one response per question is too fragile for a trend claim. Repeat runs across different days and times, retain every scheduled observation, and compare distributions rather than selecting a favorite answer. More repetitions reduce noise, but they do not repair a biased question set.

Keep a fixed benchmark panel for trend reporting. Add a rotating panel to catch new products, buyer concerns, and vocabulary, but label its results separately. This mirrors a sensible product analytics practice: protect the stable series while allowing discovery work to evolve. When a material model or product change becomes visible, annotate the series and consider a parallel run before treating the new result as comparable.

Report the number of prompts, runs, eligible answers, exclusions, and platforms. For a proportion such as mention rate, include an uncertainty interval when the sampling design supports it. Do not present a textbook confidence interval if your questions were handpicked and heavily weighted, then imply it describes the whole market. In that case, call it run-to-run variability within the benchmark panel. Precise language beats decorative statistics.

Seasonality can enter through the questions. Tax software, travel, procurement, and annual planning all have periods when buyer needs change. Preserve a year-round core, then add a labeled seasonal module. Comparing a January tax prompt set with an August general set and attributing the change to content work would mislead the board.

Geography and language are not cosmetic prompt variables. Recommendations can differ because sources, availability, regulation, and brand recognition differ. Run separate panels for markets that affect revenue. Translating an English prompt word for word does not always preserve buyer intent, so a native reviewer should confirm the scenario and expected answer type.

NIST's AI Risk Management Framework treats measurement as part of a larger cycle that includes context, governance, and management. That is the right instinct here. A technically repeatable count can still be strategically useless if the sample excludes your actual buyers or the organization has no owner for inaccurate answers.

Competitive analysis needs evidence, not a league table

Put an owner behind the metric
A fractional CTO turns AI work into an operating cadence with named decisions and accountable owners.

A competitor's higher share tells you where to investigate, not what caused the lead. Inspect the questions, answer claims, citations, and source patterns behind the aggregate. You may find that the competitor owns a clear category phrase, publishes material that answers comparison questions, earns coverage from sources the systems cite, or simply fits the chosen segment better. Each cause suggests a different response.

Walk through a common failure. A company measures 40 broad category prompts, sees a competitor mentioned in most shortlist answers, and decides to publish dozens of generic articles. Three months later its overall mention rate rises. The celebration stops when someone segments the data: the gains came from low-weight educational prompts, recommendation share stayed flat, and several new mentions repeated an obsolete integration claim. The content program improved a vanity denominator while purchase-stage visibility and accuracy did not move.

The repair is specific. Compare prompt-level changes, group them by journey stage and segment, then read the answers that moved. Map cited sources and factual defects. Decide whether the gap needs clearer owned documentation, credible third-party validation, consistent product naming, correction of stale pages, or an actual product change. AI visibility work cannot compensate for a product that does not fit the question.

Avoid reverse engineering one answer and stuffing its phrases into pages. Generated answers vary, retrieval sources change, and copied wording rarely solves weak evidence. Publish material that resolves the buyer's question with verifiable specifics. Make product names, capabilities, constraints, and update dates unambiguous. Seek independent coverage because it helps buyers assess the claim, not because a dashboard assigned a magical score to a domain.

Do not infer a platform's internal training data or ranking logic from visible citations. A citation shows what the answer presented as support in that run. It does not reveal every source that influenced generation. Report observations at that level and keep theories out of the board metric.

The board report must connect exposure to decisions

Move from signal to ownership
Month-to-month founder advisory brings experienced guidance to the decisions behind an AI visibility metric.

A board needs the trend, its business meaning, the main risk, and the decision management wants. It does not need a tour of every prompt. Use a one-page scorecard with a methods appendix that a skeptical director can inspect.

The scorecard should contain:

  • Weighted presence and preference by platform, with the prior period
  • Performance for the two or three segments tied most closely to revenue
  • Sample size, run dates, prompt-set version, and a variability indicator
  • Material accuracy defects and the owner of each response
  • One management action, cost, expected signal, and review date

Show competitors only when the comparison changes a decision. A stacked chart of ten companies consumes attention while hiding the movement that matters. I usually place the company, the strongest relevant competitor, and the category median or "none mentioned" rate on the main page. Put the complete table in the appendix.

Connect the metric to downstream evidence without claiming causation. Track referral traffic when it is observable, brand-search changes, self-reported discovery in lead forms, sales-call mentions, inclusion in shortlists, and win-loss interviews. A buyer may encounter several sources before contacting sales, and some AI surfaces send little or no traceable traffic. The honest statement is that share of model measures answer exposure within the defined panel. Pipeline evidence tells you whether that exposure appears in real buying journeys.

Set thresholds for action before seeing the quarter's result. A material accuracy defect may require immediate correction regardless of share. A sustained preference decline in a priority segment may trigger message research or product review. A small aggregate movement inside normal run variability should trigger nothing. Precommitted rules reduce the temptation to narrate every wiggle.

Write the board slide as a claim, evidence, and request. The claim might be that preference weakened among European midmarket buyers. The evidence should give the magnitude, prior period, variability, affected questions, and the answer patterns behind it. The request should name the choice, such as funding independent technical validation or correcting documentation before a product launch. If management has no request, label the slide as monitoring and keep it out of decision time.

Keep a compact methods box on the same page. A director should be able to see that the result came from, for example, a fixed prompt-set version, named product surfaces, scheduled repeated runs, and a stated scoring rule. The appendix carries the detail, but the main page must expose enough method to prevent a polished chart from acquiring more authority than its evidence supports.

Scenario analysis makes the limits concrete. Show how the headline changes without question weights, with disputed scores removed, or with one platform excluded. A stable conclusion should survive reasonable choices. If the ranking reverses under one plausible denominator, report the result as unresolved and commission more evidence. Boards handle uncertainty well when management describes it plainly; they react badly when hidden assumptions surface after a commitment.

If leadership cannot agree on the questions, segments, or decision rules, the disagreement is strategic and useful. When that conflict exposes a wider problem in AI ownership or engineering cost, oleg.is offers a five-business-day Team & AI Audit focused on team savings and AI transformation. Keep that operational work separate from the market metric, and retain an executive owner for the definition and follow-through.

Treat the metric as an operating signal

Share of model earns a place in management when one owner can trace a board number to the questions, raw answers, scoring rules, defects, and resulting work. Give ownership to a leader who spans marketing, product, and revenue decisions, then assign collection and quality control to named operators. A metric jointly owned by everyone usually gets defended by no one.

Use a monthly operating review if the market and models move quickly, but report to the board at a cadence that supports decisions, often quarterly. The team can inspect prompt-level changes more often without forcing directors to react to noise. Freeze each reporting period after review so later rescoring does not rewrite history; log corrections as corrections.

Budget for maintenance. Question sets age, products rename features, competitors enter, model surfaces change, and reviewers drift. Maintenance is part of measurement, not evidence that the metric failed. The failure is letting those changes happen invisibly while preserving a smooth chart.

The first credible report does not require a large software purchase. It requires a versioned question set, a written protocol, saved answers, an explicit rubric, and a calculation another person can reproduce. Run that benchmark before buying an AI visibility dashboard. Then evaluate any tool against your instrument: can it preserve the required surfaces, evidence, exclusions, weights, and history, or does it ask you to accept a proprietary score?

A board should reject the proprietary score. It should accept a modest metric with a clear boundary, especially when that metric exposes a decision: which buyer segment lacks credible evidence, which factual defect creates risk, and which owner will fix it before the next run.

Frequently Asked Questions

What is share of model?

Share of model measures how often and how strongly a brand appears in AI-generated answers for a controlled set of relevant buyer questions. The exact definition must state whether it counts mentions, recommendations, citations, or a weighted combination.

Is share of model the same as share of voice?

No. Share of voice counts exposure on surfaces with observable units, while AI systems synthesize answers in which a mention can be positive, negative, prominent, incidental, or cited. Treat share of model as a related but separate measurement.

How do I measure brand visibility in ChatGPT?

Create a governed question set, start a clean conversation for every prompt, record the visible product mode and settings, and save every complete answer. Score presence, prominence, recommendation, and citation under a written rubric, then repeat runs on a schedule.

Can I compare ChatGPT, Perplexity, and Gemini directly?

You can compare defined surfaces on those services under the same prompt and collection protocol. You cannot honestly claim that one consumer chat, one research mode, and one API represent three entire platforms.

How many prompts do I need for share of model?

There is no universal minimum. Use enough questions to cover priority segments and buying stages without padding the set with near duplicates, then repeat the stable panel enough to observe run-to-run variation.

Should AI citations count more than brand mentions?

Keep citations and mentions separate first. A weight can support an internal priority, but it should stay visible because it reflects your strategy rather than an objective property of AI answers.

How often should share of model be measured?

Run collection often enough to distinguish a trend from output variation, and review prompt-level findings monthly if the category moves quickly. Quarterly board reporting is usually more sensible than asking directors to react to every monthly fluctuation.

Does a higher share of model cause more sales?

The metric alone cannot establish causation. Compare it with self-reported discovery, shortlist inclusion, sales-call evidence, referral data when available, and win-loss interviews before making a commercial claim.

What should a share of model board report include?

Show weighted presence and preference, priority segments, the prior period, sample details, variability, material factual defects, and one decision management wants. Put the full competitor table, scoring rubric, and protocol in an appendix.

Do I need an AI visibility tool to start?

No. A versioned question set, saved answers, a transparent worksheet, and two reviewers on a quality sample are enough for a credible first benchmark. Buy software only after you know which evidence and controls it must preserve.

Related Posts