
Your company appears in an AI-generated shortlist on Monday and disappears when someone repeats the question on Wednesday. Another answer links to your website but describes an integration you do not offer. Meanwhile, analytics records a handful of visits from an AI platform.
These observations describe different outcomes. Combining them into one “AI visibility score” can hide the information your team needs.
A useful Generative Engine Optimization (GEO) measurement framework tracks four things separately: brand mentions, supporting citations, factual accuracy and identifiable referral traffic. It also records how each observation was collected.
Define what counts before collecting results
The definitions below are a proposed internal reporting method, not an industry standard or a platform ranking system.
- Brand mention: The answer explicitly names your company or an agreed product name. Count the brand once per answer, even if it appears several times.
- Owned-site citation: The answer includes a source link to a domain your company controls. Count this separately from a mention.
- Third-party brand coverage: A cited external page discusses your company. Record whether the answer itself names you or only the linked source does.
- Accuracy: Checkable statements about your company match current, approved evidence.
- Identifiable AI referral: A measured website visit carries a source or campaign signal associated with an AI platform. This does not capture every AI-influenced buyer.
Also label the context of a mention: recommendation, neutral description, comparison or negative assessment. Appearing in an answer is not automatically an endorsement.
Why one prompt cannot establish your market visibility
A 2026 critical survey of GEO research describes inconsistent metrics and evidence standards across the field. Within the studies it reviewed, it found no established technique with a stable, long-term, cross-platform causal effect on organic discovery or downstream behavior. This is an emerging research assessment, not official Google guidance.
A separate comparative research preprint reports differences across AI search services, languages and question wording. Its findings describe the systems and experiments studied, not a permanent rule for every platform.
The practical implication is to report your test conditions. “Mentioned in 12 of 40 sampled answers” is interpretable. “We own 30% of AI search” is not supported by that sample.
Build a question set around buying decisions
Start with questions from discovery calls, support conversations, sales objections and on-site searches. Remove personal or confidential information before using them in public tools.
For an enterprise workflow platform, a starter set might include:
- “Which workflow tools suit a distributed operations team?”
- “What should we check before choosing a platform with Salesforce integration?”
- “Which options support the procurement requirements of a UK enterprise?”
- “How does [your product] compare with [competitor] for approval workflows?”
- “Does [your product] support self-hosting?”
Keep unbranded discovery questions separate from branded evaluation questions. A prompt that names your product is useful for checking accuracy, but an answer repeating that name is weak evidence of spontaneous discovery.
Choose a fixed core set for comparisons over time. Keep newly discovered questions in an exploratory set until you deliberately revise the baseline.
Make the observations repeatable
A manageable pilot could use 20 buyer questions, three selected AI experiences and three separate runs per question: 180 observations for one market-language segment. This is a workload example, not a statistically validated sample size.
- Name the exact experience. Record the platform, model or mode when visible, and whether web search was used. Do not pool an API response with a consumer search interface as though they were identical.
- Keep conditions consistent. Use fresh conversations and record account personalization, requested geography, actual location settings where known, and language. A prompt saying “UK” does not prove that the platform localized the session.
- Spread observations across dates. Use a planned schedule instead of repeatedly regenerating until you obtain a favourable answer.
- Save the evidence. Retain the full answer, source URLs, timestamp and an internal screenshot or permitted export.
- Record incomplete observations. Log errors and unavailable features separately. If an answer does not recommend any vendors, retain that outcome rather than silently removing it.
Report completion coverage alongside visibility. If you planned 180 observations but completed 150, disclose both numbers and the reasons for the missing 30.
Use a tracking table that preserves the evidence
Illustrative template only: The rows below are fictional and do not report tests of any real brand or platform.
| Observation | Test conditions | Mention and citation | Accuracy finding | Next action |
|---|---|---|---|---|
| Q01, run 1 | Platform A, search enabled; US English | Brand absent; no owned-site citation | No brand claims to assess | Retain as a discovery baseline |
| Q01, run 2 | Platform A, same settings | Brand named; pricing page cited | One correct claim; one outdated price | Check cited pricing and cached references |
| Q07, run 1 | Platform B; UK English | Brand named; external review cited | Integration claim cannot be verified | Send claim and source to product owner |
In your working version, add the exact prompt, timestamp, answer capture, cited URLs and reviewer. Keep actual traffic in a separate analytics table; a test answer does not identify a real visitor.
Calculate rates with explicit denominators
- Mention rate: Completed, in-scope answers naming your brand divided by all completed, in-scope answers, multiplied by 100.
- Owned-site citation rate: Completed, in-scope answers linking to your domain divided by the same answer total, multiplied by 100.
- Verified-claim accuracy: Correct factual brand claims divided by all assessed, checkable brand claims, multiplied by 100. Show incorrect and unverifiable claims separately.
If no brand claims were available to assess, report accuracy as “not applicable,” not 100%. If an answer repeats the same claim, deduplicate it within that answer using a consistent rule.
Example: Across 60 completed unbranded observations on one platform, your company appears in 18 and receives an owned-site citation in nine. Mention rate is 30%; citation rate is 15%. Neither percentage estimates the share of all buyers who saw your company.
Show results by platform, market, language and question group. Repeated answers to the same question are related observations, not 60 independent buyers. Avoid presenting small movements as statistically meaningful without an appropriate analysis.
For competitor comparisons, calculate the same mention rate for each named competitor. Several brands can appear in one answer, so their rates need not sum to 100%.
Check the claims and the citations
Maintain a dated fact sheet covering product names, features, integrations, pricing, target customers and deployment options. Have product or commercial owners approve it.
For each material claim, open the cited source and ask whether it actually supports the statement. A citation beside a sentence is not enough. Record a broken link, an outdated source and an unsupported claim as different issues.
Prioritize inaccuracies that could affect a buying decision, such as a nonexistent capability or incorrect deployment option. Correct owned information where necessary and request factual corrections from external publishers when appropriate. Recheck later; an edit does not guarantee that an AI answer will change.
Connect referrals to business outcomes without inventing attribution
In GA4’s Traffic acquisition report, inspect session source and source/medium values. Build your AI-source grouping from observed, validated values rather than assuming every platform uses the same referral format.
OpenAI states that ChatGPT search referral URLs include utm_source=chatgpt.com. Its publisher FAQ explains this tracking signal. Test whether your redirects and analytics setup preserve it.
Compare identifiable referrals with engagement, completed demo requests and sales-qualified opportunities. Keep first-touch, session-level and assisted attribution definitions explicit; do not combine them into one lead total.
Google requires separate treatment. Its generative AI performance report provides impression data for AI Overviews and AI Mode. That visibility report does not establish which individual Google organic session or CRM opportunity came from an AI answer.
Use optional buyer self-reporting to add context. Record “the prospect said they used an AI assistant” separately from a technically identified referral. Neither observation proves that your latest content change caused the sale.
Turn the monthly review into specific actions
Keep four reporting sections: discovery mentions, citations, factual accuracy and business outcomes. Add test coverage and methodology changes so readers can judge the comparison.
- Frequent factual errors: Assign a product-content owner to investigate the claims and sources.
- Useful mentions without owned citations: Inspect which external sources buyers are being shown.
- Referrals without qualified enquiries: Review landing-page relevance and the next action offered.
- Improvement on one platform only: Investigate it within that platform before making a cross-platform claim.
For support with the content and search foundations behind this work, explore SEO Jetty’s GEO and SEO services.
Frequently asked questions
1. What is a good AI brand mention rate?
There is no universal threshold for this method. A rate depends on question selection, competitors, markets and platform settings. Establish your own baseline and compare like-for-like periods, with the observation count visible.
2. Can a small team start without a paid monitoring platform?
Yes. A spreadsheet, saved answers and a narrow question set can support a pilot. Paid tooling becomes useful when collection and review exceed team capacity; it does not remove the need for clear definitions and human checks.
3. What should we ask an AI visibility tool vendor?
Ask which interfaces it tests, how it selects prompts, whether it retains full answers, how it handles failed runs and how its scores are calculated. Confirm that you can export the underlying observations and distinguish simulated tests from actual audience data.
4. How should we handle a brand name shared by other companies?
Create an entity-matching rule using product names, website domains and business context. Mark ambiguous mentions for review. Counting every appearance of a common name can materially inflate visibility.
5. Should we translate the same prompts for every European market?
Use equivalent buying tasks, then have a fluent market reviewer adapt terminology and context. Keep language and country separate in the log. Compare within each segment before combining results.
6. Can AI grade the answers automatically?
It can assist with extracting names, links and potential inaccuracies. Have a person review disputed matches and material product claims. Periodically check a sample of automated labels against the original answers.
7. Should negative mentions be removed from the report?
No. Retain them in the mention count and label their context separately. Removing unfavourable answers biases the result and hides issues that product, communications or customer-success teams may need to address.
8. How should we handle a platform or model update?
Annotate the date and settings change. Where possible, collect an overlapping comparison before switching. If the experience changes substantially, start a new baseline instead of presenting the old and new scores as directly comparable.
9. What should we do when a citation links to an outdated third-party page?
Confirm the discrepancy against approved facts, contact the publisher with evidence and log the request. Monitor the source and subsequent answers separately. Correcting the source does not guarantee immediate changes across AI services.
10. How much evidence is needed before increasing GEO investment?
Look for sustained observations on relevant buyer questions, reliable brand descriptions and useful commercial signals. Set an investment review period appropriate to your sales cycle. A single favourable answer or a short-lived score increase is insufficient on its own.