Research report 002 / Version 1.0

Measuring AEO.

A practical framework for AI search visibility, citation, recommendation, and preference.

Jesse Killian580 Digital Infrastructure ResearchAugust 10, 2026

The measurement discipline

Three rules before the score.

580 Digital proposed metric

INSLRIntent-Normalized Shortlist Rate
01

Repeat

Treat generated answers as observations from a variable system, not fixed rankings.

02

Normalize

Measure buyer intents first so arbitrary prompt counts do not decide the result.

03

Separate

Keep each AI product and each measurement layer independently interpretable.

Abstract

Answer Engine Optimization is often measured with mentions, citations, share of voice, position, sentiment, referral traffic, or a single visibility score. Those measurements are useful, but they do not describe the same outcome.

A system can retrieve a page without citing it. It can cite a page without naming the brand. It can name a brand without recommending it. It can recommend a brand without preferring it. A recommendation can also influence a buyer without producing a trackable referral click.

This paper proposes the 580 Digital AEO Measurement Framework, an eight-layer model that keeps technical eligibility, retrieval, citation, information use, brand presence, recommendation, preference, and business outcome separate. It retains established measures such as Brand Presence Rate and Owned Citation Rate. It also introduces Intent-Normalized Shortlist Rate (INSLR), a proposed operational metric for measuring how consistently one AI product includes a brand across a predefined set of buyer intents without allowing an arbitrarily large number of prompt variations to dominate the result.

The framework treats generative answers as repeated observations from a variable system. It reports each AI product independently and refuses to turn unknown internal behavior into a zero. INSLR does not estimate market demand, total audience exposure, traffic, leads, or revenue. Its purpose is narrower: to make recommendation measurement more transparent, reproducible, and honest.

What this paper contributes

This paper makes four contributions:

  1. It separates eight observable and partially observable layers that are often collapsed into “AI visibility.”
  2. It defines a recommendation profile that distinguishes brand presence, shortlist inclusion, and top recommendation.
  3. It proposes INSLR as an intent-level measure of recommendation coverage.
  4. It provides a reproducible starting protocol, coding rules, a worked example, and explicit limitations.

It extends the practical Found, Matched, Proven, Shortlisted model introduced in The 580 Digital Local AI Shortlist Report by defining how the visibility and recommendation outcomes should be recorded.

The central argument is simple:

AEO requires a measurement system, not one AI visibility number.

1. The measurement problem

Traditional search measurement is built around familiar objects: queries, positions, impressions, clicks, and conversions.

Generative search changes the object being measured.

The foundational GEO paper presented at KDD 2024 treated visibility inside generated answers as a distinct optimization problem and proposed measures for source exposure within those responses.1 Later research has described generative search as a stochastic and partially observable pipeline rather than one ranking event.2

The practical distinction is this:

Visibility is not citation. Citation is not recommendation. Recommendation is not preference. Preference is not business impact.

If those outcomes are rolled into one score, the number may be easy to sell but difficult to interpret.

2. What is already measured

Commercial platforms already track useful AI visibility signals. Peec documents visibility, share of voice, sentiment, and average position.3 Ahrefs Brand Radar reports mentions, citations, estimated impressions, and AI Share of Voice.4 Semrush reports mentions, citations, position-aware share of voice, and sentiment.5

These are established measurement classes. They are not 580 Digital inventions.

Measurement Question answered
Brand presence or visibility Does the brand appear in the measured response?
Brand mentions How often is the brand named?
Mention share of voice How often is it mentioned relative to selected competitors?
Citation rate Is the measured domain cited?
Citation share What portion of measured citations belongs to the domain?
Position or prominence Where does the brand appear in the response?
Sentiment How is the brand characterized?
AI referral traffic Does an identifiable AI source send visits?
AI referral conversion Do those attributable visits produce outcomes?

The problem is not that these measurements are wrong. The problem is that they are often treated as interchangeable.

3. What Google and Microsoft report

Native platform reporting now exposes two different views of AI visibility.

Google introduced dedicated generative AI performance reports in Search Console in June 2026. The reports show impressions, pages, countries, devices, and dates for supported generative features in Search and Discover.6 Google describes those numbers as visibility reporting. They do not show whether the site was named as a business, used to support a claim, recommended, or preferred.

Microsoft introduced AI Performance in Bing Webmaster Tools in February 2026. It reports total citations, average cited pages, grounding queries, page-level citation activity, and trends across supported Microsoft AI experiences.7 Microsoft explicitly states that citation totals do not indicate placement, authority, ranking, or the role a page played in a specific answer.

That limitation matters:

Being used as a source and being recommended as a choice are different outcomes.

First-party platform data should be used. It should not be asked to prove something it does not measure.

4. One answer is an observation, not a ranking

Generative answers vary.

A March 2026 statistical study repeated queries across Perplexity, OpenAI SearchGPT, and Gemini. It found substantial citation variation and showed that apparent differences between domains could fall within the normal noise of the measurement process.8 Another 2026 paper reached the same practical conclusion: AI visibility should be measured repeatedly and treated as a distribution rather than a one-time result.9

This creates the first requirement of the framework.

Principle 1: Repeat the measurement

A single answer is one recorded observation. It is not a fixed AI ranking.

When an AEO metric is presented as evidence of system behavior, the report should disclose the number of observations and the dates on which they were collected. Where sample size permits, the report should include an interval or another honest expression of uncertainty rather than presenting a point estimate as exact.

5. One prompt is not one intent

Repeating the same prompt only measures one phrasing.

A 2026 study of production retrieval-augmented recommendation systems found that paraphrases of the same buying intent produced far less recommendation-set overlap than repeated runs of the exact same prompt.10 The wording selected by the researcher can therefore become a major source of variance.

Consider these questions:

  • Who is the best marketing agency in Lawton?
  • What marketing company should a small business in Lawton hire?
  • Recommend a local company that handles websites and local advertising in Lawton.
  • Who should I call to improve local search visibility in Lawton?

They are not identical requests. Some contain different constraints. A measurement protocol should not pretend they are all the same. It also should not automatically treat every literal prompt string as an independent unit of market demand.

Principle 2: Define intent before prompts

The primary sampling unit for recommendation measurement should be a declared buyer-intent family. Prompt formulations are observations of that intent, not independent estimates of demand.

Intent families must be defined before the results are scored. If researchers add, remove, or split intents after seeing the results, that decision must be disclosed.

6. Intent-Normalized Shortlist Rate

Classification: 580 Digital proposed operational metric

For a brand b, one AI engine or product e, and one buyer-intent family k:

  • p identifies a prompt formulation within the intent.
  • r identifies an independent repeated observation.
  • S = 1 when the brand is explicitly presented as a viable choice for the buyer's stated need.
  • S = 0 when it is not.

Per-intent shortlist rate

SLRb,e,k = (1 / PkRk) × Σp Σr Sb,e,k,p,r

Intent-Normalized Shortlist Rate gives each declared intent equal weight:

Intent-Normalized Shortlist Rate

INSLRb,e = (1 / K) × Σk=1K SLRb,e,k

An INSLR of 0.64 means:

Across the declared intent universe and measurement protocol, the brand was shortlisted in an average of 64 percent of observations per intent.

It does not mean that 64 percent of AI users saw the brand. It does not estimate prompt demand, market share, traffic, leads, or revenue.

INSLR measures recommendation coverage across a defined intent universe.

7. Why intent normalization matters

Suppose a benchmark contains three buyer intents. The number of prompt formulations is not balanced:

Buyer intent Prompt formulations Repeats per prompt Shortlisted observations Intent SLR
General local marketing 2 3 3 of 6 50.0%
Local SEO 5 3 6 of 15 40.0%
Shared direct mail 1 3 3 of 3 100.0%

A raw average across all 24 observations is 50 percent. That number gives the Local SEO intent five times as much influence as Shared Direct Mail because the researcher happened to test more phrasings.

INSLR first calculates each intent's rate, then averages the three intent rates:

Worked example

INSLR = (50.0% + 40.0% + 100.0%) / 3 = 63.3%

Neither number estimates real buyer demand. The difference is that INSLR makes its weighting rule explicit and prevents prompt-count imbalance from silently deciding the result.

Version 1.0 does not use demand weighting because reliable population-level distributions of AI buyer prompts are generally unavailable. A future demand-weighted version would require defensible exposure data, not an invented multiplier.

8. Engines must remain independent

Principle 3: Report one product at a time

Results from different AI products should be reported separately unless a documented exposure model justifies combining them.

The framework permits reporting:

  • ChatGPT INSLR: 61%
  • Gemini INSLR: 43%
  • Perplexity INSLR: 72%

It does not recommend averaging those results into an “overall AI score.”

AI products can differ in retrieval, source selection, model configuration, location handling, personalization, interface, and response generation. Current research documents meaningful differences across products.2 Without a defensible model of how the target audience uses each product, a cross-engine average adds assumptions that the observations do not contain.

9. Recommendation should be reported as a profile

A brand appearing in an answer is not the same as being recommended. A recommendation is not always a preference.

The framework records three answer-level states:

Brand presence

B = 1 when the brand is named in the response as a relevant entity.

Shortlist inclusion

S = 1 when the brand is explicitly presented as a viable choice for the buyer's stated need.

Top recommendation

T = 1 when the system explicitly identifies the brand as its preferred, leading, first, or best-fit choice for that request.

These states can be reported as a Recommendation Profile:

State Example result
Brand Presence Rate 81%
Shortlist Rate 57%
Top Recommendation Rate 24%

This preserves information without inventing numerical weights for ordinal states.

Minimum coding rules

  • A brand listed only in visible source links does not count as brand presence in the answer text.
  • A passing mention, comparison, warning, or historical reference does not count as shortlist inclusion.
  • A directory page, article, or source domain is not a business recommendation unless the response presents it that way.
  • “One option” or “worth considering” can count as shortlist inclusion when it answers the buyer's stated need.
  • Ordered placement alone does not count as top recommendation unless the wording or interface clearly communicates preference.
  • Refusals, tool errors, and empty responses are excluded from the denominator and reported separately.
  • Ambiguous cases should be marked for review rather than forced into a favorable classification.

If more than one person codes responses, the report should publish the rubric and measure agreement on a shared sample.

10. Citation is a separate dimension

Citation metrics measure source attribution. Recommendation metrics measure brand selection.

An emerging 2026 framework distinguishes citation selection from citation absorption. A page may be selected as a source while contributing little to the answer's language, facts, evidence, or structure.11 CiteEval, published at ACL 2025, also evaluates citation quality in the context of the question, generated answer, cited source, and broader retrieval set rather than treating the presence of a link as sufficient.12

These distinctions support a more careful analytical sequence:

retrieval → citation → information use → brand presence → recommendation → preference

That sequence is not a disclosed causal funnel. It is a set of categories that should not be collapsed.

11. The 580 Digital AEO Measurement Stack

The complete framework contains eight layers.

Layer Measurement question Typical evidence
1. Technical eligibility Can the information be crawled, indexed, or otherwise accessed? Index status, robots controls, platform eligibility
2. Retrieval Did the information enter the available context? Product source panel, logs, grounding-query data where available
3. Citation Was a source explicitly attributed? Visible citations, Bing citation reporting
4. Information use Did the source materially support the answer? Claim-to-source comparison, citation-quality review
5. Brand presence Was the brand named? Recorded response coding
6. Recommendation Was the brand presented as a viable choice? Shortlist coding, INSLR
7. Preference Was the brand explicitly favored? Top Recommendation Rate
8. Business outcome Did the interaction affect behavior? Referral traffic, branded search, calls, leads, sales, surveys

The stack is not a deterministic funnel. A system may rely on owned content, third-party evidence, model knowledge, or a combination. Some internal stages will remain unobservable.

12. Core metrics

Version 1.0 recommends four primary reporting metrics.

Brand Presence Rate

Classification: established measurement class

BPR = responses containing the brand / eligible measured responses

Owned Citation Rate

Classification: established measurement class

OCR = responses citing the owned domain / eligible measured responses

Intent-Normalized Shortlist Rate

Classification: 580 Digital proposed operational metric

INSLR = average of the declared intent-level shortlist rates

Top Recommendation Rate

Classification: descriptive recommendation metric

TRR = responses explicitly preferring the brand / eligible measured responses

No universal composite score is required.

13. Reproducible measurement protocol

Every AEO benchmark should disclose enough information for another researcher to understand what was measured and repeat the process.

Before collection

  1. Name the exact AI product and surface being measured.
  2. Record the relevant model or product version when visible.
  3. Define the buyer-intent universe before viewing the results.
  4. Document how each intent was selected.
  5. Write the prompt formulations assigned to each intent.
  6. Fix the number of repeated observations per prompt.
  7. Declare location, language, account state, memory state, and persona conditioning.
  8. Publish the coding rules for presence, shortlist inclusion, preference, and citations.

During collection

  1. Preserve the complete response, prompt, timestamp, visible sources, and follow-up context where platform rules permit.
  2. Record errors, refusals, and unavailable observations separately.
  3. Avoid changing the intent set or rubric in response to favorable or unfavorable results.
  4. Keep results from different AI products separate.

Reporting

  1. Publish the numerator and denominator behind every rate.
  2. Report collection dates and sample size beside the result.
  3. Distinguish observed, absent, and unobservable states.
  4. Show uncertainty when the sample supports it.
  5. Separate first-party platform data, researcher-coded observations, and business outcomes.
  6. Preserve raw observations long enough to recalculate results if the rubric changes.

A useful pilot can begin with three intent families, three prompt formulations per intent, and three repeated observations per prompt. That 27-observation design is a starting point, not a claim of statistical sufficiency. Sample requirements depend on the product, category, observed variance, and decision being made.8

14. Unknown is not zero

Commercial generative systems are only partially observable.

For an internal event such as retrieval, the framework permits:

R ∈ { 1 observed, 0 observed not to occur, ? unavailable }

If the product does not expose whether a page was retrieved, the correct value is unknown. Coding that state as zero creates false precision.

The same rule applies to downstream influence. A missing referral click does not prove that an AI answer had no effect on branded search, a later direct visit, a phone call, or an offline purchase.

15. What the framework implies for AEO strategy

Measurement does not create authority. It tells us which part of the system needs work.

Google's current guidance says the familiar SEO foundation still matters and emphasizes valuable, unique, expert-led, non-commodity information over special AEO tricks.13 Microsoft recommends clear, complete, current content with evidence supporting important claims.7

That supports a practical distinction between two kinds of authority.

Brand authority

The available evidence helps a system understand that the entity exists, belongs to the category, has real-world credibility, and is a plausible recommendation.

Information authority

The organization's content contributes original evidence, useful definitions, concrete facts, comparisons, procedures, examples, data, or first-hand expertise that can support an answer.

Technical accessibility allows systems to discover and use both.

Conceptual model, not a numerical equation

AEO strength ≈ brand authority + information authority + technical accessibility

The strategic goal is not to make an AI system see a website. It is to build an entity worth recommending and an information resource worth using.

16. Limitations

The 580 Digital AEO Measurement Framework Version 1.0 is a practitioner framework. It has not been validated as a predictive model.

  1. INSLR has not been validated as a predictor of traffic, leads, revenue, or population-level exposure.
  2. Equal intent weighting measures coverage across the selected intent universe. It does not estimate how frequently real buyers express those intents.
  3. No cross-engine aggregate is defined because the framework does not assume equivalent products or equal audience exposure.
  4. Recommendation coding can require human judgment. A formal rubric and inter-rater testing would strengthen empirical work.
  5. Retrieval, personalization, model priors, and internal source-selection behavior are not always observable.
  6. Products change. Every benchmark should identify its collection dates and the product surface measured.
  7. Several relevant 2026 studies remain preprints. They are emerging evidence, not settled consensus.
  8. The hypothetical worked example explains the calculation. It is not a reported performance result for 580 Digital or a client.

17. Research agenda

This framework creates testable questions:

  • How many paraphrases are required to estimate an intent-level shortlist rate reliably?
  • How stable is INSLR across repeated collection windows?
  • Do natural buyer prompts differ systematically from prompts derived from search keywords?
  • Does an increase in owned citation rate precede an increase in shortlist rate?
  • Does citation absorption predict recommendation better than citation frequency?
  • How do brand authority and information authority independently affect recommendation behavior?
  • Which observable AI signals predict branded search, direct visits, calls, leads, or revenue?
  • Can defensible demand data support a useful intent-weighted version of INSLR?

Those questions require recorded data, repeated tests, and honest negative results.

18. Conclusion

AEO measurement should not recreate conventional rank tracking inside a chatbot.

Generated answers vary. Prompt wording matters. Products behave differently. Citation does not prove recommendation. Recommendation does not prove preference. Referral traffic captures only one possible business outcome.

The 580 Digital framework therefore measures distinct outcomes independently:

  • Can the system access the information?
  • Does it use the information as evidence?
  • Does the brand appear?
  • Does the brand make the shortlist?
  • Does the brand become the preferred recommendation?
  • Does any of that affect business behavior?

The commercial objective may be simple: build a credible business and publish information worth retrieving, trusting, using, and recommending.

Measuring whether that is happening requires more precision than a screenshot or one visibility score can provide.

Suggested citation

Killian, Jesse. “Measuring AEO: A Practical Framework for AI Search Visibility, Citation, and Recommendation.” 580 Digital Infrastructure Research, Version 1.0, August 2026. https://580di.com/research/measuring-aeo

References

Official platform and product documentation

6Maoz, Hillel, and Moshe Samet. “Introducing Search Generative AI Performance Reports in Search Console.” Google Search Central, June 3, 2026. Official description of the dedicated generative AI performance reports, rollout, dimensions, and impression data.

13Google Search Central. “Google's Guide to Optimizing for Generative AI Features on Google Search.” Accessed August 10, 2026. Official guidance on technical eligibility, SEO foundations, unique non-commodity content, and unsupported AEO or GEO tactics.

7Madhavan, Krishna; Meenaz Merchant; Fabrice Canel; and Saral Nigam. “Introducing AI Performance in Bing Webmaster Tools Public Preview.” Microsoft Bing Blogs, February 10, 2026. Official definitions and stated limitations for total citations, average cited pages, grounding queries, and page-level citation activity.

3Peec AI. “Understanding Your Performance.” Accessed August 10, 2026. Product documentation for visibility, share of voice, sentiment, and position.

4Tan, Constance. “AI Visibility Metrics.” Ahrefs Help Center, June 26, 2026. Product definitions for mentions, citations, impressions, and AI Share of Voice.

5Loktionova, Margarita. “How to Measure AI Share of Voice Using Semrush.” Semrush, July 17, 2026. Product explanation of mentions, citations, and position-aware AI Share of Voice.

Published and preprint research

1Aggarwal, Pranjal; Vishvak Murahari; Tanmay Rajpurohit; Ashwin Kalyan; Karthik Narasimhan; and Ameet Deshpande. “GEO: Generative Engine Optimization.” Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2024. Foundational work defining GEO as an optimization problem and proposing visibility metrics for generated responses.

8Sielinski, Ronald. “Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement.” arXiv:2603.08924, March 2026. Preprint. Repeated-sampling study across Perplexity, SearchGPT, and Gemini using bootstrap confidence intervals.

9Schulte, Julius; Malte Bleeker; and Philipp Kaufmann. “Don't Measure Once: Measuring Visibility in AI Search (GEO).” arXiv:2604.07585, April 2026. Preprint. Argues for repeated measurements and distributional reporting.

10Jack, Will; Noah Lehman; Keller Maloney; and Sarah Xu. “Paraphrase Brittleness in Production Retrieval-Augmented Commercial Recommendation: Reproducibility Below the Rerun-Stability Baseline.” arXiv:2605.27440, May 2026. Preprint. Studies recommendation variation across paraphrases and same-prompt reruns.

11Zhang, Kai; He Xinyue; and Yao Jingang. “From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms.” arXiv:2604.25707, April 2026. Preprint. Distinguishes source selection from answer-level information absorption.

12Xu, Yumo; Peng Qi; Jifan Chen; Kunlun Liu; Rujun Han; Lan Liu; Bonan Min; Vittorio Castelli; Arshit Gupta; and Zhiguo Wang. “CiteEval: Principle-Driven Citation Evaluation for Source Attribution.” Proceedings of ACL 2025, pages 32759-32778. DOI: 10.18653/v1/2025.acl-long.1574.

2Martinez, Olivier. “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026).” arXiv:2607.14035, July 2026. Preprint. Reviews 45 studies and proposes a multistage, partially observable model of generative search.

Measure the right outcome.

The framework is open for scrutiny. If you use it, publish the intent set, observation count, coding rules, product, dates, and limitations beside the result.