Protocol / One standard for proof you can audit

The iLuvSEO AI Search Evidence Protocol

A practical way to separate official documentation, live site data, controlled tests, interpretation, and business results in AI search work.

When I review AI search visibility, I do not treat a citation screenshot, a crawler hit, and a conversion as equal proof. This protocol shows what each record can support, how it was collected, when it was checked, and where the claim stops.

Rules for every review

  • Start with the business decision. Choose the metric after that.
  • Keep eligibility, exposure, citation, visit, assisted conversion, and revenue separate.
  • Save the exact test conditions. Repeat the test before calling it a pattern.
  • Use primary documentation for platform behavior. Label interpretation as interpretation.
  • Put the limits and change history next to the findings.

Evidence matrix

Match every outcome to its proof surface

One artifact cannot prove every stage of AI search performance. Use the right proof for the question.

Evidence required for each outcome and what that evidence does not prove.
OutcomePreferred evidenceDoes not prove
Crawl and index eligibilityCheck the HTTP response, headers, robots and CDN behavior, rendered content, canonical state, and the relevant webmaster inspection tool.A route opening successfully in one human browser.
Google generative-AI exposureUse the Search Console generative-AI performance report when the property has access. Keep Google's dimensions intact.An exact prompt ranking or a complete record of every generated answer.
Bing citation activityUse Bing Webmaster Tools AI Performance citations, cited pages, trends, and sampled grounding phrases. Keep the report's stated limits attached.Authority, answer placement, endorsement, or revenue.
ChatGPT-referred visitUse analytics referral data with the documented ChatGPT source parameter. Connect it to the landing page and the next action.A GPTBot or OAI-SearchBot request in a server log.
Agent task completionRepeat the task. Verify the final interface state or business receipt outside the agent.The agent loading a page, filling one field, or beginning an action.

01

Start with the decision, not the dashboard

Start with the decision. Are we fixing crawler access, strengthening a source page, changing a product feed, or testing a new discovery surface? Write that decision down before opening a dashboard.

Use a metric only if it can change the decision. Citation volume can help diagnose source reuse. Qualified visits and revenue answer a commercial question. Mixing them produces a neat report and a bad recommendation.

  • Name the decision owner and deadline.
  • Choose one primary measure. Keep diagnostic measures secondary.
  • Lock the comparison window, market, and pages in scope.

02

Separate the evidence layers

Keep six layers separate: technical eligibility, platform exposure, citation or brand mention, referred visit, assisted conversion, and completed business outcome. Improvement in one layer does not prove improvement in the next.

Platform documentation tells us how the platform says a feature works. Property data shows recorded activity. A controlled prompt test gives us one bounded observation. Any connection we draw is inference unless the system exposes direct proof.

  • Platform statement: link the primary documentation and review date.
  • Property data: save its dimensions, filters, and reporting window.
  • Controlled observation: save the exact prompt and test environment.
  • Inference or hypothesis: mark it separately from observed fact.

03

Record the environment behind every observation

Generated answers can change by surface, time, locale, device, account state, personalization, and prompt wording. Record the exact prompt, timestamp and time zone, country and language, device class, signed-in state, and the product or model label shown in the interface.

Save the answer, visible sources, and destination URLs. Also record whether the brand was cited, described accurately, recommended, or omitted. Those are different outcomes. Do not merge them into one score without showing the formula.

  • Keep verbatim prompts and meaningful paraphrases.
  • Preserve visible source URLs rather than domain names alone.
  • Record factual errors and uncertainty, not only favorable mentions.

04

Repeat before describing a pattern

One response is one observation. It is not a stable rank. Set the repeat count and collection dates before looking at the results. Hold controlled variables fixed. Report counts, ranges, and recurrence instead of showcasing the best screenshot.

When testing a prompt family, save the original request and every declared paraphrase. Keep each engine and product surface separate. They expose different data and can use different retrieval or answer systems.

  • Declare the sample plan before reviewing the result.
  • Report missing answers and failed runs with the successful runs.
  • Do not convert a citation count into an invented cross-platform rank.

05

Connect visibility to accountable outcomes

AI visibility can influence discovery without producing an immediate click. Track that, but do not stop there. Keep referral sessions, qualified actions, assisted conversions, sales, and revenue separate.

If referral data exists, connect the session to the landing page and next action. If it does not, say so. Branded demand, direct visits, and sales changes add context. They do not automatically belong to AI visibility.

06

Verify agent usability as an end-to-end task

Crawler access is only the first gate. Browser agents can use screenshots, HTML, and accessibility information. Test the layout, semantic controls, labels, visible state changes, and final completion state.

For actions that matter, also test narrow permissions, explicit approval, argument validation, and an independent receipt. Reaching the final screen does not prove that a reservation, form submission, purchase, or publication completed.

  • Test the same meaningful task more than once.
  • Capture the failure point and interface state, not only pass or fail.
  • Confirm completed actions against the authoritative business system.

Method

Methodology

Use this workflow for a one-page diagnosis, a platform baseline, or a long-running benchmark. The scale changes. The evidence chain does not.

  1. Frame the decision

    Name the owner, scope, decision date, primary measure, and the action each possible result would trigger.

  2. Establish source authority

    Use current primary platform documentation for platform-specific behavior and record when each source was reviewed.

  3. Validate delivery

    Check the real response, crawl controls, rendered information, canonical state, and consistency between visible facts and machine-readable data.

  4. Collect controlled observations

    Preserve prompts, conditions, answers, source URLs, errors, and repeated samples without discarding inconvenient results.

  5. Reconcile measurement layers

    Keep platform exposure, citation, referral, conversion, and revenue distinct before interpreting how they relate.

  6. Publish limits and changes

    State what the evidence cannot establish, add the verification date, and update the change log when the method or source guidance changes.

Limits

What this does not prove

This protocol makes each claim easier to audit. It does not reveal private ranking or retrieval systems. It cannot guarantee crawling, indexing, citation, recommendation, traffic, or revenue.

Platform reports measure different things. Their availability, aggregation, dimensions, and definitions can change. Each finding applies only to the recorded conditions and dates.

  • Crawler access establishes possibility, not inclusion or preference.
  • Repeated prompt observations describe the sampled surfaces, not every user experience.
  • A citation does not by itself establish endorsement, placement, factual influence, or commercial value.
  • A before-and-after change does not establish causation when other conditions also changed.
  • A successful agent task does not prove that every agent, device, account, or user journey will succeed.

Primary sources

Primary documentation reviewed

  1. Optimizing your website for generative AI features on Google SearchGoogle Search Central, reviewed Opens in a new tab
  2. Introducing Search Generative AI performance reports in Search ConsoleGoogle Search Central, reviewed Opens in a new tab
  3. Introducing AI Performance in Bing Webmaster Tools Public PreviewMicrosoft Bing, reviewed Opens in a new tab
  4. Publishers and Developers FAQOpenAI Help Center, reviewed Opens in a new tab
  5. Build agent-friendly websitesweb.dev, reviewed Opens in a new tab

Record

Change log

  1. Published the first version. It defines the evidence layers, collection workflow, reporting limits, and reviewed primary sources.