Product

Topic Engine Content Engine AEO Rank Research-Driven Content Autopilot Visibility Tracking Website & Migration Content Factory Audits Rankings Pricing

Resources

Browse all resources → Case Studies Blog FAQ Use Cases Knowledge Base Research Docs

Run a free AEO audit twice: which half of the result should move

Two printed website audit reports compared side by side on a desk beside a laptop

Key Points

  • Readiness rows read a site's own files, such as robots.txt rules for GPTBot, ClaudeBot and PerplexityBot, schema and llms.txt, so they should match on a second run unless the site changed.
  • In 2025, a study by Oshen Davidson, first published by AirOps, found only 30% of brands kept their visibility from one run of an identical query to the very next.
  • Ethan Smith of Graphite, on Lenny's Podcast in 2025, described a test design of 200 questions, 100 left untouched and 100 changed, tracked before and after the change.
Three things site owners believe about a second audit. Myth or fact?
Call each one, then see how other readers called it.
1 A second run showing fewer AI mentions proves your fixes failed.
2 A readiness row that flips between runs with nothing deployed points to the checker, not the site.
3 Being the first citation in an AI answer wins it, as the top Google result does.
Two printed website audit reports compared side by side on a desk beside a laptop

Two runs of the same free AEO audit: the readiness rows should match, the visibility rows may not.

Quick Answer

Two runs of a free AEO audit should agree on readiness rows unless you deployed; AI-visibility rows move on their own and deserve reading as a rate across weekly runs.

Readiness rows read your own files: robots.txt rules for AI crawlers, schema markup, llms.txt and heading structure. A flip with nothing shipped points to the checker. Visibility rows sample live answers from ChatGPT, Perplexity and Gemini, so hold the prompt set and account state fixed before comparing. Any AEO firm worth hiring reports the two halves apart.

Did this answer your question?

Even seasoned practitioners concede that much AEO advice goes untested, so the second run of a free audit, not the first, tells you whether a change did anything.

It is worthy of notice how candid the field can be on this point. One AEO adviser, interviewed on a 2025 product podcast, allowed that maybe half of their own advice worked and urged listeners to run experiments rather than trust blog posts. The same adviser held that there is no premium version of rank tracking, and that the cheapest answer tracker which does the job is enough.

The sections below take the matter in order:

  • Which checks should repeat exactly on a second run, and what a flip with nothing deployed means.
  • Why the AI-visibility half moves when nothing on the site changed.
  • How to compare two runs without mistaking turnover for progress.

Even the readiness half needs keeping current. An April 2026 walkthrough reported that OAI-SearchBot had surpassed GPTBot as OpenAI's dominant crawler, so a robots.txt rule written with one bot in mind may no longer describe how ChatGPT reaches a site. The same walkthrough priced one major paid AEO tool at $900 per month. That is reason enough to learn what the free version shows before paying for more.

Owners weighing outside help also ask which companies specialize in answer engine optimization. I come to that question last, with a single test for any firm that answers it.

Run a free AEO audit twice and one half should move: the readiness checks that read your site hold still, while the sampled AI-visibility findings shift of their own accord.

A free AEO audit is a one-time scan that reads a site's crawler permissions, schema and content structure, and in some tools samples live answers from engines such as ChatGPT besides. The two halves obey different rules. One reads files that change only when you change them. The other does not.

The second half was described with some precision in a 2025 study by Oshen Davidson, published by AirOps: "every query is a fresh retrieval, every response is a fresh selection." It is remarkable how much hangs on the definition of a hit.

In that 2025 dataset, mentions without citations were 3× more common than citations without mentions. In the same 2025 work, a brand both cited and mentioned on its visible run was 40% more likely to resurface later than one only cited. A tool counting only links will report less presence than one that also counts names.

Practitioners in the AEO community advise running each prompt 5 or 10 times, since a single run of a non-deterministic model wanders considerably. Upon closer inquiry, a recorded miss sometimes proves to be an answer that named the product and not the company. Neither finding touches the readiness rows.

I take the view that the readiness rows deserve the strictness of an inspection, and the visibility rows the patience of a tally kept across many runs. Which rows belong to which half is the first matter to settle.

12-24 months Visibility Outlook

How AI answer audits will separate signal from noise

Forecasts on how audit tools, agencies and site owners will treat fixed site checks versus sampled AI answers over the next two years.

3 sources analyzed2 community discussions1 web source
A

What changes in AI answer measurement

Use each forecast to decide which audit numbers to act on now and which to re-sample before drawing conclusions.

61/100
Medium confidence 12-24 months

Audit reports will state exactly what counts as a hit, such as company name versus product names and citation versus mention. They will also separate unverified or unmatched results from failures, because definitional choices alone can swing a score.

Contrarian Take
61/100
Medium confidence 12-24 months

Buyers will move toward windowed measurement rather than more frequent checks. They will run multi-run series spread over weeks, judge presence by whether a brand reappears, and treat a one-day disappearance as noise rather than a loss.

Emerging, Not Established Practitioners already recommend running each prompt 5x or 10x. They also recommend keeping a fixed set of 20-50 prompts on a schedule and recording whether the answer changed. In one tracked run for a large consumer electronics brand, about half of the recorded misses were answers that named the brand's products rather than the company. Separately, the SearchD free audit already labels rows UNKNOWN, UNVERIFIED or NOT RUN instead of counting them as passes. The agency GrowthSpree reports answers shifting within 24 hours, sometimes between 9 am and 6 pm on the same prompt. A study of runs issued within a 23-day window found that ~57% of brands reappeared in a later, non-consecutive run.

B

Practitioner and study sources behind each call

Public practitioner threads and a citation-volatility study, with the line from each that a forecast relies on.

Source What it states Forecasts it backs
Are you actually tracking AI search visibility, or just checking it [Community / Forum] Commenter 2 recommends a fixed set of 20-50 prompts run on a regular schedule. Record brand mentions, citations, competitors, source types, and whether the answer changed. “Curious what people are actually doing in practice rather than what the tools say we *should* be doing.”
About half of the "misses" were not misses. The answers named the brand's products rather than the company, and the run counted only the company name.
Repeated runs replace single snapshots
Match rules become a reported setting
What are you actually using to audit your AEO/AI visibility right now? [Community / Forum] Commenter 4 recommends running each prompt 5x or 10x, because LLMs are non-deterministic and single runs "move around a lot.". “You need to measure on a decent sample size per prompt.”
Commenter 10 (GrowthSpree, a B2B SaaS marketing agency) reports that answers shift within 24 hours, "sometimes 9 am versus 6 pm on the same prompt.".
Repeated runs replace single snapshots
Daily drops will count for less
Measuring Citation Volatility in AI Search - Oshen Davidson [Web source] ~30% of brands stayed visible from one run to the very next run of the identical query. “The 'fixed #1 position' era is over.”
Mentions without citations were 3× more common than citations without mentions.
~57% of brands reappeared in a later, non-consecutive run.
Repeated runs replace single snapshots
Match rules become a reported setting
Daily drops will count for less
The sources behind the forecasts above: what each one states, and which forecasts lean on it.
C

What would make AI answers hold still

Scenarios in which AI answers stabilize or measurement standards arrive faster, weakening the case for repeated sampling.

Built-In Uncertainty

Weigh “Repeated runs replace single snapshots” more heavily than the rest, and keep an eye on “Daily drops will count for less” as the forecast least protected by current evidence.

  • The case for repeated sampling and reappearance rates would weaken if AI assistants began retrieving and citing sources by default on every query, or if answers to identical prompts stopped shifting within a day.
  • Either change would make single-run checks reliable again.
Methodology Each forecast is scored 0-100 from the public sources shown for it: how many there are and how authoritative they are.

Where can you run a free AEO audit worth running twice?

Run the Free AEO Readiness Audit from the team behind 26,577 real AI-visibility audits and 11,000+ domains scored across 15 sectors and 28 categories, then run it again after your next deploy.

The platform is engineered by a full-stack engineer and content infrastructure architect with 20 years of building enterprise systems. I previously founded an INC 5000 inductee (#84, 2015-2018). Bring your change log.

Which checks in a free AEO audit should repeat exactly on a second run?

Every check that reads your own files and markup should repeat exactly: crawler permissions in robots.txt, llms.txt presence, schema, heading structure and citability signals, unless someone changed the site between runs.

Before comparing a second report with the first, I would put four plain questions to each readiness row:

  1. Did the row read something on the site itself, such as robots.txt, llms.txt, sitemap.xml or a schema block?
  2. Was anything deployed, edited or republished between the two runs?
  3. If nothing changed, does the file still load at the same address when you open it in a browser?
  4. If the file loads and the row still flipped, has the checker simply erred?

GPTBot, ClaudeBot and PerplexityBot are either permitted in robots.txt or they are not, and one free checker shared on Reddit tests precisely that among the 18 AEO-specific checks its builder says it runs. A deterministic check is a test that reads the site's own code or content and returns the same verdict whenever those inputs are unchanged. The same checker, by its builder's account, also looks for FAQ, HowTo, Article and Organization schema, for heading hierarchy and answer-first paragraphs, and for author bios, dates and source links. Nothing in that list consults an answer engine. Each item is read from your pages, and a second reading of the same pages ought to agree with the first.

It is worthy of notice how regularly these families recur. Set 5 published walkthroughs of free audit tools side by side and the same groups appear in nearly every one: crawler access, structured data, content structure and signals of citability. One vendor's demonstration scores a page between 0 and 100 across SEO, AEO, schema, performance and content sections; another sorts its fixes into crawler accessibility, content quality and structured data; a third, run on its maker's own landing page, flagged missing H1 and H2 structure and an unclear definition of the service. A fourth builder claims to go beyond robots.txt into content length and copy, and reports lifting their own site to 81/100 after a bit of work. One exception deserves a note, since the first of those vendors describes its performance section as a real-time load test taken at the moment the report runs. I am inclined to think a speed reading of that sort belongs to neither half, and should be read on its own terms.

The common assumption is that any difference between two runs must reflect the site. In the fixed half, a flip without a deploy may just as well reflect the instrument. On that same Reddit checker, one commenter reported that it declared llms.txt missing on a site that had one, and did the same with sitemap.xml. A file is present or absent. When a yes-or-no row disagrees with itself across two runs, the fault lies with the tool rather than the page. Fix the instrument before you touch the page.

At AEO Content we undertake that a client will be named in AI answers in 90 days, or we work free until they are, and a commitment of that kind can only be judged against a site whose readiness rows hold still from one run to the next. Readers who wish to verify any single row by hand will find each check described in our AEO knowledge base. The other half of the report keeps no such discipline, and the cause lies not in your pages but in how the engines choose their sources on each separate occasion.

Why does the AI-visibility half of an audit move when nothing on the site changed?

Answer engines reselect their sources on every run, so a visibility sample can differ from last week's even when not one line of your site has changed.

In 2025, a study by Oshen Davidson, first published by AirOps, found that across more than 45,000 citations only 30% of brands kept their visibility from one run of an identical query to the very next. In that 2025 study, about 20% held across all 5 runs in a series. Yet in the same 2025 data, about 57% of brands reappeared in a later, non-consecutive run, so that most disappearances proved to be absences rather than losses. The study treats this drift as a structural property of how models select citations from a retrieval pool that varies from run to run, and it states the matter plainly: "Reappearance is the rule, not the exception." It is remarkable that the query never changed. Only the draw did. The runs all fell within a 23-day window on a single model, the consumer ChatGPT interface running GPT-4.1, so how far the turnover extends from one month to the next is a question the evidence leaves open.

Practitioners report the same restlessness at a finer grain. GrowthSpree, a B2B SaaS marketing agency, says in an r/aeo discussion that answers shift within 24 hours, "sometimes 9 am versus 6 pm on the same prompt." Another commenter in that thread enumerates the conditions that alter a reply to an identical question:

  • Session state: a logged-in account versus a fresh session.
  • Memory: switched on or switched off.
  • Timing: the hour at which the prompt is sent.

None of these is a property of your website. Each belongs to the moment of asking.

The engines also differ in how closely they lean on the ordinary web index. A study of thousands of questions, described in a 2025 interview, found ChatGPT's citations overlapping with Google results about 35% of the time, against about 70% for Perplexity, and practitioners who run formal tests note that answers vary noticeably even with no intervention at all. A report that pools several engines therefore pools several different rates of drift, each with its own habits and its own season of change.

Not every free audit samples live answers, and the distinction matters when reading a second run. One vendor's free grader, demonstrated in an April 2026 walkthrough, scores how ChatGPT, Perplexity and Gemini characterize a brand based on their training data. A reading drawn from training data answers a different question from a live sample, and ought not to be set against one as though the two were alike. A drop in a sampled row is weak evidence of anything. The sample moved; your site did not.

Owners who would rather follow the sampled half across several weeks than judge it from a single run can see how ongoing visibility tracking works. Which leaves a practical difficulty on the desk: two reports, one visibility figure lower than before, and the need to decide whether anything you did has caused it.

How do you compare two free AEO audit runs to separate real progress from turnover?

Diff the readiness rows against a dated change log, read visibility as a rate across repeated runs, and let booked calls settle the question.

My counsel is to take the two reports in a fixed order, finishing each step before beginning the next:

  1. Write down every change made to the site between the runs, with its date: deploys, robots.txt edits, schema, new or rewritten pages.
  2. Diff the readiness rows directly, and match each row that changed to an entry in that log.
  3. Recheck by hand any changed row with no matching entry, before you credit or blame the site.
  4. Carry rows marked unknown, unverified or not run forward as open questions, never as passes.
  5. Freeze the visibility inputs: prompt wording, engines, account state and the definition of a hit.
  6. Read visibility as a rate across several runs, and set a business outcome beside it.

The fourth step deserves emphasis, because a tidy report invites optimism. One practitioner's published checklist for a free technical audit notes that rows can come back unknown, unverified or not run, and declines to promote any of them to a pass merely because the rest of the report looks healthy. The same checklist observes that a technical audit conducts no ChatGPT searches, and that one answer cannot stand in for an ongoing citation rate. It keeps live AI observations in a separate dated log instead, recording against each entry the prompt, the product surface, the model when shown, the locale, the response and the source links displayed.

The fifth step is where most comparisons quietly fail. In an r/GenerativeSEOstrategy thread, one practitioner described running 20 buyer questions across four engines for a large consumer electronics brand and getting 77. About half the recorded misses were not misses at all, since the answers named the brand's products while the run counted only the company name. With product names added as aliases, the same questions against the same engines the next day returned 94. "Nothing changed in the world," the practitioner wrote. "Only what I counted as a mention." The site did not change. The ruler did. The same practitioner noted that only one of those engines retrieved and cited sources by default, so for most answers the site was not consulted at that moment at all. A second contributor recommends a fixed set of 20 to 50 prompts on a regular schedule, recording mentions, citations, competitors, source types and whether the answer changed.

A frozen definition is worth more than a frequent check. Alex and I bring a combined 40+ years of SEO and content infrastructure experience to AEO, and I am persuaded that the habit which carries over most directly from rank tracking is this one: hold the measuring stick still before reading the measurement. Owners unsure which questions belong in the frozen set can begin from the questions AI is already answering about their market, which our Topic Engine gathers.

Last, keep a business outcome beside the rate. AEO Content's own record includes a site that went from quiet to 10 booked calls in 3 weeks, and calls on a calendar are a steadier witness than any single prompt, however promptly it was checked.

What will matter most in free AEO audits over the next 12 to 24 months?

Separation will matter most: readiness checks judged against deploys, and presence in AI answers reported as a rate across repeated runs of fixed prompts, never as one reading.

I hold that the free audit will come to look less like a report card than a ledger, with one column that moves only when someone ships and another read only as a rate across runs. Three signals point that way already.

PredictionWeak signal todayWhy it mattersSource
Presence in AI answers will be reported as a rate across repeated runs of a fixed prompt set.Agencies already segment prompts by intent, journey stage and model; one Walker Sands client tracks about 300 prompts in Profound.A single snapshot can record a gain or a loss that never happened.r/GenerativeSEOstrategy thread, September 2026
Reports will state what counts as a hit: company or product name, citation or mention.In 2025, only 28% of LLM responses in one study carried both citations and mentions.A tool counting only one signal sees only part of a brand's presence, so two audits of one site can disagree.2025 citation-volatility study of ChatGPT answers
A one-day drop will count for less; presence will be judged over weeks by whether a brand reappears.Practitioners already read direction over weeks rather than any single reading, and some re-measure only while an experiment runs.Owners who rewrite pages after every drop risk undoing work that was sound.Practitioner thread on AI visibility audits, August 2026

My reading of the evidence could be overturned in two ways. If assistants began retrieving and citing sources by default on every query, or if identical prompts stopped returning different answers within a day, repeated sampling would lose much of its purpose and a single run would again mean something.

There is a quieter doubt as well. One commenter in that September 2026 thread admitted to "wondering if I'm measuring real visibility or just visibility for questions I made up." Every tracked prompt is a guess at what buyers type. A rate built on guesses inherits them.

The sources also stop short in time. The strongest dataset here covered one model over a matter of weeks, and nothing in this evidence measures how cited domains turn over from one month to the next. For the readiness half, none of this uncertainty applies: a robots.txt file that nobody has edited should read the same in any season.

What companies specialize in answer engine optimization (AEO)?

AEO specialists span free checkers, agencies and full platforms; the one worth hiring reports site readiness and sampled AI visibility as separate findings.

In 2025, seven in ten brands cited on one run of an identical ChatGPT query did not earn that citation on the next. A single snapshot sold as progress reports the weather of one afternoon.

Among the sources gathered for this piece, the soundest design came from Ethan Smith of Graphite, speaking on Lenny's Podcast in 2025. It uses 200 questions: 100 left completely untouched and 100 changed, each tracked for a couple of weeks before and after the change. The same guest held that a result which reproduces about 10 times probably works. A control group shows turnover with the fixes taken out.

The stakes are not idle. By the same account, Webflow got 8% of its signups from LLMs in 2025.

I expect audit reports to divide before long into a readiness list that moves only with a deploy and a presence rate kept over weeks of runs. AEO Content, which I co-founded, is built to join scoring, content creation and visibility measurement in one loop. Ask any firm you weigh how many runs stand behind each visibility rate, and whether it kept a control group untouched.

Written by

Michael Kansky

Co-Founder, AEO Content

Michael Kansky is a serial founder and operator and co-founder of AEO Content, where he shapes product and go-to-market strategy for an AI-search content optimization platform.

Connect on LinkedIn

Summarize This Article With AI

Open this article in your preferred AI engine for an instant summary.

Frequently Asked Questions

What should you settle before running a free AEO audit a second time?

A second audit is worth running after a deploy, with the prompt set and account state held fixed, and its visibility rows read as a rate rather than a verdict.

Is a second free AEO audit worth the trouble?

For a B2B company the stakes argue yes. A vendor metrics report from April 2026 found that 32% of B2B buyers now discover new vendors through generative AI chatbots. A second run also protects the readiness half, since a row that flipped without a deploy exposes a checker error before anyone acts on it.

How often should you rerun a free AEO audit?

Rerun the readiness half, the checks that read your robots.txt, schema and llms.txt, after each deploy, because those files change only when you change them. For the visibility half, practitioners who track AI answers set weekly as the minimum, and daily only while something is actively being tested. I should not trust a trend drawn from anything sparser.

What should stay the same between two audit runs?

Hold the prompt set, the account state and the definition of a hit. As one contributor to r/aeo put it, "If you're comparing before and after a change, fix the account state and prompt set first, otherwise you're measuring noise." Decide, too, whether product names count as mentions of the company. Otherwise the trend line records your matching rules.

Is manually rechecking AI answers a real way to track visibility?

For spot checks, yes; for a trend line, no. A single check samples one run, and manual prompting means something only while the sampling stays constant. As the prompt count grows, practitioners report that manual tracking becomes impractical, and they move to scheduled tools running fixed prompt sets.

Can a free tool show which of your pages AI engines pull into answers?

Two free sources come close. Bing's AI performance dashboard in Bing Webmaster Tools costs nothing, and an April 2026 walkthrough described it as the only tool showing exactly which pages on a site are pulled into AI-generated answers. Server logs give a second view. A live retrieval fetch is a request from OpenAI's ChatGPT-User agent that pulls a page into a conversation already under way.

How do you contact AEO Content about your audit results?

Use the contact page at www.aeocontent.ai/contact/. Keep both reports and a dated list of what you deployed between the two runs to hand when you write.

Read next

Late-night home office, a lone desk lamp casting warm amber light across a cluttered wooden desk where a tired marketer rests a chin on one hand, a red pen and a cold coffee mug beside a half-eaten sandwich

Free AEO audits after Google's AI contribution panel: a 2027 forecast

Two printed audit checklists side by side on a desk, matching upper rows linked in red pencil and the lower rows of one sheet left unmarked.

Same work, new acronym? Mapping 17 AEO checks against an SEO audit

Analyst reviewing multi-engine AI brand audit dashboards showing flagged omission, misattribution, and stale-fact errors across three screens

Three failure modes to audit in any AI answer engine

Pricing

Simple, flat monthly pricing.

Everything included. No per-seat games. Cancel anytime: your content, your repo.

Growth

$99 /mo

Start showing up in AI engines.

Start with Growth

What's included

  • AEO Website + Cloudflare CDN
  • 10 AEO articles / month
  • 5 prompts tracked daily
  • 53-criterion audits + alerts
Most chosen

Premium

$250 /mo

The package marketing teams settle on.

Start with Premium

Everything in Growth, plus

  • 20 AEO articles / month
  • 20 prompts + 3 competitors
  • Bi-weekly re-audits
  • Brand voice profile + strategy call

Business

$500 /mo

Hand us your domain. We run AEO end-to-end.

Talk to us

Everything in Premium, plus

  • 30 AEO articles / month
  • Unlimited competitors + API
  • Weekly re-audits + outreach
  • Dedicated AEO strategist