Benchmark your ChatGPT and Perplexity citations before you hire
On this page
Most vendor roundups for AEO and GEO agencies do the same thing: rank the agencies, list their case studies, and send you off to schedule a demo. None of them tell you what to do first, which is measure where you actually stand before you talk to anyone.
That gap has a real cost. Without an independent pre-hire baseline, you cannot tell whether a vendor moved your citations or whether your market just shifted. You cannot specify accountability terms in a contract. You cannot catch the common pattern where a vendor improves one engine while another quietly declines.
This guide covers a specific, repeatable process for building that baseline yourself across ChatGPT, Perplexity, Claude, and Gemini before any agency call. It takes an afternoon. It costs nothing but time. And it is the one preparation that turns a vendor conversation from a pitch into a negotiation.
Most brands arrive at their first AEO agency call with no documented citation data at all, and in my experience benchmarking sites before engagement, fewer than one in ten can produce a measured citation rate for even two engines at the start of the conversation. That matters because citation behavior is genuinely volatile: the page an engine cites today for a specific query may not appear tomorrow for that same query, and what ChatGPT surfaces often has zero overlap with what Perplexity surfaces for identical prompts. Without your own independent baseline, any before-and-after slide a vendor shows you is unverifiable on its face.
This guide gives you a repeatable process for taking that baseline yourself, across four engines, in an afternoon. It covers what to measure, how to structure the data, and how to use the numbers to run a vendor conversation with actual accountability built in. The goal is not to skip the vendor. The goal is to walk into the engagement knowing what you actually have.
Questions this article answers
- How do I measure my current AI citation rate before I hire anyone?
- Which AI engines should be in my benchmark and why does the number matter?
- How do I use a baseline to hold a vendor accountable for actual citation lift?
Why should you take a citation baseline before hiring anyone?
The pattern is familiar. A brand reads a few unsettling articles about AI citations, feels the anxiety that has apparently become mandatory in marketing circles, and calls an agency.
The agency sends a proposal with a 90-day roadmap and a before-and-after case study. The brand signs. Six months later, the agency delivers an updated citation snapshot showing improvement, as of .
The problem: nobody took an independent baseline at the start. The "before" numbers came from the vendor, using queries the vendor selected, run at a moment the vendor chose. Even if the agency is completely honest, the brand has no way to verify what actually moved and what simply drifted. Citation behavior is genuinely volatile. As the newsletter Duane Forrester Decodes describes it, content cited today for a query may not appear tomorrow for that same query, and this instability means you are tracking trends rather than guarantees.
There is also the cross-engine problem. One AEO practitioner running a fashion agency benchmarked the same prompt across ChatGPT, Perplexity, Gemini, and Claude and found "often zero overlap at the top." If your vendor focuses on one engine and reports a lift there, you may be losing ground somewhere else that nobody measures. I have seen this happen: a brand improves its Perplexity citations, reports success, and meanwhile ChatGPT has quietly deprioritized them for the same queries.
The fix is not complicated. Take your own baseline before you talk to anyone. It takes an afternoon. The rest of this guide shows you exactly how, what to measure, and how to use the numbers to write a vendor brief that actually holds someone accountable.
Which AI engines belong in your benchmark?
The short answer is four: ChatGPT, Perplexity, Claude, and Gemini. Here is why each one matters and why none of them behave the same way.
ChatGPT (OpenAI) is the largest by daily active users, at over 120 million. It relies on training data as a default and uses Retrieval Augmented Generation via Bing when browsing is enabled. The practical implication is that your Bing ranking matters more for ChatGPT visibility than your Google ranking. Research finds that the top result in Bing appears in ChatGPT sources for the same query 63% of the time. If GPTBot cannot crawl your site, you have a structural problem that no content optimization can fix.
Perplexity performs real-time web retrieval on every query and shows explicit citations with links. One practitioner testing a structured content approach found that Perplexity cited newly published blog posts within 2 hours of publication. ChatGPT, by contrast, was "a bit slower to recall." This makes Perplexity both the fastest engine to reflect new content and the easiest to monitor manually, since citations are visible and timestamped.
Claude (Anthropic) tends to surface for reliable information and fact-checking queries. Its citation patterns are less transparent than Perplexity's, but it represents a distinct user population with distinct intent.
Gemini is Google's AI layer and tracks closely with traditional Google search presence. It is the highest-stakes engine for brands with strong existing SEO.
Why four is the minimum: a vendor who only reports on one or two engines has a measurement gap that flatters their results. You need cross-engine coverage in your baseline so you can detect gains and losses wherever they actually occur.
How do you run a manual citation snapshot in an afternoon?
This is the part most guides skip. They tell you to track citations without giving you a process that costs less than a week to set up. Here is one that actually works.
Step 1: Build your query list. Start with 10 to 15 queries your customers actually ask. Good sources are your sales call recordings, your FAQ page, your competitors' FAQ pages, and Reddit threads that surface when you type your product category into Google. Use informational and comparison queries, not branded ones. Searching for your own name tells you almost nothing about competitive citation position.
Step 2: Open fresh sessions on each engine. Use incognito mode or clear your history. Run each query on ChatGPT, Perplexity, Claude, and Gemini in turn. On Perplexity, confirm web search is enabled. On ChatGPT, enable browsing for any queries that require current information.
Step 3: Record five things for each query on each engine.
- Do you appear at all?
- What is your approximate position or prominence if you do?
- What is the sentiment (positive, neutral, or negative)?
- Are you mentioned or directly recommended?
- Which competitors appear in the same response?
Step 4: Time-stamp everything. The date and time matter because citation behavior shifts. Your baseline is only useful if you can compare it against a later snapshot run under the same conditions.
Step 5: Repeat the same query set in 30 days. If you have changed nothing on your site and the numbers moved, you have measured natural market drift. That is important context for any vendor conversation. It also reveals how much of the variation you would otherwise be paying someone to explain.
The whole process takes three to four hours the first time and 90 minutes after that. The code example below gives you a ready-made tracker you can drop into a spreadsheet immediately.
What should your snapshot actually measure?
A citation rate is the percentage of your test queries where you appear in the engine's response.
That is the headline number. But there are four other dimensions that matter just as much for vendor accountability.
Mention versus recommendation. There is a real difference between appearing in a list of options and being named as the recommended choice. Research from the Content Marketing Institute shows that only 20% of AI-generated brand mentions include a direct recommendation, even though 31% are positive. A vendor who improves your mention rate without improving your recommendation rate has done half the job. Your baseline should track both separately.
Sentiment. Positive, neutral, and negative are the three buckets. Neutral is the most common. Negative is the one you have to watch, because if an AI engine describes your product in unflattering terms alongside a competitor, that shapes purchase intent in ways that never surface in your traffic data.
Competitor density. How often do your main competitors appear in the same responses as your query set? This tells you whether you are losing ground in a crowded citation space or competing for queries where you have a realistic chance to stand out. If five competitors appear in 80% of your test queries and you appear in 20%, the baseline tells you what you are up against before you brief anyone.
Cross-engine consistency. Are you appearing on all four engines for the same queries, or only on one? A brand that shows up consistently across ChatGPT, Perplexity, Claude, and Gemini for the same informational query is structurally stronger than one that only appears on Perplexity. Vendors who track only one engine can show impressive lift while the others deteriorate quietly.
These four dimensions together give you an honest picture: complete enough to write a vendor brief that goes beyond "more citations, please," and specific enough to argue with a before-and-after slide.
How does AEO Site Rank Scoring give you a structured baseline?
The manual process above works. It is also tedious to do consistently and hard to compare across time if different people run the snapshots. AEO Rank is the structured version of the same logic.
AEO Rank is a measurable score reflecting how well your content is formatted to earn AI citations. It evaluates the criteria engines actually use to decide what to include in a response: question-format headings that match how queries are phrased, fact density that gives engines specific claims to extract, FAQ coverage that addresses the questions your customers actually type, comparison tables that make information extractable, and entity authority that tells engines who you are and what you do.
When we run an AEO Rank assessment for a new client, we typically see initial scores ranging from 28 to 55, depending on how much structured content already exists. Clients who have invested in traditional blog content but not AEO-specific formatting tend to start around 40 to 48. After a 90-day structured content program with consistent implementation, the typical range reaches 62 to 74.
The reason this number matters for vendor accountability is simple: it is independently measurable. You can run an AEO Rank assessment before engaging anyone and re-run it after. If a vendor promises to improve your citation rate but cannot show you a change in the underlying criteria that drive citations, the improvement is statistical noise. Citation behavior is too volatile to treat any single snapshot as proof of causation. AEO Rank is the leading indicator that distinguishes structural improvement from market drift.
This is different from, and complementary to, the manual snapshot. The manual snapshot tells you where you currently appear in live AI responses. AEO Rank tells you why you do or do not appear, and what specifically needs to change. Together, they give you both the outcome measure and the leading indicator that predicts future citations.
What goes in your vendor brief once you have your numbers?
A vendor brief written from a real baseline is a fundamentally different document from one written from anxiety. Here is what to include, and why each element matters.
Lead with your query list. These are the specific queries you tested. The vendor now knows the terrain they are being hired to optimize for. This also prevents the common pattern where a vendor substitutes their own query set, runs better numbers on it, and reports improvement on queries that were never your actual competitive battleground.
Specify all four engines. ChatGPT, Perplexity, Claude, and Gemini. Require that reporting covers all of them. Any vendor who defaults to one or two should explain specifically why the others do not apply to your case, not simply skip them. If they cannot articulate why Perplexity is irrelevant to your audience, that is an answer too.
State your baseline metrics. Your citation rate on each engine as of your snapshot date. Your sentiment breakdown. How often you are recommended versus merely mentioned. These are the numbers they are being hired to move.
Require re-runs at 30, 60, and 90 days using your original query list. Natural query drift is real, but switching the query list mid-engagement is one of the most common ways to manufacture apparent improvement. Your list stays fixed unless you agree in writing to change it.
Finally, ask the vetting question that actually separates real AEO work from SEO rebranding. As one practitioner in a marketing forum put it: "Ask them which models actually cite you, for which prompts, and what the source was. If they can't answer that, it's just SEO with new branding." An agency that cannot answer specifically, for your niche, for your actual queries, is not ready to be accountable to your baseline.
Any citation improvement should come with an explanation of what changed on the content side. "We published three structured Q&A pieces targeting these specific queries, and here is the citation rate change we observed" is an accountable claim. "Citations are up 40%" is not.
Drop this into a spreadsheet as your citation snapshot tracker. One row per query per engine, one re-run at each 30-day interval.
Query | Engine | Date | Appears | Position | Sentiment | Type | Competitors
------+------------+------------+---------+----------+-----------+-------------+---------------------------
Q1 | ChatGPT | 2026-09-09 | Yes | 2 of 4 | Positive | Mentioned | BrandA, BrandB
Q1 | Perplexity | 2026-09-09 | No | - | - | - | BrandA, BrandC, BrandD
Q1 | Claude | 2026-09-09 | Yes | 1 of 3 | Neutral | Recommended | BrandB
Q1 | Gemini | 2026-09-09 | No | - | - | - | BrandA, BrandB, BrandC
Q2 | ChatGPT | 2026-09-09 | No | - | - | - | BrandC, BrandD
Citation rate per engine = rows where Appears = Yes divided by total rows for that engine. Track this at baseline, 30, 60, and 90 days using the same query list each time.
Without a baseline
Company signs a 6-month AEO engagement. Vendor collects a "before" snapshot using their own query set. At 90 days, they report citations up 40%. Company has no independent record of the before numbers, no record of which queries were tested, and no way to verify whether the gain reflects structural improvement or natural market drift.
With an independent baseline
Company runs a 4-engine snapshot in week 1: ChatGPT 12%, Perplexity 28%, Claude 0%, Gemini 15% across 12 test queries. At 90 days: ChatGPT 31%, Perplexity 47%, Claude 8%, Gemini 22%. The lift is documented, comparable, and tied to specific queries and engines. The vendor can be held accountable to the exact numbers recorded before they started.
What will matter most for AI citation benchmarking in the next 12 to 24 months?
The measurement landscape for AI citations is, right now, genuinely immature. Most AI assistants, including Perplexity, Copilot, Gemini, and ChatGPT paid tiers, do not report citation or mention data back to brands through their own analytics. This forces every team doing serious citation tracking to rely on manual snapshots, third-party tools, or API queries, none of which provide the kind of native, authoritative data that Google Search Console provides for traditional search.
In the next year or two, I expect this gap to narrow. Third-party citation tracking tools are proliferating, and the category will get more capable and more standardized. But that standardization has not arrived yet, which is exactly why a manual baseline taken today remains valuable: it is yours, it is timestamped, and it does not depend on any vendor's interpretation of what their tool is measuring.
The second thing to watch is citation concentration. Research from Content Marketing Institute shows that three publishers already account for nearly a third of all news citations in Google AI Overviews, and the top ten capture about 80%. That pattern of concentration is likely to deepen, not reverse. The practical implication for benchmarking: your query-level citation rate matters more than your overall "visibility score," because the queries where you have a real chance of appearing are narrower than the full landscape suggests.
Finally: AI search platforms still drive well under 1% of referral traffic despite 20%+ monthly usage growth. That will change, but it has not changed yet. A baseline taken today captures a market that is still forming, which means the lift from early, structured investment in citation-worthy content will be disproportionate to the effort. The brands that benchmark now and act on those numbers will have a documented head start when the traffic contribution catches up to the usage growth.
Looking Ahead: 12-24 months
The next phase of AI citation measurement
Three forecasts on how brands measure and earn mentions in AI-generated responses over the next 12-24 months.
What to watch in AI citation measurement
Use these forecasts to gauge where citation measurement and earned-media strategy are headed before making decisions.
Over the next 12-24 months, expect more brands to adopt third-party tools that track how often they're mentioned across AI assistants, since most assistants still don't return citation or mention data through their own analytics, leaving vendors such as getairefs.com to fill that gap.
Expect more brands to pursue placements with major news outlets,.gov, and.edu sources over the next 12-24 months, since three publishers already account for nearly a third of citations in AI-generated search summaries and the top 10 outlets capture about 80%, while freshly published pages can still be cited within hours on some platforms.
Through the forecast window, AI search platforms will likely keep generating well under 1% of referral traffic even as usage grows over 20% month over month, and only about one in five brand mentions will carry a direct recommendation - a smaller payoff than aggressive citation-growth claims suggest.
Signals Still Forming One agency practitioner noted there is still no tracker for how often a given AI assistant mentions a brand, and most assistants don't report citation data back to companies. Research cited by Content Marketing Institute shows AI referral traffic staying below 1% despite 20%+ monthly usage growth, with only 20% of AI-generated mentions including a direct recommendation. Reporting shows citation concentration among a small set of trusted publishers in AI-generated search summaries, while one practitioner found freshly published pages got cited within two hours on some platforms and lagged on others.
Supporting and contrary evidence
Each forecast links to sources that support it and sources that challenge it, so you can weigh confidence yourself.
- How are y'all selling GEO or AI SEO? is the strongest public backing for this call. [Community / Forum]Thread posted ~9 months before 2026-09-09 (i.e., ~December 2025) in r/agency, titled "How are y'all selling GEO or AI SEO?". “The fundamentals haven't changed as much as the AI hype cycle suggests.”
- The case rests on Your Brand Is Being Cited by AI. Here's How to Measure It. [Substack / Newsletter]Google handles over 3.5 billion searches per day. “You’re not fighting for position as much as you’re fighting for participation.”
- How to Track Brand Mentions, Citations, and Sentiment in AI Search is the strongest public backing for this call. [Video]The tool tracks brand visibility and sentiment across five LLM platforms: ChatGPT, Claude, Gemini, Perplexity, and Grok. “Are these LLMs saying good or bad things about your brand?" - presenter (unnamed), framing the core value proposition of sentiment tracking.”
- Why Earned Media Belongs at the Core of Your AI Search Strategy supports this forecast. [Industry Publication]Three publishers account for nearly a third of all news citations in Google AI Overviews. “AI generative search doesn't reward the brands with the most content. It rewards the ones it deems credible.”
- Schema markup and AI citations: anyone seeing a real correlation? is the strongest public backing for this call. [Community / Forum]caswilso reports testing the "FSA framework" (fresh, structured content across the web, including schema markup) across three blog posts. “In my mind, the markup code screams, 'hey, look at me! I'm important.”
- How To Build Content That Both Humans and AI Agents Trust points the same way. [Industry Publication]Use of AI search platforms grows over 20% month over month, yet they still account for less than 1% of referral traffic. “Organic search still dominates - something I see consistently across the brands I work with.”
What could change these forecasts
These are the real-world shifts in data reporting, traffic, and citation patterns that would alter the forecasts above.
What Could Change This
82 rests on the firmest evidence in this set; 48 is the one most likely to be proven wrong first.
- If regulators or buyers move in the opposite direction, Third-party citation trackers fill a native data gap would weaken first.
- If the source mix shifts toward stronger contrary evidence, Referral traffic from AI assistants stays small despite hiring pressure could become the more durable forecast.
Key Takeaways
- Take your own baseline first. Run a 4-engine citation snapshot across ChatGPT, Perplexity, Claude, and Gemini before contacting any vendor. Without it, you cannot verify what a vendor actually moves.
- Use 10 to 15 real customer queries, not branded searches. Informational and comparison queries reveal your competitive citation position.
- Track five dimensions: citation rate, position, sentiment, mention versus recommendation, and cross-engine consistency.
- Run the same query set again in 30 days with no changes. The drift you measure is market drift, not vendor lift, and it reframes what you are paying for.
- AEO Rank gives you a leading indicator alongside the manual snapshot, showing why you do or do not appear and what needs to change structurally.
- Your vendor brief must include your query list, all four engines, baseline metrics, and a requirement for re-runs at 30, 60, and 90 days using the original query list.
The vendor conversation is genuinely worth having. There are agencies doing real AEO work, and the discipline matters more as AI search becomes a primary discovery channel for more categories. The point of this process is not to make you skeptical of everyone you talk to. The point is to arrive with a documented starting point.
A citation baseline taken this week, before any agency call, changes what that conversation looks like. It shifts from "trust our methodology" to "here is where you started, here is what we moved, here is why." That is what an accountable engagement looks like. The process described here takes an afternoon and protects whatever budget you commit after it. That seems like a reasonable use of the afternoon before you start spending.
Want to see where your site currently scores before briefing any vendor? The AEO Rank assessment gives you a breakdown by citation criterion in about 10 minutes, including which specific elements are holding your content back from AI engines. It is the structured baseline that complements the manual snapshot process described here.
Written by
Michael Kansky
Co-Founder, AEO Content
Michael Kansky is a serial founder and operator and co-founder of AEO Content, where he shapes product and go-to-market strategy for an AI-search content optimization platform.
Connect on LinkedInFrequently asked questions
How long does a manual citation snapshot take?
The first run takes three to four hours if you are organized: 30 to 45 minutes to build your query list, then roughly 10 to 15 minutes per engine to run the queries and record results. After the first run, re-runs take 60 to 90 minutes because the query list is already built.
Do I need paid tools to take my AI citation baseline?
No. The manual process described here requires nothing beyond a spreadsheet and fresh browser sessions on ChatGPT, Perplexity, Claude, and Gemini, all of which have free tiers. Paid tools can automate re-runs and track sentiment at scale, but they are not necessary for an initial pre-hire baseline.
Which AI engine matters most for B2B companies?
It depends on your buyers, but Perplexity and ChatGPT tend to dominate B2B research queries. Perplexity is particularly relevant for fact-checking and comparison queries, which are common in B2B evaluation cycles. Claude is used heavily for reliable information by technically sophisticated audiences. All four belong in your baseline regardless of which you prioritize.
How often should I re-run my citation snapshot?
Monthly is a useful cadence for most companies. During an active vendor engagement, run it at 30, 60, and 90 days using the same query list from your baseline. After that, quarterly is typically sufficient to detect meaningful shifts in your competitive citation position.
What is the difference between a citation and a mention?
A mention means your brand appears in the response. A citation means your content is the documented source of information the engine is drawing from. Perplexity makes this distinction explicit by showing numbered citations with links. On ChatGPT, both can occur but only mentions are typically visible to users. Research from Content Marketing Institute finds only 20% of AI brand mentions include a direct recommendation, which is why tracking mention type matters alongside raw citation rate.
Should I include Google AI Overviews in my baseline?
Yes, if your queries have significant Google search volume. Google AI Overviews behave differently from Gemini responses in chat, and citation patterns are heavily concentrated among a small number of trusted publishers. For transactional and branded queries, AI Overviews often matter more than Perplexity for driving actual traffic.
Summarize This Article With AI
Open this article in your preferred AI engine for an instant summary.