Cross-engine consistency audit cover for ChatGPT, Perplexity and Gemini on UK queries.

TL;DR

  • A cross-engine consistency audit measures how far ChatGPT, Perplexity and Gemini agree with each other, and with themselves over time, on the same UK queries.
  • Agreement between engines is low. Only about 11% of cited domains overlap between ChatGPT and Perplexity across roughly 680 million citations (Profound, 2026).
  • Self-consistency is also weak. Only around 30% of brands appear in back-to-back AI responses to the same prompt (AI visibility trackers, 2026).
  • We score three dimensions: citation-count consistency, recommended-firm consistency and recency-of-source consistency, across 50 UK queries, three engines and four weeks.
  • The output is a stability map that tells you which queries you can rely on and which are too volatile to report as wins.
A cross-engine consistency audit runs the same commercial queries across ChatGPT, Perplexity and Gemini, repeatedly over several weeks, and measures how consistent the results are. It captures two kinds of instability: engines disagreeing with each other, and each engine disagreeing with itself run to run. Both are high in 2026. The audit turns that instability into a score per query, so a firm knows which visibility claims are durable and which are noise.

Key facts

  • Only about 11% of cited domains overlap between ChatGPT and Perplexity across roughly 680 million citations (Profound, 2026).
  • Only around 30% of brands remain visible in back-to-back AI responses to the same query (AI visibility trackers, 2026).
  • A citation share above about 60% across repeat runs signals durable visibility, while under about 20% signals volatility-driven noise (AI visibility trackers, 2026).
  • Independent 2026 testing that repeated prompts across engines found agreement on some trust queries and wide divergence on commercial ones (Rank4AI, 2026).
  • Gemini is frequently flagged as the least consistent of the three, strong when it works and erratic when it pulls from weaker sources (Rank4AI, 2026).
  • Structure and attribution raise the odds of a stable citation on every engine (Aggarwal et al., arXiv, 2023).

Why consistency is a metric, not a footnote

Most AI visibility reporting takes a single snapshot. You run a query once on one engine, see your brand, and record a win. The problem is that AI answers are unstable in two directions at once. Different engines return different sources for the same query, and the same engine returns different sources on repeat runs of that query. A snapshot cannot tell the difference between a durable position and a lucky roll of the dice, which means it cannot be trusted to guide spend.

Consistency reframes the measurement. Instead of asking whether you appear, it asks how reliably you appear, across engines and across time. A query where you show up on all three engines every week is a genuine asset. A query where you flicker in and out on one engine is noise that should not be reported as a result. The audit exists to separate the two.

The two kinds of inconsistency

Cross-engine inconsistency is disagreement between engines. It is large. Profound’s analysis of about 680 million citations found only around 11% of cited domains overlap between ChatGPT and Perplexity, and independent 2026 testing that repeated prompts across ChatGPT, Gemini and Perplexity found agreement on some trust-oriented queries but wide divergence on commercial ones. Within-engine inconsistency is disagreement of an engine with itself on repeat runs. It is also large. Only around 30% of brands remain visible in back-to-back responses to the same prompt, so roughly two in three visible brands are not reliably visible.

Bar chart showing illustrative share of UK queries where three, two or one engines recommend the same firm
Illustrative cross-engine agreement on recommended firm across UK queries. Directional pattern, not exact figures.

Dimension one: citation-count consistency

The first thing to score is whether the number of sources an engine cites for a query stays stable. Perplexity typically cites many sources per answer and ChatGPT far fewer, but within each engine the count should be reasonably steady for a given query if the underlying retrieval is stable. A wildly swinging citation count week to week is an early signal that the query sits in a volatile part of the index, where your appearance, if any, is fragile. We record the count per engine per week and measure its variance.

The second dimension is the one buyers care about. For a query like “best employment solicitor in Bristol”, which firms does each engine name, and do those names persist across runs and across engines. This is where the 30% back-to-back figure bites. A firm named once may not be named on the next run, so we track how often each recommended firm reappears. A firm that shows up in most runs on most engines holds a real position. A firm that appears once is inside the noise band and should not be counted.

Line chart of illustrative week-to-week retention of cited firms by engine over four weeks
Illustrative share of week-one cited firms still cited in later weeks, by engine. Directional pattern, not exact figures.

Dimension three: recency-of-source consistency

The third dimension checks whether the age profile of the cited sources holds steady. If an engine cites month-old sources one week and two-year-old sources the next for the same query, the topic is churning and any position on it is temporary. Perplexity’s strong recency weighting makes this dimension especially informative there, because a sudden shift toward older sources can signal that fresh coverage has dried up and the query is about to destabilise. Stable recency profiles, by contrast, tend to accompany stable citations.

Turning three dimensions into a stability score

Each query gets a score on all three dimensions per engine, then a blended cross-engine consistency score. The banding is simple and borrows from how practitioners already read citation share. A query where your presence holds above about 60% of runs is durable and can be reported as a win. Between about 20% and 60% is contested and worth targeted effort. Below about 20% is noise, where a single appearance means little. The score converts a wall of run-by-run data into three actions: defend, invest, or ignore.

Bar chart comparing cross-engine domain overlap and within-engine back-to-back brand visibility
Cross-engine domain overlap and within-engine back-to-back visibility, two measured sources of instability (Profound; AI visibility trackers, 2026).

What the audit changes about client reporting

The consistency lens changes what an honest agency puts in a report. A snapshot invites cherry-picking: run a query enough times on enough engines and something flattering eventually appears, which can be screenshotted and presented as a result. A consistency audit removes that temptation by design, because a position only counts if it survives repetition. That protects the client from paying for noise and protects the agency from the awkward month when a screenshot-based win quietly vanishes.

It also reframes the conversation around targets. Rather than promising a brand will appear in ChatGPT, which is unfalsifiable on a single run, an agency can commit to moving a query from the contested band into the durable band, a claim that is measurable, time-bound and defensible. That is a healthier basis for a retainer than a gallery of one-off screenshots, and it aligns the incentive with real, repeatable visibility.

How to run a lightweight version yourself

You do not need 50 queries to start. Take 15 commercial queries that matter, run each on ChatGPT, Perplexity and Gemini once a week for four weeks, and record three things each time: how many sources were cited, which firms were named, and roughly how old the top sources were. At the end you can see, per query, whether your presence is durable, contested or noise. The discipline that keeps this honest is repetition. As the original GEO research showed, structure and attribution improve your odds on every engine, but only repeat measurement tells you whether an improvement actually held.

Frequently asked questions

What is a cross-engine consistency audit?

It is a method that runs the same commercial queries across ChatGPT, Perplexity and Gemini, repeatedly over several weeks, and measures how consistent the results are. It captures cross-engine disagreement, where engines cite different sources for the same query, and within-engine disagreement, where one engine cites different sources on repeat runs. The output is a stability score per query that shows which visibility positions are durable and which are noise.

How much do AI engines disagree with each other?

A lot. Profound’s 2026 analysis of about 680 million citations found only around 11% of cited domains overlap between ChatGPT and Perplexity. Independent 2026 testing that repeated prompts across engines found agreement on some trust queries but wide divergence on commercial ones. The practical consequence is that a strong position on one engine tells you very little about your position on another, so each engine has to be measured separately.

Why does the same engine give different answers to the same query?

Retrieval and ranking in AI answers carry real randomness, and the underlying index changes constantly. The result is that only around 30% of brands remain visible in back-to-back responses to the same prompt in 2026. So a single appearance is weak evidence. Repeating the query several times and measuring how often you reappear is the only way to tell a durable position from a one-off that happened to surface on that run.

What score means a position is durable?

As a working band, presence above about 60% of repeat runs signals durable visibility, between about 20% and 60% is contested and worth targeted effort, and below about 20% is noise where a single appearance means little. These thresholds mirror how practitioners already read citation share across repeat runs. The point is to stop reporting a lucky single run as a win and to report only positions that hold across repeats.

Is Gemini less consistent than the others?

It is frequently flagged that way in 2026 testing, described as strong when it works but erratic when it pulls from weaker sources. That does not make it unimportant, since it has wide reach through Google’s ecosystem, but it does mean Gemini positions need more repeat measurement before you trust them. Treat a single strong Gemini result with more caution than a single strong Perplexity result, and weight the consistency dimension more heavily there.

How many queries and weeks do I need?

A useful audit can start with 15 commercial queries run weekly for four weeks across the three engines, which is enough to separate durable from noisy positions. A fuller version uses around 50 queries for more stable per-segment reads. What matters more than volume is repetition over time, because the instability the audit measures only appears when you run the same query more than once and compare.

Sources and references

  1. Most-cited domains across AI answer engines (approx. 680 million citations). Profound, 2026
  2. How do AI search results compare across ChatGPT, Gemini and Perplexity. Rank4AI, 2026
  3. Visibility volatility in AI search. Omnia, 2026
  4. Ranking volatility and AI answer instability. Digital Applied, 2026
  5. GEO: Generative Engine Optimization. arXiv (Aggarwal et al.), 2023
  6. Which sources AI Overviews and chat engines cite. Search Engine Land, 2026

See where your firm is cited across engines today, the baseline for tracking consistency.

Get your AI visibility report

Change log

  • 2026-07-13: Initial publication.