AEO Apps

A/B Testing Content Formats for AI Citation Rate

Different AI engines cite the same content differently, so test each platform separately.

Staff Writer · · 11 min read
Cover illustration for “A/B Testing Content Formats for AI Citation Rate”
AEO Measurement · October 6, 2026 · 11 min read · 2,416 words

AI answer engines have replaced the ranked list of links with a single synthesized answer, and that answer names specific sources by a logic that has little to do with domain authority alone. A majority of B2B software buyers now begin vendor research in an AI chatbot, a share that jumped sharply from the prior year according to G2's "The Answer Economy: How AI Search Is Rewiring B2B Software Buying" report, so a brand that goes unnamed in that first answer is absent before it ever reaches a website, an ad, or a sales call.

The wrong metric to accept passively

The familiar move, building a page for Google's organic rankings and assuming the AI engines will follow along, rests on an assumption that no longer holds. Only a minority of pages cited in AI answers also rank in the organic top ten. The two contests for visibility have come apart and now run on separate tracks. A page can sit on page one of Google and never once get named inside an AI Overview, a ChatGPT answer, or a Perplexity summary, because the engine generating that answer is running its own retrieval and ranking logic, not consulting the search results page a marketer spent a year climbing. That gap matters because it turns citation rate into something a team can act on. The rate at which a page or a format earns a citation is not handed down by domain reputation the way organic rank often is. It moves in response to specific, nameable choices about how content is built, structured, and evidenced, and that responsiveness to deliberate change is what makes it something a team can test, measure, and improve on purpose.

Citation outcomes across platforms for the same content change

Treating "AI search" as one target is the first design error a test can make, because each major engine retrieves and ranks sources through a genuinely different mechanism. Gemini can ground its answers in Google's live organic index, but that grounding is an opt-in behavior a developer or product surface has to trigger rather than something that happens by default, so an organic ranking gain carries over to Gemini only when grounding fires. Perplexity, by contrast, crawls in close to real time and consistently returns clickable inline citations, with a limited exception in its Writing focus mode, which makes it one of the more reliably trackable platforms for referral traffic inside GA4. ChatGPT leans on web search results only when it chooses to use them, and most of its responses rely on training data instead. Copilot generates GA4-trackable referral traffic as well, alongside ChatGPT, but each of these platforms behaves differently in how it retrieves and ranks sources. MaxAEO's analysis of 61,400 citations spanning ChatGPT, Gemini, Perplexity, Claude, Copilot, Google AI Mode, and AI Overviews found that citation rates for the same content format vary meaningfully from one engine to the next, which rules out the idea of a single universal format winner. A test built without naming the platform in advance will produce a number nobody can interpret, because the same content change can lift citations on one engine while doing nothing, or actively hurting the page, on another. A result that holds across two or more platforms is a signal to act on. A result that shows up on one platform and nowhere else is most likely an artifact of that platform's retrieval mechanics, rather than evidence that the content change itself worked.

Citation rate versus citation share in test design

Diagram: Citation Rate vs. Citation Share: Two Different Wins. Visualizes: Visualize the contrast between two metrics that measure different things: citation rate (efficiency per page) and citation share (total volume across the web).

Two numbers get used interchangeably in most conversations about AI visibility, and they measure different things. Citation rate is the share of in-play pages of a given format that earn at least one citation, a measure of efficiency per page. Citation share is the fraction of all AI citations across the entire web that a format captures in total, a measure of volume. A listicle that gets cited in one of every thirteen relevant prompts, but exists in vast numbers across the open web, can out-produce in raw citation share an original-data study that gets cited in seven of every ten relevant prompts but exists as a single page. Both formats are winning, just at different things, and MaxAEO's dataset of 61,400 citations drawn from thousands of B2B- and SaaS-intent prompts shows the pattern clearly: original-data studies lead on citation rate because they are the most efficient per page, while ranked listicles dominate citation share because the open web is saturated with them and shortlist-style queries pull them in by the dozen. Picking the wrong metric to optimize for breaks a test before it starts. A team building one flagship asset, a single study or guide meant to become the definitive source on a topic, should be testing for citation rate. A team trying to build broad share of voice across hundreds of long-tail queries should be testing for citation share instead. Running a rate-based test design against a share-optimization goal, or the reverse, produces a result that cannot be read against the original question. The most common objection to running a rate-based test is a practical one: most teams don't have enough pages in any single format to make the comparison meaningful. That objection has an answer. The control population for a rate test does not need to be freshly built. It can be a matched set of existing pages at the same domain-authority tier, because what the test measures is divergence from that baseline, not an absolute volume of pages.

The format variables that citation data says are worth isolating in a test

Format is a strategic choice that shapes citation probability directly, not a cosmetic wrapper around content that would get cited anyway. In MaxAEO's citation-rate analysis, original-data and statistics-driven studies rank first, for a structural reason: a specific number is the one thing an AI engine cannot synthesize on its own from other sources it has already seen. When a model needs a figure, it has to attribute that figure to wherever it came from, and often only one page on the web carries it. These pages also compound their advantage over time, as other articles quote the original figure and build a second-order citation trail back to it. This format carries a real cost: it demands an actual dataset and a method that holds up, since a "study" built with no underlying data earns no trust from either the engine or the reader. Comparison and versus pages rank second on citation rate in the same dataset, because "X vs Y" queries are shaped exactly like a shortlist request, and the engine needs some page to credit when it lays out the trade-offs between two options. Gemini and Perplexity in particular pull feature matrices out of comparison pages almost word for word when the page includes an actual table, though the page has to include real criteria, honest trade-offs, and named alternatives to earn that treatment; a thinly disguised sales page does not qualify. Ranked listicles and best-of roundups rank third on citation rate but lead on citation share by a wide margin, a pattern Outwrite's AEO content strategy analysis attributes to the sheer size of the shortlist query pool and to a format that mirrors the shape of the engine's own output. FAQ and Q&A content earns its place for a related reason: question-shaped headings hand a retrieval system an answer unit already matched to a specific prompt, which makes it easy to extract and cite. For a team choosing where to spend a test cycle, three conversions carry the clearest testable lift: turning a standard blog post into a comparison page on the same topic, adding original first-party data to an existing high-authority guide, and converting a prose answer into a structured FAQ block. Each of these isolates one format variable against a natural, already-existing control.

The structural elements inside any format that act as independent test variables

Below the level of format sits a set of choices a team can change on a page without rebuilding it, and those choices move citation outcomes on their own, independent of which format the page uses. The GEO study run by researchers from Princeton, the Allen Institute for AI, Georgia Tech, and IIT Delhi (Aggarwal et al., published at ACM KDD 2024) tested nine distinct optimization methods across a large corpus of queries, and two stood out. Adding statistics to a page improved AI visibility by a substantial margin, and citing external sources produced the single largest lift of any method tested for lower-ranked content, with gains up to 115% for pages ranked around position five, even as the same change produced a slight decrease for pages that were already ranked at the top. That asymmetry matters for test selection: a page already performing well has less to gain, and possibly something to lose, from adding more external citations, while a page struggling for visibility has the most to gain from exactly that change. Position on the page is a separate, independently testable lever. A large share of ChatGPT citations draw from content sitting in the first third of a page, which suggests that moving the most answerable, most extractable block of content higher up is a cheap test to run before committing to a full rewrite of the page. Schema markup, the Article, FAQPage, and HowTo structured-data types, behaves differently from either of these. Outwrite's analysis treats schema as structural hygiene rather than a direct driver of citations, and that reading lines up with testing from Ahrefs, which found no statistically significant change in AI citations after adding JSON-LD schema to a page. Schema is still worth testing on its own, by adding it to one page in a matched pair while leaving an equivalent page unmarked, precisely because the expectation of no effect is itself a testable claim. Each of these three variables, statistics and external citations, position on the page, and schema, can move independently of which format the page uses. That independence is what makes it possible to run a test that isolates one of them cleanly enough to produce a result worth trusting.

How to structure a GEO A/B test

A GEO test does not work the way a web conversion test does. There is no real-time split of traffic between two live versions of a page, because citation behavior is not a single visit to be routed one way or another. The workable structure instead splits the topics a team already tracks into a test group, where the change under study gets applied, and a control group that is left untouched, and then compares how citation rates for the two groups diverge over the following weeks. The result worth reading is that divergence itself, not the raw citation count on either side. The control group has to be matched to the test group on format, domain-authority tier, and query type before the test ever starts. That matching is what lets any gap between the two groups be attributed to the change itself rather than to pre-existing differences between the groups. Once the observation period closes, the divergence gets checked with standard proportions testing, either a chi-square test or a two-proportion z-test run on the citation counts, with a p-value below 0.05 treated as the threshold for saying the gap is unlikely to be chance. These are the same statistical mechanics used for years in ordinary conversion-rate testing, applied here to citation counts. Teams sitting on enough historical data can reach for a more rigorous alternative, the synthetic control method, which builds a statistical model of what a topic's citation metrics would have looked like without any change at all, based on the historical behavior of topics that were never touched. The gap between that modeled baseline and what actually happened after the change is the causal lift estimate, and it is a more defensible number than a simple split can produce, but it demands both a deep history of prior data and real data science capacity to build. For a content team running two or three tests in a quarter, that overhead outweighs the benefit, and the matched-topic split is the better fit for the resources most teams actually have. Synthetic control earns its keep at a different scale, among large organizations running dozens of experiments in parallel and needing econometric rigor to defend the results internally. The one rule that holds regardless of which method a team picks is to change a single variable per test. Converting a page to a comparison format at the same time as adding schema to those same pages leaves no way to say which change, if either, caused whatever gap appears in the data afterward. Isolating one variable at a time is what separates a test that produces a usable answer from an exercise that only produces a story someone tells about it later.

Metrics to measure and tools that expose the signal

Every test design above runs into the same underlying obstacle: AI citation behavior is non-deterministic, and the same prompt put to the same engine on two different occasions will rarely return the same set of brand mentions. That instability rules out rank as a usable metric, since there is no stable rank to track from one query to the next. Frequency, measured across a large enough volume of prompts, is the measure that actually holds up under that instability. The operative number for a GEO test is visibility rate, the percentage of relevant prompts across a statistically meaningful volume that mention a given brand, and a single prompt checked once tells a team nothing beyond what that engine happened to generate at that moment. Translating visibility rate into trackable traffic runs into the platform divergence described earlier in a very concrete way. GA4 can track ChatGPT referral traffic, but only when a user actually clicks a citation link, and the majority of ChatGPT brand mentions never surface as a clickable link at all, which leaves GA4 picking up an incomplete slice of ChatGPT activity. Perplexity is the platform where that measurement gap closes, since every citation it generates is clickable, which makes GA4 a complete signal there in a way it simply is not for ChatGPT. Reading a GEO test well means pairing that referral data with direct, volume-based prompt tracking across engines, so that the visibility rate a team calculates reflects actual citation frequency rather than only the subset of citations that happened to come with a working link.

Sources

  1. Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)
  2. From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms
Filed underAEO Measurement

More in AEO Measurement