Your AI Visibility Number Is Not Comparable to Anyone Else's
AI visibility tools report figures more than four times apart on the same metric, so treat your own citation data as a trend in one tool, not a benchmark.
If your content plan still has listicles and comparison pages in it, you have probably been told in the last two weeks to cut them. The instruction traces back to a single number: the share of the sources ChatGPT credits in its answers that are listicles fell by half after its 5.6 release. That number is now written into a public library of instruction files that marketers load into their AI assistants, which is how a measurement turns into an industry position without anyone voting on it.
It might be right. It was also tested once, in public, by someone working from their own dataset, and it came out the other way. The obvious response is the one every commentator reaches for: stop arguing about other people's numbers and go measure your own. That response is correct and it is not sufficient, because the tools you would measure with report figures more than four times apart on the same metric over the same period.
The caveats did not travel with the number
Where the number came from turns out to matter more than what it says. The finding is Peec AI's, reported through Tomek Rudzki and Lily Ray in August 2026: listicles fell from 15.77 percent of ChatGPT citations to 7.80 percent, a drop of 50.5 percent, and comparison pages went from 9.08 percent to 6.17 percent. Peec AI sells AI-visibility tracking, which makes it the source of a headline figure about the thing it sells. That does not make the finding wrong, but it is worth keeping in view. On September 5 the largest of those skill libraries encoded the finding in a new format volatility reference (release notes).
To be fair to the library, it did not launder the number. Its own file records that no sample size was stated for the underlying data, that the measurement covers ChatGPT and no other platform, that every figure should be read as a dated snapshot, and that you should "verify against your own citation monitoring before betting budget" (format-volatility.md).
That is a properly hedged reference file. The trouble is that the action list is the part people actually use: stop citing AI citation gains to justify listicle production, treat owned pages as the rising format. An instruction gets repeated in meetings while a caveat sits in a reference file two directories down, so by the time this reaches you from a colleague or a consultant it has compressed to three words: ChatGPT killed listicles.
Someone re-ran it and got the opposite sign
The file's own advice, verify against your own monitoring, is exactly what one researcher did. Nicolas Sitter has been capturing 616 destination-hotel prompts across 56 destinations every week since December 2025, and compared the two weeks before ChatGPT 5.6 became the consumer default on August 6 against the two weeks after (Nicolas Sitter).
In that corpus the listicle share of citations rose from 33.8 percent to 36.4 percent. Fan-out queries, the extra searches ChatGPT fires off behind a single question, stayed flat at 0.7 percent growth against a reported 154 percent rise. The site: operator, which restricts a search to a single domain, appeared in none of the captured chats, against the 18.4 percent the original analysis reported.
Sitter is careful about what this proves: 616 hotel prompts against roughly a million across many categories, a prompt library that leaves out the "X vs Y" queries where a comparison-page effect would be most likely to surface, and a citation classifier not built the same way as Peec AI's. Sitter's conclusion is narrow and I think correct: "Don't assume the 'listicles are dying' narrative applies to your category without checking your own vertical's data."
So the honest reading is that the effect is probably real in some verticals and absent in others, and nothing in the version of the number that reached you tells you which side of that line your market sits on.
Your instruments do not agree on what a citation is
Checking your own vertical is the right instinct, and it is also where the second problem starts. 97th Floor lined up three measurements of one quantity, Reddit's share of ChatGPT citations, across the same stretch of summer. Promptwatch put the pre-drop share at 3.83 percent. Ahrefs, measuring the same period, put it at 16.7 percent. Gizmodo cited 4.5 percent (97th Floor).
Those are three answers to one question, and the largest is more than four times the smallest. 97th Floor's explanation is that the tools run "different prompt sets, different databases, different definitions of what counts as a 'citation' in the first place."
The definition is the part that should worry you, because it is the only one you cannot fix by working harder. Run-to-run noise has a known remedy, and the same skill library ships it: three to five runs per query, a mention rate reported with its sample size, rates tracked over time instead of one run against the next. That is good practice and it does nothing here. Averaging more runs of a tool that counts a link in a collapsed source tray as a citation will never converge on the number from a tool that does not.
Warning
If a slide compares your AI citation share against a published industry benchmark, check whether both figures came out of the same tool. If they did not, most of the gap you are looking at is two vendors disagreeing about what a citation is.
97th Floor lands on treating citation share as a "directional signal, not a KPI." I would go narrower: watch the direction inside one tool, on your own prompts, over time, and treat the absolute number as something that cannot be set beside anyone else's.
What a number you can act on looks like
That narrower rule is one we owe you. A fortnight ago we argued that AI visibility has drifted far enough from search to deserve its own channel and its own reporting (Promptafire). That still holds, and this is the half we left out: a channel you track needs an instrument you can trust, and this one is not there yet.
Kevin Indig's listicle study, published the same week, shows what the alternative looks like in practice. It found that vendor-owned B2B comparison pages put their own product first 74.7 percent of the time, and that those self-promoting pages gained traffic more often than the benchmark, 90.4 percent against 67.3 percent, with a median gain of 186 percent against 20 percent among pages that already had traffic (Growth Memo). That is exactly the shape of a statistic that gets screenshotted into "self-promotion works now."
Indig then did the thing that stops it. After adjusting for starting traffic, industry, listicle type, current rank and repeated pages from the same publisher, the self-promotion advantage fell to about 4 percent, with a range running from 51 percent lower to 120 percent higher. The data cannot tell the two groups apart, Indig says so plainly, and offers a different reading of the growth: companies put more work into listicles after noticing they mattered in AI search, rather than gaining anything from ranking themselves first.
The same dataset and the same author give you two completely different articles depending on which paragraph you stop at. What makes a number usable is not the size of the effect but whether its author showed you what happened when they tried to make it go away.
So before you cut a format, find three things: the vertical the finding was measured in, the run count and sample size behind it, and whether anyone outside the original vendor has reproduced it. If any of the three is missing, you have a reason to start watching something, not a reason to change a plan.
This costs something real, because it is slower and slower sometimes loses. If the demotion is real in software and B2B, the teams that cut listicles in August on one unreplicated number are ahead of the teams still waiting for a second measurement. Insisting on replication costs you every case where the rumour turned out to be true. What it buys is a content calendar that survives the next release note.
