Why this exists. Every roundup of AI visibility tools ranks the tools. None of them measured what the engines actually cite. We ran the checks and logged every source URL, so here is the data instead of an opinion. Use it, quote it, argue with it. Every number carries its sample size.
From 99 citations across 5 real dated business checks, asking the question a customer would ask, such as who is the best roofer in a given city.
| Source type | Share of citations |
|---|---|
| The business's own website | 49.5 percent |
| Directory and profile listings | 20.2 percent |
| Community and social | 15.2 percent |
| Roundup listicles | 2.0 percent |
Inside the directory bucket, the specific sources cited for the businesses that won: BBB cited 4 times, Yelp 3, city government and utility pages 6, plus Nextdoor and Mapquest. Inside community, Facebook pages carried 3 citations for winners. Every one of those is free to claim.
From 93 citations on the other kind of question entirely, the best X tool kind, which is how buyers find software.
| Source type | Share of citations |
|---|---|
| Vendor websites | 40.9 percent |
| Community | 24.7 percent |
| Roundup listicles | 17.2 percent |
| News | 7.5 percent |
The two segments do not behave the same way, and it is not close. Roundup listicles carry roughly eight times more weight on a software or vendor search than on a local business search: 17.2 percent versus 2.0 percent. Community more than doubles too, 24.7 versus 15.2.
That is why generic AEO advice fails. A checklist that tells a plumber to chase listicle placements is spending their time on the source type that carried 2 percent of the citations that named their competitors. A checklist that tells a SaaS founder their own site is nearly everything is understating how much the community and the lists decide it.
Practical version. Local business: answer-shaped pages on your own domain, then claim the free directory ring. That is roughly 70 percent of what named the winners. Software or vendor: your own site still leads, but the lists and the forums are where you are actually chosen, and being absent from them is invisible in a way no amount of on-site work fixes.
| Measure | Result | Sample |
|---|---|---|
| Engines disagreed on whether the business was named at all | 20.0 percent of checks | 2 of 10 businesses |
| Of the companies the terser engine named, share the other engine also named | 44.1 percent | 10 businesses |
| Checks where the two engines shared no company at all | 2 | of 10 |
In both disagreements the pattern ran the same way: Perplexity named the business and Gemini did not. Two real businesses would have been told they are visible by one tool and invisible by another, on the same day, from the same question.
| Measure | Result |
|---|---|
| Answers that were byte-identical between the two runs | 0 of 22 |
| Checks where the named / not-named verdict flipped between runs | 1 of 22, 4.5 percent |
The same engine, asked the same question on the same day, never gave the same answer twice. Not once in 22 comparisons. And in one case the variation was large enough to flip the verdict: Gemini did not name World Martial Arts Academy on the first run and did name it on the second, minutes later.
This is the most important thing on this page, because it applies to every other number here including ours. A single AI-visibility check is a sample, not a measurement. A score generated from one run of one engine is a coin landing, reported as a fact. That is true of every tool in this category, and it is true of the free check on this site.
It is also the honest reason a dated transcript beats a score. A transcript carries its own timestamp and its own caveat. A number out of 100 hides both.
An earlier version of this page reported engine disagreement at 27.3 percent. On the next day's run the same measurement gave 20.0 percent. We did not quietly swap it. Both figures were correct for the run that produced them, and the gap between them is Finding 4 doing its work on our own research.
We also deleted two figures entirely. We first measured engine agreement as the overlap between the sets of companies each engine named, and got 2.5 percent, then 4.2 percent after normalising company names. Both looked dramatic. Both were wrong. Gemini writes 4,279 characters per answer to Perplexity's 1,920. When one engine names 27 companies and the other names 8, perfect agreement on all 8 still caps set overlap near 30 percent. The metric had a ceiling set by verbosity, not by disagreement.
The corrected measure asks a directional question: of the companies the terser engine named, how many did the other one also name? That moved the result from 4.2 percent to 44.1 percent, an order of magnitude. The alarming number was the bug.
We publish the corrections because a research page that only shows its surviving numbers is an advertisement. If you cite this work, you should know which of our figures died and what killed them.
Buyer-shaped questions were asked on live AI engines and every cited source URL in each answer was recorded and classified by source type. Totals are counts of citations, not counts of answers, so an answer citing four directories contributes four rows. Checks are dated individually. The businesses in the local sample are real and their results include the ones that were named by nobody.
On Findings 3 and 4 specifically: one day, one prompt variant, two engines, 10 businesses for agreement and 22 comparable rows for the stability test. It is a first measurement, not a constant, and repeated dated runs could move it. Company-name matching is mechanical, using canonical forms and whole-word containment, so a company written two very different ways could still be counted as two. Errors run toward under-counting agreement, which means the true agreement figure is more likely higher than 44.1 percent than lower.
This is first-party data, not a third-party study. The local sample is 5 business checks producing 99 citations, which is small. Some captures are single-engine rather than all five. The category sample of 93 citations comes from our own checks on our own category, so we are inside the sample we are measuring. AI answers also personalize and drift over time, so any of these numbers can move. We publish the method so you can run it yourself and disagree with us using data instead of vibes.
Namebeam checks whether ChatGPT, Claude, Perplexity, Gemini and Google AI Overview name a business when a customer asks for a recommendation, and returns the dated raw transcript rather than a score.
We are our own first customer, and we are losing. Namebeam scored zero out of five on its own first check and we published that, dated. We have zero paying customers as of 2026-08-04. Our running record, wins and losses printed the same size, is public at proof.namebeam.ai. We think you should audit a vendor before you pay one, so we made ourselves auditable first.
Free to quote with attribution. Suggested line: Namebeam AI Citation Study, 2026, namebeam.ai/ai-citation-study.html. If you are maintaining a roundup of AI visibility tools and want the underlying method, write to [email protected] and we will send it.
Run the free check on your own business