Measured, not guessed

What do the law firm websites AI cites have in common?

The short answer: the sites AI engines cite look systematically different from the sites they never touch, and the gaps are far too large to be chance. Cited sites carry almost three times the homepage text, two and a half times the legal structured data, and sit on domains registered a median of five years earlier. But the same comparison contains a warning for anyone selling a quick fix: the cited sites also have more of the things AI cannot see at all, which means this is a portrait of the firms that win, not a checklist that makes you win.

The setup

We track the questions New Yorkers type when they need a lawyer, across ChatGPT, Gemini, and Google AI Overviews, and record every page each answer cites. Over seven weeks ending August 21, 2026, that produced 8,935 completed answer samples. We joined those citations against our August census of 2,626 New York firm websites, which fingerprinted each reachable homepage for platform, tracking, and structured data. Of the 2,365 sites the census could read, 583 were cited at least once in an AI answer during the window. 1,782 never were. The hundred most-cited sites each appeared at least 73 times.

That split lets us ask a clean question: on the signals we measured, do the top hundred look different from the sites the engines never cite? With groups this size, differences of a few points would be noise. These are not a few points.

Cited sites have more of everything, including what AI cannot seeTop-100 cited (n=100) vs never-cited (n=1,782) NYC law firm homepages, July to August 2026
Top-100 citedNever cited
Signals AI engines read
Legal-service structured data59%22%
FAQ structured data16%3%
Meta description91%72%
Homepage text, median words2,127781
Signals AI engines cannot see
Google Analytics or Tag Manager93%57%
Call-tracking service29%6%
Meta (Facebook) pixel18%7%
Beyond the page
Domain registered, median year20072012
Registered before 201583%61%

Share of each group's homepages carrying the signal, unless the row says otherwise. Every gap shown is statistically significant at p < 0.01 or far beyond; methodology and signal definitions are in the census explainer.

The gaps that jump out

  • Content volume is the widest gap of all. The median top-100 homepage carries 2,127 words of visible text. The median never-cited homepage carries 781. An engine assembling an answer simply has three times as much to work with, quote, and cite.
  • Legal structured data separates cited from invisible. 59 percent of top-cited sites mark themselves up as a legal service, against 22 percent of never-cited sites. Notably, the share is nearly identical between the top hundred and the other 483 cited sites, so markup tracks whether a site gets cited at all, not how often.
  • Age compounds. The median top-100 domain was registered in 2007; the median never-cited domain in 2012. 83 percent of the top hundred predate 2015, against 61 percent of the never-cited. Nothing on this list is harder to imitate: you can add markup this afternoon, but you cannot register your domain in 2007 this afternoon.

The warning in the middle of the table

Now look at the signals AI engines cannot see. A call-tracking script swaps phone numbers for human visitors; an analytics tag measures them; an ad pixel retargets them. No answer engine reads any of that, yet all three are dramatically more common on cited sites, and call tracking is one of the single strongest separators we measured. Invisible signals cannot be causing citations. They are markers of something else: firms that invest seriously in their web presence end up with the visible signals and the invisible ones together.

That is why we will not tell you this data proves markup gets you cited. It does not, and we have been critical of exactly that style of overclaim when others do it. What the data shows is a consistent portrait: the firms AI engines cite run substantial, well-maintained, machine-legible websites, usually on domains with a long history. Which parts of that portrait cause the citations, and which merely travel with them, cannot be untangled from a snapshot, however large. Untangling it takes watching sites change and answers move, which is what our tracking does day by day.

What this means if you are not cited yet

  • Start with the parts of the portrait you can control this quarter. Content depth and structured data are the two visible signals with the largest gaps, and both are ordinary website work. The census explainer's checklist covers the mechanics, including the robots.txt check that costs nothing.
  • Treat age as a reason to start, not a verdict. You cannot backdate your domain, which is precisely why the controllable signals matter more for you, and why switching domains casually is a bigger decision than it looks.
  • Distrust anyone selling the portrait as a recipe. If a vendor cites correlations like these as proof their service gets you cited, remember the call-tracking row. Winning firms have more of everything; that is not evidence any one thing works.
  • Measure the outcome, not the inputs. The only way to know whether AI engines cite you is to ask them the questions your clients ask, repeatedly, and watch. That is what we measure, every day, per engine.

Usual caveats: one metro, homepages only, one seven-week citation window, and "never cited" means never cited in the questions we track, not never cited anywhere. The full methodology, signal definitions, and the census this builds on are in the companion explainer.

See your own numbers

Ayo tracks how often ChatGPT, Gemini, and Google AI Overviews recommend your firm, question by question, updated daily, with the specific changes that move it. 7-day free trial, no card required.

Track your firm