Measured, not guessed

Do legal directories control which lawyers AI recommends?

The short answer: not in the questions real clients ask. A claim circulating in AI-search circles says the AI engines are mirrors of the legal industry's existing authority structure: that Chambers, Super Lawyers, and the peer-reviewed directories decide who gets recommended, and that in large markets a firm's own website is essentially invisible to AI. Our measurement says otherwise. Over the last 30 days, across 24,345 citations in AI answers to NYC lawyer questions, the single most cited domain was a law firm's own website, seven of the top eight cited domains were firm websites, and every directory combined accounted for about 4 percent of citations. The mirror story is tidy and quotable. It does not survive contact with repeated measurement.

The claim, and where it comes from

The studies behind this claim tend to share a method: take a question like "best personal injury lawyer in Chicago", run it once per engine across a handful of markets and practice areas, note who appears, and read the pattern as the ranking. From one such pass, the conclusions write themselves: the engines reward decades of directory recognition, firm websites only matter in small markets, and the industry's old authority structure has simply been re-projected onto a new surface.

The problem is not the ambition. It is that a single run per query is a single draw from a distribution with enormous variance, taken on one day of a system that changes week to week. Both halves of that sentence are measurable, and we measure them daily.

How we know

Ayo asks ChatGPT, Gemini, and Google AI Overviews the questions New Yorkers type when they need a lawyer, every day, many times per question, and records every firm each answer names and every page each answer cites. We ask the consumer products themselves, not the vendor APIs (why that distinction matters). The numbers below are from the 30-day window ending August 12, 2026: 5,773 completed samples, of which 4,836 cited at least one source, for 24,345 citations in total.

What 24,345 citations say about directories

New York is the largest legal market in the country, which under the mirror theory should make it the clearest case of directories dominating and firm websites disappearing. Here is what the engines actually cited:

  • The most cited domain is a firm's own website. Jacoby & Meyers' site leads all sources at 3.9 percent of every citation we recorded, ahead of every directory, every court site, and every news outlet.
  • Seven of the eight most cited domains are firm websites. The one exception is nyc.gov. The first legal directory, Justia, appears tenth at 1.4 percent, with Super Lawyers close behind at 1.2 percent.
  • All directories combined are a rounding story, not a power structure. We tallied seventeen well-known legal directories, from Avvo and FindLaw to Martindale, Best Lawyers, and Chambers. Together they account for 3.8 percent of citations, and the picture holds engine by engine: about 5 percent of ChatGPT's citations, 4 percent of Gemini's, and 3 percent of AI Overviews'. Add general review platforms like Yelp and Reddit and the combined share is still under 5 percent.
  • The pattern is unanimous across engines. On each of the three engines separately, the three most cited domains are all law firm websites.

A firm website being the top cited source in the country's largest legal market is the direct opposite of "firm websites are invisible in large markets". Whatever is true of nine queries run once, it is not true of six thousand answers measured over a month.

Citations measure what the engines read and show, not everything that goes into a name, so it is fair to ask whether directories could still steer the names invisibly. Who gets named points the same way: the firms each engine favors track that engine's own taste, a neighborhood firm on ChatGPT, a household brand on Gemini, an SEO-strong local practice on AI Overviews (three engines, three different front-runners), not any directory's rankings. And those front-runners flip within weeks, which a stable mirror of decades-old peer review cannot do.

Why one-time studies see a different world

When we sample the same lawyer question only a handful of times in a day, the day-to-day "leader" flips constantly: on typical question-engine series in one 19-day stretch we measured, the daily front-runner changed on 16 of 18 consecutive day pairs, with single-day swings as large as 47 points. That is not the market moving. That is what variance looks like at a tiny sample size, and one run per engine is the tiniest sample there is. It takes on the order of 100 samples per question per engine before the error bars shrink to around 9 points and a leader claim means anything at all.

And even a properly sampled snapshot expires. When we deep-sampled the same questions twice, 18 days apart, the most-named firm changed on all three engines, by margins far outside the error bars, with model releases moving citation patterns overnight (the three mechanisms behind the churn). A study that queries each engine once, on one day, is a photograph of traffic: real cars, real positions, no information about where any of them is going.

What directories are actually good for

None of this means directories are worthless. They show up as citations on every engine, steadily, in the low single digits. And a complete, accurate directory presence is still cheap insurance for a different reason: directories rank well in the ordinary Google results that AI Overviews is assembled from, so they can help you get named even where they are rarely cited. The honest reading of our data is narrower than the mirror theory: directories are one modest source among many, not the gate. The engines also disagree with each other about everything else, including which firms they favor and which pages they trust (same question, three different front-runners), so a single lever that explains all three engines at once should be suspect on its face.

One honest caveat in the other direction: we measure the questions consumers actually type, in one metro. Segments we do not track, like the corporate and private-wealth work where Chambers rankings live, could behave differently. If someone measures those segments with real sample sizes over a stated window, we will read it with interest. That is the standard the claim needs to meet.

How to read the next AI visibility study

  • Ask how many times each query ran. Once is an anecdote. A ranking needs enough samples for error bars, and the error bars need to be smaller than the gaps between firms.
  • Ask for the date window. AI answers change measurably within weeks. An undated finding is a finding about a day someone will not tell you.
  • Ask which surface was measured. The developer APIs and the consumer products give different answers, and third-party tools often measure the former while claiming the latter.
  • Then check yourself. Our public leaderboards show which firms the engines actually recommend in NYC, per engine, with the window stated on every page.

See your own numbers

Ayo tracks how often ChatGPT, Gemini, and Google AI Overviews recommend your firm, question by question, updated daily, with the specific changes that move it. 7-day free trial, no card required.

Track your firm