Is your law firm's website built for AI search?
The short answer: probably not, and probably not for the reason you'd guess. In August 2026 we fetched the homepages of 2,626 New York law firm websites from our tracking registry and checked what each one is actually built to do. The pattern is consistent. Firms have invested heavily in measuring the visitors they get: analytics on 64 percent of reachable sites, plus chat widgets, call tracking, and ad pixels. Meanwhile the machine-readable signals that AI answer engines read are nearly absent; only 6 percent mark up their FAQs. And about one site in sixteen tells AI crawlers, in its robots.txt file, to go away entirely. We'd bet most of those firms don't know they're doing it.
How we checked
Our registry lists the New York firms and lawyers that AI engines name and cite when New Yorkers ask for a lawyer. For each of the 2,626 firm websites in it, we fetched the homepage once, plus the site's robots.txt file, and looked for standard platform and markup fingerprints: what the site runs on, what tracking it loads, and what structured data it publishes. 2,365 homepages came back readable. The rest split two ways, and both failure modes are findings in their own right. We'll come back to them.
The good news: your stack is not the problem
Among the sites whose platform we could identify, 84 percent run WordPress, with Squarespace and Wix covering most of the rest. Agency-branded proprietary builds are a small minority. That matters because it means almost no firm is trapped: whatever your site is missing, you are one plugin or one afternoon of edits away on a platform every web developer knows. The gap we found is not a technology gap. It is an attention gap, and it points in one direction.
Instrumented for clicks, bare for answers
Here is the shape of it. The first group is the tools firms install to measure and convert the visitors they already get. The second group is the structured-data signals that help machines understand who the firm is, what it does, and what questions it answers. Search engines read those signals, and the AI answer engines are built on top of what search engines read.
Built to measure visitors
Built to feed AI answers
Structured data = schema.org JSON-LD by type. 78% of sites emit some JSON-LD, but most of it is SEO-plugin boilerplate (breadcrumbs, image objects); the shares above count markup that actually describes a law practice.
None of the first group is wasted money. Analytics, chat, and call tracking are how a firm learns whether its website earns its keep, in a world where clients arrive by clicking. The problem is what the contrast reveals about where attention went. When a potential client asks ChatGPT or Google's AI for a lawyer, the "visit" happens on the engine's side: it reads your pages, composes an answer, and most of the time nobody clicks through at all. Every tool in the first group is blind to that interaction. The signals that matter for it are the ones in the second group, and those are the ones almost nobody ships.
What the answer engines actually read
Two honest cautions before the advice. First, structured data is not a magic citation lever; no one outside the engine companies can prove "add FAQ markup, get cited." What it verifiably does is make your facts machine-legible: schema.org markup is how you state, in a format Google documents and parses, that you are a law firm, in this borough, handling these practice areas, answering these questions. AI Overviews is assembled on top of Google's index, and ChatGPT's browsing reads the same pages. Second, missing markup does not make you invisible. In fact, firm websites are already the most-cited sources in AI answers about NYC lawyers; the engines lean on them heavily even though, as we just measured, most are barely optimized. Read those two facts together and the picture is not doom. It is a low bar: the engines want to read firm websites, and almost every firm is making it needlessly hard.
The sites AI cannot read at all
Three failure modes from our census are worth naming, because the firms involved almost certainly don't know.
- About 6 percent of sites tell AI, in robots.txt, to leave.3.0 percent of the robots.txt files we fetched block at least one AI crawler by name; GPTBot, ClaudeBot, and Google-Extended are the usual trio. Another 3.1 percent block every crawler, Google included, with a blanket disallow rule. The name-block pattern matches what security services like Cloudflare offer as a one-toggle "block AI bots" feature. We know how easy it is to have this on without deciding to: we found the same managed toggle enabled on our own site this month and turned it off. A firm paying for SEO while its robots.txt turns AI crawlers away is paying to be legible to one kind of machine while locking out the other.
- 88 sites render as empty shells without JavaScript. Their homepages arrive as a script bundle and nothing else; any reader that does not execute JavaScript, which includes most AI crawlers most of the time, sees a blank page. Another 5 percent of readable homepages carry fewer than 150 words of visible text: technically present, effectively silent.
- 173 of 2,626 registry domains are dead. That is 6.6 percent of the websites that directories and public records still list for these firms, pointing at domains that no longer resolve at all. An AI engine that follows a stale listing to a dead domain learns nothing about the firm, and there is nobody home to notice.
One smaller neglect signal in the same spirit: 17 percent of readable sites still ship the legacy Universal Analytics tag that Google switched off in 2023. Three years of collecting nothing. Websites drift when nobody is watching them; the answer era just raises the price of not watching.
What to actually do
- Read your own robots.txt. Open
yourfirm.com/robots.txtand look for GPTBot, ClaudeBot, Google-Extended, or a bareDisallow: /. If your site sits behind Cloudflare or a security service, check its dashboard for an AI-bot-blocking toggle someone may have enabled by default. - Check what a machine sees. View your homepage's source, or turn JavaScript off and reload. If the text of who you are and what you do isn't there, AI readers get the blank page, not the site you paid for.
- Mark up what you already wrote. If you have an FAQ page, FAQ structured data is an afternoon of work on WordPress, and 94 percent of your competitors haven't done it. The same goes for stating your practice areas and location as legal-service markup rather than only prose.
- Then measure the surface your analytics can't see. Everything above makes you more legible to AI engines; none of it tells you what they currently say about you. That is a different measurement. You can spot-check it by hand, or see which firms the engines already recommend on our public leaderboards.
The usual caveats apply to our numbers: one metro, homepages only, fetched in one August 2026 pass, and presence of a signal is not proof it moves any particular answer. But the aggregate picture doesn't need fine print. The legal web is instrumented, thoroughly, for a funnel that AI answers are starting to bypass, while the machine-readable layer those answers are built from sits mostly empty. For now, that emptiness is an opportunity for whoever fills it first.
See your own numbers
Ayo tracks how often ChatGPT, Gemini, and Google AI Overviews recommend your firm, question by question, updated daily, with the specific changes that move it. 7-day free trial, no card required.