We requested /robots.txt from 3742 US dental practice domains in 49 states on 3 October 2026 and parsed the 2503 files that came back. The short answer: practices are not blocking AI crawlers — only 1.5% of the files block one, while 10.3% explicitly allow one. The real damage is elsewhere: 3.6% block their whole site from every crawler, and 10.4% of domains refuse unknown crawlers at the server, where no robots.txt line can fix it.
What we asked and what we did not
The frame is the same as our study of 832 US dental practice websites: dentist listings in OpenStreetMap that publish a website. For this one we fetched a single file per domain — /robots.txt, nothing else — and parsed it by the RFC 9309 rules, counting a crawler as blocked when its own group disallows the site root. We did not judge whether blocking AI is right; we measured what the files say. No practice is named.
What the domains answered
| Response | Share of domains |
|---|---|
| Served a robots.txt file | 66.9% |
| Server refused an unknown crawler (401/403) | 10.4% |
| Host unreachable, DNS or TLS failure | 16.2% |
| No robots.txt (404) | 3.3% |
| Other response (redirect chain, 202, 405) | 3.1% |
One HTTPS request per domain, nothing else fetched.
Almost nobody blocks AI crawlers
Across the 2503 files, the most-blocked crawlers are not the assistants at all: Bytespider at 1.2% and Common Crawl at 1.1%, both of which arrive inside default block lists. GPTBot is blocked by 0.4% and ClaudeBot by 0.5%. Meanwhile 10.1% of files name Google-Extended and 7.4% name GPTBot — almost always to allow them. If you were told dental practices are shutting AI out, the files disagree.
Who is named, who is blocked
| Crawler | Named in the file | Blocked |
|---|---|---|
| GPTBot (OpenAI training) | 7.4% | 0.4% |
| OAI-SearchBot (ChatGPT search) | 5.2% | 0.0% |
| ChatGPT-User (live fetch) | 4.6% | 0.1% |
| ClaudeBot (Anthropic) | 7.5% | 0.5% |
| PerplexityBot | 2.7% | 0.1% |
| Perplexity-User (live fetch) | 4.1% | 0.0% |
| Google-Extended (Gemini) | 10.1% | 0.3% |
| Applebot-Extended | 5.8% | 0.4% |
| CCBot (Common Crawl) | 6.7% | 1.1% |
| Bytespider (ByteDance) | 5.9% | 1.2% |
| Meta-ExternalAgent | 5.8% | 0.5% |
| Googlebot (search) | 6.1% | 0.0% |
“Named” means the file has a group for that crawler at all — most of those groups allow it.
The two settings that actually cost visibility
3.6% of files disallow the entire site for every crawler. That is one line — Disallow: / under User-agent: * — usually left over from a staging site that went live. It removes the practice from Google as surely as from ChatGPT. 10.4% of domains refused us at the server with a 401 or 403 before robots.txt was even read. That is bot protection, and it usually keeps verified Googlebot while turning away crawlers it does not recognise — including the live fetchers assistants use when a person asks about your practice. Neither setting shows up in an SEO report, and both stay invisible until someone checks.
Inside the files
| What we found | Share of files |
|---|---|
| Declares a sitemap | 80.1% |
| Names at least one AI crawler | 10.9% |
| Explicitly allows at least one AI crawler | 10.3% |
| Blocks at least one AI crawler | 1.5% |
| Blocks the whole site for every crawler | 3.6% |
| No rules at all | 6.2% |
Median file size: 177 bytes — most are the default a theme or plugin wrote.
Training crawler, search crawler, live fetcher — three different decisions
The names in that table are not interchangeable, and treating them as one group is why most files say something their owner never meant. GPTBot and Google-Extended collect text that trains a model. OAI-SearchBot and PerplexityBot build the index an assistant searches. ChatGPT-User, Claude-User and Perplexity-User fetch your page in the moment a person asks about your practice. A clinic can refuse training and still want to be found: allow the search and live-fetch agents, disallow the training ones. Almost none of the files we read make that distinction — one group, one decision, usually written by a plugin that shipped a block list two years ago.
What a workable robots.txt looks like for a practice
- One group for everyone — User-agent: * with Allow: / — and disallow only what should never be indexed: the cart, the staff login, internal search results.
- A Sitemap: line with the real sitemap URL, because 19.9% of files skip it and make crawlers find the pages the slow way.
- A deliberate line on AI: name the training crawlers if you want them out, and leave the search and live-fetch agents allowed if you want to appear in assistant answers.
- Nothing else. The median file we read is 177 bytes, and the ones doing damage are rarely the long ones.
What to check on your own site this week
- Open yourdomain.com/robots.txt in a browser. If you see Disallow: / under User-agent: *, that is an emergency, not a nuance.
- Confirm the file declares your sitemap — one in five does not.
- Decide about AI crawlers deliberately: allow them if you want to be quotable in assistants, block them if you do not, but do it on purpose rather than by plugin default.
- Ask whoever runs your firewall or CDN which crawlers get a 403. Bot protection is where most invisible blocking happens.
- Re-check after every redesign or host migration — that is when a staging rule ships to production.
Allowing crawlers is step one, being quotable is the rest
Access only gets an assistant to the page. What it quotes is the clear answer, the sourced number, the FAQ it can lift — that part is covered in AI search optimisation for clinics, and it is the work behind our search and AI visibility service for dental practices.
Method, limits and the raw data
Recorded 3 October 2026. Frame: dentist listings with a website in OpenStreetMap across the contiguous US; 3742 domains attempted, 2503 files parsed, 49 states represented. Limits worth stating: the domains that refused us at the server are missing from the parsed set, so the file-level shares describe the more open half of the market; robots.txt is a request rather than an access control, and a crawler can ignore it; server-level protection can block a crawler the file allows, and allow one the file blocks; and the OpenStreetMap frame skews toward practices someone bothered to map. Download the aggregates as CSV or JSON, published under CC BY 4.0 — use any figure with a link to this page, citing “Tepexa, Do dental websites block AI crawlers? 2026”.