The ai-discovery radar
We measure what machines actually find on websites — and publish it, including the uncomfortable parts. A great deal is claimed about AI visibility and very little is measured. Yet the routes a website offers machines, so that it can be found and read correctly, can simply be counted. That is what we do: monthly, in public, with the method laid open.
Dataset
AI Discovery Radar — adoption measurements for machine-readable discovery and consent routes
Monthly measurements of how many websites in a frozen, stratified sample (Tranco + Chrome UX Report country lists) publish machine-readable discovery and consent files (robots.txt, llms.txt, TDMRep, agent-card.json and others). Shares over observed hosts only, Wilson 95% intervals; full ruleset with version history in RULESET.md. Data files carry run identifier and ruleset version; numbers across ruleset versions are not comparable. Published by Berger+Team, Bolzano — project page: https://btlabs.dev/en/ai-discovery-radar
Repository recordDownload file · 43 KBMethod report (DOI)
How to cite
Berger, Florian (2026). AI Discovery Radar — adoption measurements for machine-readable discovery and consent routes (v2026.09) [Dataset]. Zenodo. https://doi.org/10.5281/zenodo.22178282
What gets measured
The radar checks which machine-readable files a website makes available to machines: robots.txt, which for thirty years has governed where automated visitors may go. llms.txt, a table of contents written for language models. ai-catalog.json, which describes the content and services a domain offers AI agents. 34 such routes are in the running measurement.
Then there is what we deliberately do not measure. Another 24 routes sit under observation: known to us, but not yet widespread enough to justify a request on every host in the sample. And 48 we examined and rejected, with a reason recorded entry by entry. A catalogue does not get better by admitting every format someone happens to mention.
Only public configuration files meant for machines are fetched — no page content, no images, no text. The measurement obeys each domain's robots.txt and identifies itself by name.
What the panel shows
As of September 2026, measured across a stratified sample of 853 reachable domains:
- 86.2% serve a
robots.txt. The oldest standard is the only one nearly everyone honours. - 13.5% serve an
llms.txt. The format is young, and adoption reflects that. - 0.0% serve an
ai-catalog.json. Of the 34 routes probed, 16 returned not a single response across the entire panel.
Those zeros are not a measurement error. They are the result. Many of these formats are discussed as though they were established. On the open web, they are not.
Since September 2026 the same panel of 1,000 sources is measured again every month; the figures above are the latest run and comparable month over month. The August 2026 exploration series across 16,554 domains remains documented as the baseline in the GitHub repository.
Why we publish this
We build websites meant to be found by people and by AI systems alike. That work needs numbers instead of assumptions: which route actually carries weight today, and which is a bet on tomorrow? Anyone who does not measure is selling guesswork.
So the method is public — sample, ruleset, confidence intervals, and the routes that came back at zero stay in the table. Anyone who doubts one of these figures can recompute it. That is the whole point.
The full method — classification model, denominator rule, sampling design, verification procedure and the measurement pitfalls that cost us data ourselves — is in the Technical Report v1.0 (Zenodo, DOI 10.5281/zenodo.22769680, CC BY 4.0). Anyone citing a share should have read it first.
Who does not want to be measured
No questions, no reason needed. An email with the domain is enough, or a single line in your own robots.txt. Excluded domains are skipped before any request is made.
Questions about the measurement?
Write to us — about the method, about a single figure, or if your domain should not be measured.
- Reply usually within 24 hours
- Opt out with no questions asked
- Method fully disclosed