Minimalist mint-green-on-black illustration: a single document icon on the bottom step of a staircase rising into a more elaborate structure of connected nodes, data layers and a radiant endpoint node.

llms.txt is only step one — the architecture after

btlabs Core · Aug 3, 2026

At its core, llms.txt is nothing more than a table of contents pointing to Markdown files — a starting point, not a destination. That's how Duane Forrester, nearly 30 years in the search industry and involved in Schema.org's launch during his time at Bing, puts it in a widely discussed article on Search Engine Journal. Set up llms.txt once and consider the job done, and you've solved the easiest part of the task — and missed the bigger one.

What llms.txt actually is

llms.txt is a lightweight file at the root of your site that offers AI systems the most important content as a tidy list of links — without the navigation, adverts or script overhead around it. That's useful: an AI agent finds what matters faster, instead of having to pick its way through HTML. But a table of contents doesn't explain how your products relate to each other, which service belongs to which audience, or which claim is still current and which is outdated. That, according to Forrester, is the real gap: most website architectures were never built to give AI systems clean, authoritative access to brand information — llms.txt does little to change that if the same unstructured foundation sits underneath it.

The architecture that comes next

Forrester sketches out what has to follow the table of contents if you want an AI to understand you reliably:

  • Structured fact sheets (JSON-LD). Precise, machine-readable markup for organisation, services and reviews — not as a box-ticking exercise, but with genuinely accurate attributes.
  • Entity relationship mapping. How does a service connect to a category, a use case, or another service? Without that link, even well-marked-up facts stay isolated islands rather than a network.
  • Programmatic content access. Versioned interfaces for FAQs, documentation and specifications — in the direction standards like the Model Context Protocol (MCP) are already pointing.
  • Provenance metadata. Timestamps, authorship and traceable sources on every fact an AI might cite — without it, every claim stays unverifiable.

The point, in Forrester's words: this infrastructure turns your content from "something the AI read somewhere" into something verifiable.

Why this matters for the bottom line

For decision-makers, the practical question isn't whether llms.txt does any harm — it doesn't, it costs almost nothing and it's hard to get wrong. The question is whether a single file is enough for an AI to represent your business correctly and completely when someone asks. It isn't. Without structured facts, entity clarity and traceable provenance, an AI is left guessing — and guesswork is where most mis-citations, outdated claims and wrongly attributed services in an AI answer come from.

Building the full architecture is manageable if it's designed in from the start — and expensive if it has to be grafted onto a website that's already grown organically. That's exactly why btlabs Core builds this layer as the default state, not as an add-on step tacked on later: JSON-LD, entity relationships and an MCP endpoint are generated from the same underlying data as the human-facing site, not maintained separately and caught up later. One source of truth, two output layers — maintained together for people and machines alike.

An honest assessment

No major AI provider has formally committed to weighting llms.txt, or any specific architecture behind it, as a fixed signal — the standards around AI visibility are still evolving. That's not a reason to wait; it's a reason to invest in the cheap, robust building blocks: a correct data foundation with structured facts is useful to you regardless of how heavily any one AI provider weights it tomorrow. It's also, largely, the same work that classical discoverability already requires — clearly named entities, consistent information, traceable sources. Invest here and, worst case, you lose nothing; most likely, you gain an answer that's actually correct when someone asks.

The recommendation

Set up llms.txt if you haven't already — the effort is minimal. But treat it for what it is: a first, small building block, not the end of the task. The real work lies in a data foundation that automatically feeds structured facts, entity relationships and provenance metadata — for every service, every page, every language, maintained continuously rather than set up once. Get that right, and an AI asked about a business like yours can give a correct answer instead of a guessed one.

How much that granularity actually delivers is something we worked out in 2.3× more AI citations from granular structure.

Want to know where your digital foundation currently stands in this architecture? Get in touch — we'll take a look together.

Source

  • "Llms.txt Was Step One. Here's The Architecture That Comes Next" — Duane Forrester, Search Engine Journal, 2026. searchenginejournal.com

Frequently asked questions.

What is the difference between SEO, GEO and LLMO?

The three terms build on each other rather than replacing one another. SEO (Search Engine Optimisation) makes sure a page ranks well in classic search results. GEO (Generative Engine Optimisation) extends that goal to being cited as a source in AI-generated answers — in a ChatGPT or Perplexity answer, for example.

LLMO (Large Language Model Optimisation) goes a step further and gets more technical: it means specifically understanding and influencing how individual language models process, weigh and reproduce content — essentially reverse-engineering the model in question. In practice: SEO remains the entry ticket, GEO secures broad citability, and LLMO fine-tunes the details for the systems that matter most to a given business.

Do I absolutely need an llms.txt?

It is not a Google ranking signal (Google says so itself) — but it takes minutes to create and is useful as orientation for AI agents. A cheap addition, not a miracle cure.

Next step

Want to look at it together?

No pitch, no standard package — an honest take on your project. A reply, usually within 24h.

Start a conversation
Blogverzeichnis Bloggerei.de - Wissenschaftsblogs