---
title: "The plumbing nobody explained: five files that decide if an AI can book you"
description: "robots.txt, structured data, llms.txt, MCP and WebMCP, in plain English, and which ones a gym, studio or spa should fix first. Part 2 of Beyond the Four Walls."
section: AI and Agents
author: David Steel
published: 2026-10-10T19:43:07.329Z
url: https://one.sneeze.it/blog/the-plumbing-nobody-explained
tags: ["beyond-the-four-walls", "robots-txt", "structured-data", "webmcp", "mcp"]
---

# The plumbing nobody explained: five files that decide if an AI can book you

_robots.txt, structured data, llms.txt, MCP and WebMCP, in plain English, and which ones a gym, studio or spa should fix first. Part 2 of Beyond the Four Walls._

Every business owner knows the feeling. The studio looks great, the front desk is staffed, and then someone mentions that the hot water has been off in the back since Tuesday. The part customers see was never the problem. The plumbing was.

Websites now have the same split. The part you paid a designer for is the part people see. Behind it sit a handful of small files and settings that decide whether an AI assistant can get in, understand what you sell, and book something. Most owners have never been told these exist, because until this year they barely mattered.

In [Part 1](/blog/your-next-customer-wont-visit-your-website) we made the case that customers are starting to send software to shop for them. This part opens the wall and names the pipes. There are five worth knowing. You do not need to install any of them yourself, but you should know what to ask for, and in what order.

## Pipe one: robots.txt, the sign on the door

Every website can have a plain text file at `yoursite.com/robots.txt`. It is a sign on the door that tells automated visitors which rooms they may enter. On many sites it was written by a web developer years ago and never read since.

What changed is who is knocking. AI companies now send two very different kinds of visitor, and their own documentation keeps them apart.

**Training crawlers** read the web in bulk to build future AI models. OpenAI's is called GPTBot. Anthropic's is ClaudeBot. Common Crawl, a nonprofit that publishes an open archive of the web, runs CCBot. Google handles training differently: there is no separate crawler, but a setting called Google-Extended tells Google whether pages it has already crawled may be used to train its Gemini models.

**Live assistants** visit because a person just asked for something. ChatGPT-User, Claude-User and Perplexity-User are the ones a customer sends when they say "find me a facial next Thursday." Alongside them sit the search indexers that decide whether you show up in AI answers at all: OAI-SearchBot for ChatGPT search, Claude-SearchBot for Claude, PerplexityBot for Perplexity.

The companies are clear that these are separate switches. OpenAI's bot documentation says of its search and training crawlers that "each setting is independent of the others." Google says Google-Extended "does not impact a site's inclusion in Google Search nor is it used as a ranking signal." Anthropic spells out what happens if you turn away its live assistant:

:::pullquote
Disabling Claude-User on your site prevents our system from retrieving your content in response to a user query.
-- Anthropic, on its crawlers
:::

So blocking training crawlers is a business decision you are entitled to make. Blocking the live assistants is turning away a customer who is standing at the door with their phone out. One honest footnote: OpenAI says robots.txt rules "may not apply" to ChatGPT-User, and Perplexity says its user fetcher "generally ignores robots.txt rules," because a person asked for the page. That does not make the sign pointless. It means the bigger blocker is often not robots.txt at all but a security wall that shows a "checking your browser" page no assistant can pass, the problem we flagged in Part 1.

## Pipe two: structured data, your business card in machine print

Your hours are probably on your website. The question is whether they are in a form a machine can trust. A picture of your opening hours, or hours drawn on the page by a script after it loads, is invisible to most software.

Structured data fixes this. It is a short block of code, usually in a format called JSON-LD, that states your facts in a vocabulary maintained by schema.org. The core type is LocalBusiness, which schema.org defines as "a particular physical business or branch of an organization." There are more specific types that fit this industry closely: DaySpa, HealthClub, HairSalon, BeautySalon and NailSalon all sit under HealthAndBeautyBusiness, and gyms fit under SportsActivityLocation. The block carries your name, address, phone, opening hours and price range.

This one is not new and not speculative. Google's own guide says local business markup lets you "tell Google about business hours, different departments within a business," and every example on that page is written in JSON-LD. AI assistants read the same block. If you run several locations, each location page needs its own, with that location's real phone and hours.

## Pipe three: llms.txt, the cheat sheet

llms.txt is a short Markdown file at `yoursite.com/llms.txt` that briefs an AI on what your site is and where the important pages are. It was proposed by Jeremy Howard in September 2024, in plain words: "We propose adding a /llms.txt markdown file to websites to provide LLM-friendly content."

It is a proposal, not a standard, and the evidence on it is mixed. Adoption is growing among large sites. Whether anything reads the file is a different question.

:::stats
11.5% | Top 10,000 websites with a valid llms.txt, September 2026
97% | llms.txt files that got zero requests in May 2026
10,000+ | Active public MCP servers, December 2025
source: [Underneath](https://underneath.agency/research/llms-txt-adoption-study), [Ahrefs](https://ahrefs.com/blog/llmstxt-study/), [Anthropic](https://www.anthropic.com/news/donating-the-model-context-protocol-and-establishing-of-the-agentic-ai-foundation)
:::

Ahrefs looked at about 38,000 sites that publish the file and found 97% of those files were never requested in a month. We still recommend one, because it takes twenty minutes to write and costs nothing to host. Just do not expect it to move anything on its own. It is the cheat sheet you leave on the counter, not the reason anyone comes in.

## Pipes four and five: MCP and WebMCP, from reading to doing

The first three pipes help an assistant find you and understand you. The last two let it do something.

**MCP**, the Model Context Protocol, was introduced by Anthropic in November 2024 as "a new standard for connecting AI assistants to the systems where data lives." Its own site calls it a USB-C port for AI applications: one standard plug instead of a custom cable for every pairing. In December 2025 Anthropic gave MCP to the Agentic AI Foundation, a fund under the Linux Foundation co-founded by Anthropic, Block and OpenAI. At that point Anthropic counted more than 10,000 active public MCP servers and said MCP had been adopted by ChatGPT, Cursor, Gemini, Microsoft Copilot and Visual Studio Code. The specification is still being revised; the current documentation is versioned July 28, 2026.

For a local business, MCP mostly arrives through your software vendors. It is how an assistant might talk to a booking system, a CRM or a payments platform. You will rarely build one, but you should ask your booking platform whether they have one.

**WebMCP** brings the same idea to an ordinary web page. Instead of an assistant guessing which button means "book," the page hands it a short menu of named actions, like check availability or book a class, each with the fields it needs. Chrome's documentation describes it as "a proposed web standard to help you build and expose structured tools for AI agents." It is early: Chrome offers it as an origin trial from Chrome 149, and the page warns it "is under active discussion and subject to change." Stripe has already documented how agents can use WebMCP on Stripe's own checkout pages, while labeling it "an experimental browser capability" and telling agents to get the person's confirmation before submitting a payment. More on that in the next part.

:::figure wide
![A cutaway of a spa reception wall: five pipes of different widths run behind the plaster, four in grey and one in magenta, leading from the desk to a small robot hand holding a calendar](/blog/media/535d64767ca2aaa27dda7c29.jpg)
Five pipes behind the wall. The thin ones let an assistant in; the wide one lets it book.
:::

## Which pipe to fix first

Here is the whole stack in one table, ranked for a local or multi-location business. The order follows what an assistant needs first: get in, understand, act.

| Order | Layer | What it does | Status, October 2026 | Who fixes it |
|---|---|---|---|---|
| 1 | robots.txt and firewall | Lets live assistants in | Long established | Your web developer or host |
| 2 | Structured data (JSON-LD) | States hours, address, phone, prices | Established, used by Google | Your web developer |
| 3 | Plain text facts and real links | Phone and email as tappable links, hours as text | Basic web practice | Whoever edits your site |
| 4 | llms.txt | A briefing file for AI | Proposal; read rarely | You can write it |
| 5 | WebMCP tools | Lets an assistant book on your page | Chrome origin trial | Developer, or your booking vendor |
| 6 | MCP server | Lets an assistant talk to your systems | Widely adopted by AI apps | Usually your software vendor |

Our scoring tool, Agent Ready, weighs the same ground. It puts the most points on whether an assistant can act, because that is the step where a booking happens or does not.

:::chart column "Agent Ready: points per pillar, out of 100" highlight="Act"
Reach | 25
Comprehend | 25
Act | 35
Complete | 15
source: [Agent Ready methodology](https://agentready.sneeze.it/methodology). Reach covers robots.txt and access; Comprehend covers structured data and llms.txt; Act covers WebMCP, forms and booking.
:::

We tested this order on ourselves. Our own new site scored 39 out of 100 on October 9. The first pass of plumbing, the first four rows plus a sitemap and a real form, took about an hour and brought it to 89. Adding real WebMCP tools brought it to 97. [The full story is here](/blog/we-scored-our-own-website-39). The lesson for a gym or spa: the cheap plumbing gets you most of the way, and the last stretch is about letting an assistant actually book.

:::takeaways "What to ask for this month"
- Ask your web developer to confirm ChatGPT-User, Claude-User, Perplexity-User and the AI search bots are allowed, and that no security wall blocks them.
- Decide on training crawlers (GPTBot, ClaudeBot, CCBot, Google-Extended) as a separate question. Blocking them does not remove you from Google Search.
- Add one LocalBusiness block per location with the real phone, address and hours, using the closest type such as DaySpa or HealthClub.
- Put your phone and email in the page as real links, and your hours as text.
- Write a short llms.txt, but treat it as a nice extra.
- Ask your booking platform two questions: do you support MCP, and are you working on WebMCP?
:::

Next, in Part 3, we follow the money: [Stripe is rebuilding checkout for robots](/blog/stripe-is-rebuilding-checkout-for-robots), and what that means for a membership sale or a class pack bought by an assistant.

:::cta button="Scan your website free" href="https://agentready.sneeze.it"
Which of the five pipes is leaking on your website?
Agent Ready checks all of them in about a minute and lists the fixes in order of points.
:::

---

#### Sources

- OpenAI, "Overview of OpenAI Crawlers," accessed October 2026. [developers.openai.com](https://developers.openai.com/api/docs/bots)
- Anthropic, "Does Anthropic crawl data from the web, and how can site owners block the crawler?", April 7, 2026. [support.claude.com](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler)
- Perplexity, "Perplexity Crawlers," accessed October 2026. [docs.perplexity.ai](https://docs.perplexity.ai/guides/bots)
- Google Search Central, "Google's common crawlers" (Google-Extended), accessed October 2026. [developers.google.com](https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers)
- Common Crawl, "CCBot," accessed October 2026. [commoncrawl.org](https://commoncrawl.org/ccbot)
- schema.org, "LocalBusiness" and "HealthAndBeautyBusiness." [schema.org/LocalBusiness](https://schema.org/LocalBusiness), [schema.org/HealthAndBeautyBusiness](https://schema.org/HealthAndBeautyBusiness)
- Google Search Central, "Local Business (LocalBusiness) structured data," updated September 8, 2026. [developers.google.com](https://developers.google.com/search/docs/appearance/structured-data/local-business)
- Jeremy Howard, "The /llms.txt file," published September 3, 2024. [llmstxt.org](https://llmstxt.org/)
- Underneath, "How many websites have an llms.txt file? 2026 adoption data," September 26, 2026. [underneath.agency](https://underneath.agency/research/llms-txt-adoption-study)
- Ahrefs, "We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read," June 15, 2026. [ahrefs.com](https://ahrefs.com/blog/llmstxt-study/)
- Anthropic, "Introducing the Model Context Protocol," November 25, 2024. [anthropic.com](https://www.anthropic.com/news/model-context-protocol)
- Anthropic, "Donating the Model Context Protocol and establishing the Agentic AI Foundation," December 9, 2025. [anthropic.com](https://www.anthropic.com/news/donating-the-model-context-protocol-and-establishing-of-the-agentic-ai-foundation)
- Model Context Protocol, "What is the Model Context Protocol?", accessed October 2026. [modelcontextprotocol.io](https://modelcontextprotocol.io/)
- Chrome for Developers, "WebMCP," published May 18, 2026, updated October 7, 2026. [developer.chrome.com](https://developer.chrome.com/docs/ai/webmcp)
- Stripe, "Use WebMCP to complete Stripe payments in a browser," accessed October 2026. [docs.stripe.com](https://docs.stripe.com/agentic-commerce/for-agents/webmcp)
- Sneeze It, "Agent Ready methodology," accessed October 2026. [agentready.sneeze.it](https://agentready.sneeze.it/methodology)
