Blog AI and Agents · Beyond the Four Walls · Part 2
The plumbing nobody explained: five files that decide if an AI can book you
robots.txt, structured data, llms.txt, MCP and WebMCP, in plain English, and which ones a gym, studio or spa should fix first. Part 2 of Beyond the Four Walls.

Every business owner knows the feeling. The studio looks great, the front desk is staffed, and then someone mentions that the hot water has been off in the back since Tuesday. The part customers see was never the problem. The plumbing was.
Websites now have the same split. The part you paid a designer for is the part people see. Behind it sit a handful of small files and settings that decide whether an AI assistant can get in, understand what you sell, and book something. Most owners have never been told these exist, because until this year they barely mattered.
In Part 1 we made the case that customers are starting to send software to shop for them. This part opens the wall and names the pipes. There are five worth knowing. You do not need to install any of them yourself, but you should know what to ask for, and in what order.
Pipe one: robots.txt, the sign on the door
Every website can have a plain text file at yoursite.com/robots.txt. It is a sign on the door that tells automated visitors which rooms they may enter. On many sites it was written by a web developer years ago and never read since.
What changed is who is knocking. AI companies now send two very different kinds of visitor, and their own documentation keeps them apart.
Training crawlers read the web in bulk to build future AI models. OpenAI's is called GPTBot. Anthropic's is ClaudeBot. Common Crawl, a nonprofit that publishes an open archive of the web, runs CCBot. Google handles training differently: there is no separate crawler, but a setting called Google-Extended tells Google whether pages it has already crawled may be used to train its Gemini models.
Live assistants visit because a person just asked for something. ChatGPT-User, Claude-User and Perplexity-User are the ones a customer sends when they say "find me a facial next Thursday." Alongside them sit the search indexers that decide whether you show up in AI answers at all: OAI-SearchBot for ChatGPT search, Claude-SearchBot for Claude, PerplexityBot for Perplexity.
The companies are clear that these are separate switches. OpenAI's bot documentation says of its search and training crawlers that "each setting is independent of the others." Google says Google-Extended "does not impact a site's inclusion in Google Search nor is it used as a ranking signal." Anthropic spells out what happens if you turn away its live assistant:
Disabling Claude-User on your site prevents our system from retrieving your content in response to a user query.
So blocking training crawlers is a business decision you are entitled to make. Blocking the live assistants is turning away a customer who is standing at the door with their phone out. One honest footnote: OpenAI says robots.txt rules "may not apply" to ChatGPT-User, and Perplexity says its user fetcher "generally ignores robots.txt rules," because a person asked for the page. That does not make the sign pointless. It means the bigger blocker is often not robots.txt at all but a security wall that shows a "checking your browser" page no assistant can pass, the problem we flagged in Part 1.
Pipe two: structured data, your business card in machine print
Your hours are probably on your website. The question is whether they are in a form a machine can trust. A picture of your opening hours, or hours drawn on the page by a script after it loads, is invisible to most software.
Structured data fixes this. It is a short block of code, usually in a format called JSON-LD, that states your facts in a vocabulary maintained by schema.org. The core type is LocalBusiness, which schema.org defines as "a particular physical business or branch of an organization." There are more specific types that fit this industry closely: DaySpa, HealthClub, HairSalon, BeautySalon and NailSalon all sit under HealthAndBeautyBusiness, and gyms fit under SportsActivityLocation. The block carries your name, address, phone, opening hours and price range.
This one is not new and not speculative. Google's own guide says local business markup lets you "tell Google about business hours, different departments within a business," and every example on that page is written in JSON-LD. AI assistants read the same block. If you run several locations, each location page needs its own, with that location's real phone and hours.
Pipe three: llms.txt, the cheat sheet
llms.txt is a short Markdown file at yoursite.com/llms.txt that briefs an AI on what your site is and where the important pages are. It was proposed by Jeremy Howard in September 2024, in plain words: "We propose adding a /llms.txt markdown file to websites to provide LLM-friendly content."
It is a proposal, not a standard, and the evidence on it is mixed. Adoption is growing among large sites. Whether anything reads the file is a different question.
Source: Underneath, Ahrefs, Anthropic
Ahrefs looked at about 38,000 sites that publish the file and found 97% of those files were never requested in a month. We still recommend one, because it takes twenty minutes to write and costs nothing to host. Just do not expect it to move anything on its own. It is the cheat sheet you leave on the counter, not the reason anyone comes in.
Pipes four and five: MCP and WebMCP, from reading to doing
The first three pipes help an assistant find you and understand you. The last two let it do something.
MCP, the Model Context Protocol, was introduced by Anthropic in November 2024 as "a new standard for connecting AI assistants to the systems where data lives." Its own site calls it a USB-C port for AI applications: one standard plug instead of a custom cable for every pairing. In December 2025 Anthropic gave MCP to the Agentic AI Foundation, a fund under the Linux Foundation co-founded by Anthropic, Block and OpenAI. At that point Anthropic counted more than 10,000 active public MCP servers and said MCP had been adopted by ChatGPT, Cursor, Gemini, Microsoft Copilot and Visual Studio Code. The specification is still being revised; the current documentation is versioned July 28, 2026.
For a local business, MCP mostly arrives through your software vendors. It is how an assistant might talk to a booking system, a CRM or a payments platform. You will rarely build one, but you should ask your booking platform whether they have one.
WebMCP brings the same idea to an ordinary web page. Instead of an assistant guessing which button means "book," the page hands it a short menu of named actions, like check availability or book a class, each with the fields it needs. Chrome's documentation describes it as "a proposed web standard to help you build and expose structured tools for AI agents." It is early: Chrome offers it as an origin trial from Chrome 149, and the page warns it "is under active discussion and subject to change." Stripe has already documented how agents can use WebMCP on Stripe's own checkout pages, while labeling it "an experimental browser capability" and telling agents to get the person's confirmation before submitting a payment. More on that in the next part.

Which pipe to fix first
Here is the whole stack in one table, ranked for a local or multi-location business. The order follows what an assistant needs first: get in, understand, act.
| Order | Layer | What it does | Status, October 2026 | Who fixes it |
|---|---|---|---|---|
| 1 | robots.txt and firewall | Lets live assistants in | Long established | Your web developer or host |
| 2 | Structured data (JSON-LD) | States hours, address, phone, prices | Established, used by Google | Your web developer |
| 3 | Plain text facts and real links | Phone and email as tappable links, hours as text | Basic web practice | Whoever edits your site |
| 4 | llms.txt | A briefing file for AI | Proposal; read rarely | You can write it |
| 5 | WebMCP tools | Lets an assistant book on your page | Chrome origin trial | Developer, or your booking vendor |
| 6 | MCP server | Lets an assistant talk to your systems | Widely adopted by AI apps | Usually your software vendor |
Our scoring tool, Agent Ready, weighs the same ground. It puts the most points on whether an assistant can act, because that is the step where a booking happens or does not.
Source: Agent Ready methodology. Reach covers robots.txt and access; Comprehend covers structured data and llms.txt; Act covers WebMCP, forms and booking.
We tested this order on ourselves. Our own new site scored 39 out of 100 on October 9. The first pass of plumbing, the first four rows plus a sitemap and a real form, took about an hour and brought it to 89. Adding real WebMCP tools brought it to 97. The full story is here. The lesson for a gym or spa: the cheap plumbing gets you most of the way, and the last stretch is about letting an assistant actually book.
Next, in Part 3, we follow the money: Stripe is rebuilding checkout for robots, and what that means for a membership sale or a class pack bought by an assistant.
Sources
- OpenAI, "Overview of OpenAI Crawlers," accessed October 2026. developers.openai.com
- Anthropic, "Does Anthropic crawl data from the web, and how can site owners block the crawler?", April 7, 2026. support.claude.com
- Perplexity, "Perplexity Crawlers," accessed October 2026. docs.perplexity.ai
- Google Search Central, "Google's common crawlers" (Google-Extended), accessed October 2026. developers.google.com
- Common Crawl, "CCBot," accessed October 2026. commoncrawl.org
- schema.org, "LocalBusiness" and "HealthAndBeautyBusiness." schema.org/LocalBusiness, schema.org/HealthAndBeautyBusiness
- Google Search Central, "Local Business (LocalBusiness) structured data," updated September 8, 2026. developers.google.com
- Jeremy Howard, "The /llms.txt file," published September 3, 2024. llmstxt.org
- Underneath, "How many websites have an llms.txt file? 2026 adoption data," September 26, 2026. underneath.agency
- Ahrefs, "We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read," June 15, 2026. ahrefs.com
- Anthropic, "Introducing the Model Context Protocol," November 25, 2024. anthropic.com
- Anthropic, "Donating the Model Context Protocol and establishing the Agentic AI Foundation," December 9, 2025. anthropic.com
- Model Context Protocol, "What is the Model Context Protocol?", accessed October 2026. modelcontextprotocol.io
- Chrome for Developers, "WebMCP," published May 18, 2026, updated October 7, 2026. developer.chrome.com
- Stripe, "Use WebMCP to complete Stripe payments in a browser," accessed October 2026. docs.stripe.com
- Sneeze It, "Agent Ready methodology," accessed October 2026. agentready.sneeze.it

