9 October 2026International edition
Vol. I · No.
9 October 2026
AI in Fashion
DAILY
The daily briefing on AI in the fashion business
Where fashion meets artificial intelligence.
Commerce & Marketing · How-to

Managing AI crawlers on a fashion site: robots.txt, llms.txt and licensing

A practical guide to deciding which AI crawlers may fetch a fashion site, what robots.txt can and cannot do, and where llms.txt fits.

a close up of a network with wires connected to it
Photo: Albert Stoynov / Unsplash

KEY TAKEAWAYS Summary by the editors

  1. OpenAI runs separate crawlers for search (OAI-SearchBot), model training (GPTBot) and user-triggered fetches (ChatGPT-User), and each can be managed independently.
  2. Google-Extended is a robots.txt token that controls whether crawled content may be used for Gemini training and grounding, and it does not affect inclusion in Google Search.
  3. Google says it needs no llms.txt or other special AI files to include a page in AI Overviews or AI Mode.
  4. llms.txt is a proposal from September 2024 for a Markdown file that gives AI agents a curated site overview, and it is not a formally ratified standard.
  5. Web Bot Auth, based on signed HTTP requests, gives site owners a way to verify that a bot is who it claims to be, which robots.txt alone cannot do.

Which AI crawlers should a fashion site know about?

A fashion site is visited by several kinds of automated agent, and the right policy depends on the purpose of each. Search crawlers fetch pages so that a product or article can be cited in an AI answer. Training crawlers collect content that may be used to build models. User-triggered fetchers load a page because a person asked an assistant about it. Treating these as one thing leads to either over-blocking, which removes the brand from AI answers, or under-managing, which gives away content on terms the business never chose.

Crawler types named in OpenAI and Google documentation
NameOperatorPurpose as documentedControlled by
OAI-SearchBotOpenAISurfaces sites in ChatGPT search featuresrobots.txt, with about 24 hours to take effect
GPTBotOpenAICollects content that may be used to train modelsrobots.txt
ChatGPT-UserOpenAIFetches pages when a user asks a questionOpenAI says robots.txt rules may not apply
GooglebotGoogleCrawls for Search, including AI featuresrobots.txt, plus snippet controls
Google-ExtendedGoogleToken for Gemini training and grounding userobots.txt token, no effect on Search inclusion

Fashion sites have particular assets at stake. Editorial content, lookbooks and campaign photography carry creative value, product pages carry commercial value and customer reviews may carry both. Different parts of a site can justify different rules, since robots.txt works by path. A retailer might allow discovery crawlers on product and category pages while taking a stricter line on image directories or archived editorial, if the business decides that is wise.

How does robots.txt work for AI crawlers?

Robots.txt is a voluntary convention: a file at the root of a domain that tells compliant crawlers which paths they may fetch. OpenAI says that disallowing GPTBot signals that content should not be used for training, that sites opting out of OAI-SearchBot will not appear in ChatGPT search answers although they may still show as navigational links, and that each setting is independent.

There are limits. Fetches triggered by a user, such as ChatGPT-User, may not follow robots.txt, because a person initiated the visit. Robots.txt also cannot stop a non-compliant scraper, and it cannot prove who is making a request.

round white LED light
Read also
Getting found in ChatGPT shopping and AI assistants: what fashion brands can control

What does Google-Extended do and not do?

Google describes Google-Extended as a standalone product token that publishers can use to control whether content Google crawls may be used to train future Gemini models and for grounding, where Search index content is supplied to models at prompt time. It states that the token does not affect a site's inclusion in Google Search and is not a ranking signal. It has no separate user agent string, so crawling still happens through existing Google user agents.

A separate page states that Googlebot's robots.txt directives control crawling for Search, including AI features, and that snippet controls such as nosnippet and max-snippet limit what is shown. Blocking Googlebot to keep content out of AI features would therefore also remove it from Search, which is rarely what a retailer wants.

Is llms.txt worth adding?

The llms.txt proposal, published by Jeremy Howard of Answer.AI on 3 September 2024, describes a Markdown file at the site root with a title, a short summary and curated links to more detailed pages. The proposal's own page calls it open for community input and an informal overview, not a ratified standard. It says thousands of sites publish the file, and that OpenAI, Anthropic and Google's Gemini publish one for their developer documentation.

For a fashion retailer the evidence of benefit is thin. Google says that to appear in AI Overviews or AI Mode, a site does not need new machine-readable files, AI text files or special markup. OpenAI's published crawler guidance does not mention the file either. A reasonable view is that llms.txt is cheap to produce and harmless, but it should not be prioritised ahead of crawl access, clean product data and fast, indexable pages.

How can a site verify that a bot is genuine?

Because user agent strings can be copied, verification matters more as agents begin to place orders. Cloudflare documents Web Bot Auth, a method in which a bot signs its HTTP requests with a private key and publishes the public key in a directory on its own domain. The site, or its CDN, checks the signature against the registered key. Cloudflare recommends short expiry values to limit replay.

This sits with security and platform teams. For a retailer, the benefit is the ability to allow verified agents through while treating unverified scripts with suspicion, which is hard to do with robots.txt and IP lists alone.

Logging is the practical starting point. Server logs show which user agents request which paths, how often, and from where. Comparing those with the published crawler names shows quickly whether the documented crawlers are reaching the site and whether unknown agents are scraping product pages. That evidence also informs the policy discussion, because decisions are easier when they rest on observed traffic and not on assumptions.

woman in black and white polka dot dress sitting on brown wooden chair
Read also
How to audit your fashion brand's visibility in ChatGPT, Gemini and Perplexity

What steps should a fashion site take, in order?

  1. List the AI crawlers you want to allow for discovery, such as OAI-SearchBot and Googlebot, and write that policy down with the legal and brand teams.
  2. Decide separately about training access, using GPTBot and Google-Extended rules.
  3. Check CDN and firewall rules so allowed crawlers are not blocked.
  4. Review access logs monthly for unknown agents and consider signature verification for those you rely on.
  5. Add llms.txt only after the basics are in place, and keep it short and accurate.
  6. Where content such as lookbooks or editorial photography has licensing value, involve legal counsel before deciding on blanket access.

Licensing deserves a note. Robots.txt expresses a preference and is not a contract. Brands that want to be paid for training use, or that want to restrict it by agreement, need terms of use and, where relevant, direct negotiation, and the law differs by jurisdiction.

Frequently asked questions

Should I block GPTBot on my fashion website?

That depends on whether you want your content used for model training. OpenAI says disallowing GPTBot signals that content should not be used for training, and that this is independent of OAI-SearchBot, which governs inclusion in ChatGPT search.

Does blocking Google-Extended hurt my Google rankings?

Google says it does not. Google-Extended is a token that controls use of content for Gemini training and grounding, and it does not affect inclusion in Google Search or act as a ranking signal.

What is llms.txt and do I need it?

It is a proposed Markdown file at the site root that summarises a site for AI agents. Google says no special AI files are needed for AI Overviews, so it is optional and secondary to crawl access and product data.

Does robots.txt stop ChatGPT from reading my page?

Not always. OpenAI says that ChatGPT-User fetches are initiated by people, so robots.txt rules may not apply, while GPTBot and OAI-SearchBot are managed through robots.txt.

GuideThe complete guide to AI in fashion e-commerce, marketing and retailRead the complete guide
Get the Daily

One edition every weekday morning. Read in five minutes. Free for industry professionals.

Newsletter

More on Technology

View all