Overview

OpenAI announced that as of February 5, 2025, ChatGPT Search is available to everyone in regions where ChatGPT is available (source).

This is visible every time you see ChatGPT cites a web source link in its response.

While OpenAI did not explicitly state their intention of building their own web indexing system, there are evidences suggesting that is the case like:

  • Advising people not to block their OAI-SearchBot if they want their information to be cited in the ChatGPT Atlas (their browser) – here
  • Stating that using Offline mode will only allow users to access ChatGPT’s indexed and cached web links – here.
  • Even sharing more technical details for their 3 types of crawlers for developers here.

For this reason, to have your web page more accessible for ChatGPT specifically, it is worth looking into what types of crawlers they are using and how they work.


How to Optimize for ChatGPT Crawlers

There are 4 types of crawlers that OpenAI deploys for ChatGPT:

  • OAI-SearchBot – crawls web for search results
  • GPTBot – crawls web for training generator AI foundation models
  • ChatGPT-User – for certain actions by user in ChatGPT, CustomGPT, or external apps built via GPT Actions
  • OAI-AdsBot – specific to landing pages for ads in ChatGPT – makes sure they are compliant

The short answer is – you do not need to do anything to be crawled by these bots.

These bots use Robots.txt file – which acts more in an opt-out fashion.

Meaning, unless you instruct bots in the Robots.txt file not to crawl your website, they will do so.

Remember that the point of OpenAI’s crawlers or any other crawlers meant to best index and retrieve the world wide web should be able to crawl most forms of web content unless otherwise inaccessible for other reasons like password protected, denied access by robots.txt, or restricted by privacy policy in some way.

Further note on this – is that an ambitious web knowledge system that seeks to help users find the most relevant and accurate answer as best as possible would not have the motif to share the details on “how to rank better” on their platform. That can activate bad actors with technical knowledge to overpower others that may be the true authority on a subject matter (like universities, hospitals, etc.) just because they have the ability to “program better”.


LLMs.txt – How Effective Is It?

llms.txt Overview

The llms.txt file was proposed by Jeremy Howard, a prominent figure in AI research and co-founder of fast.ai and Answer.AI.

He is widely known for his contributions to practical deep learning education, open-source AI tools, and applied machine learning. He is also the author of the proposed llms.txt specification.

In September 2024, Jeremy proposed the llms.txt to help website owners have their web pages be better understood by LLMs.

The point of llms.txt is simple – to make web pages and content more digestible for large language models to comprehend.

The web, including software and apps, are comprised of all kinds of different languages. HTML, Javascript, PHP, Typescript, Python, R, and so on…

How would a large language model be able to understand all that?

This was a very logical question that surfaced, in which Jeremy Howard took initiative on.

His point was to make the web content more readable and parseable – meaning, to write some kind of “how to navigate our content” guide LLMs visiting their web content.

Let’s say that ChatGPT primarily speaks English. It is its native language. But ChatGPT is learning Spanish, French, Chinese, Korean, and Arabic. But it’s not quite fluent in all other languages as there are many.

So Jeremy Howards’ idea was to put it all into a “universal language” for LLMs – the Markdown language.

Truth About llms.txt

When you think of the nature of llm.txt – it is meant to help a digital system that is not as capable of reading other formats of web programming language understand the content better.

However, when it comes to search engines, it is rather counterintuitive to how they are meant to crawl the web.

Google, Bing, and even ChatGPT’s goals in indexing the vast web knowledge is to be able to help users find the most relevant and best answer to each user’s search query.

Meaning, that the user may be looking for a specific software, a local restaurant, a product, a book, research paper, or even an answer to a long-tail question for which the answer exists by a credible authority on the web.

This means that web content could surface in any or in combination of programming languages.

It is in search engines’ best interest to crawl the web and index to fit into the how information is naturally surfaced on the world wide web, instead of relying on their technical knowledge to bring information to them.

Search engines are meant to find images, videos, code, niche-content, and even long-tail content like software help files to best serve their core mission.

With the integration of AI, this capability is likely to improve, not worsen.

For this reason, Narr Theory has checked for each of our target AI LLMs to see what they publicly stated as the first-party source of authority on the matter.

  • Google claims that its AI does not need llms.txt file to discover content here.
  • OpenAI has not publicly made a statement about llms.txt files, but they have shown evidence of building their own web index as indicated in the above section of this post.
  • Claude does not explicitly state their usage in llms.txt file either in their document here.

In other words, while the fundamental intent of llms.txt is useful theoretically, for web indexing systems that aim to crawl and index as widely as possible in the worldwide web, it logically would be counterintuitive to assume that llms.txt would be an essential component to having AI systems like ChatGPT or Google AI discover their content.


Advantages of llms.txt

There are, however, some advantages of llms.txt.

The most prominent use case seems to be fore AI agents – for instance Claude has their own llms.txt page found here.

This makes sense, as Claude is often used for programming and their position seems to be very prominent in this sector, likely interacting a lot with agentic AI.

ChatGPT also has adopted llms.txt for their agentic AI purposes – their llms.txt can be found here and here.

Google has published some content on llms.txt – this is in relation to their Lighthouse product, which is an open-source website auditing tool developed by Google. It helps developers evaluate the quality of a website—not by ranking it, but by testing it against a set of best practices.

Lighthouse is like a website inspector that runs automated audits and generates a report for things like:

  • Performance
    • Page load speed
    • Largest Contentful Paint (LCP)
    • Cumulative Layout Shift (CLS)
    • JavaScript execution
  • Accessibility
    • Color contrast
    • Alt text
    • ARIA labels
    • Keyboard navigation
  • SEO
    • Meta descriptions
    • Crawlability
    • Mobile friendliness
    • Structured metadata
  • Best Practices
    • HTTPS
    • Security issues
    • Deprecated APIs
    • Image optimization

So llms.txt can absolutely be helpful – but again, this is meant for agentic AI’s navigation for web pages, so it still pertains to that utility.


Conclusion to llms.txt for AIPR

Specifically for AI Product Recommendation (AIPR), as of July 2026, it can be said that while it is not necessary for web search purposes, it could be helpful for guiding agentic AI to interact with the client’s web content.

Whether or not that should occupy a significant investment in overall AIPR Strategy is dependent on how the everyday consumer adopts agentic AI in their day-to-day adoption of asking questions.

For now, Narr Theory’s position on llms.txt is that while they can be helpful, it is not as important as understanding the overall AIPR strategy, Corroboration in AIPR, or tactics specific to ChatGPT Product Recommendation or Google AI Product Recommendation.


Cloud Browser – Allowlisting

Further in context of AIPR, OpenAI’s posted a formal introduction of their shopping research, which includes allowlisting.

Allowlisting is used by ChatGPT’s Cloud Browser, which is similar to ChatGPT’s agentic AI. It is also used by ChatGPT Work.

Cloud Browser is meant to perform tasks on a website on behalf of a human. It can do things like:

  • Check items in stock
  • Click different store locations
  • Make a reservation
  • Contact businesses for quotes
  • Track a package

It is in ChatGPT’s motif to take more control of the shopping experience in the long-term. In fact, it is stated directly by OpenAI in this post.

OpenAI’s instructions for allowlisting is posted here.

Leave a Reply

Discover more from Narr Theory

Subscribe now to keep reading and get access to the full archive.

Continue reading