Knowledge System Entry Points
When someone asks ChatGPT:
“What are the best basketball shoes for jumping?”
ChatGPT does not simply open a master database of basketball shoes, compare every product, and choose a winner.
Information about brands and products can reach ChatGPT through several different pathways.
At Narr Theory, we can reduce those pathways into three major layers through which a brand can enter ChatGPT’s accessible knowledge environment:
- Parametric Knowledge — information learned through model training.
- Web Search — information retrieved from the web when ChatGPT answers a question.
- Commerce Infrastructure — structured product information supplied through systems such as Shopify Catalog and the Agentic Commerce Protocol.
These layers perform different functions.
A brand appearing in one does not necessarily mean it appears in another. And being present across all three does not guarantee that ChatGPT will recommend the brand.
Understanding these layers helps explain a much larger question:
How does ChatGPT know enough about a brand or product to consider recommending it in the first place?
Layer 1: Parametric Knowledge
The first layer is information learned during model training.
OpenAI explains that its models learn patterns from large amounts of information. Text is processed through smaller units called tokens, and during training the model adjusts its internal parameters based on patterns and relationships found across the information it processes.
Importantly, OpenAI explains that its models do not simply store copies of the documents on which they were trained. Instead, the model’s weights or parameters are adjusted to reflect learned patterns.
A useful abstraction is:
Training Data → Model Parameters → Parametric Knowledge
If information about Nike, Adidas, ASICS, or a much smaller brand appears throughout the information used during training, relationships involving those brands can potentially become represented within the model’s parameters.
Narr Theory refers to this learned information as parametric knowledge.
Where Does ChatGPT’s Training Information Come From?
OpenAI currently describes three major sources of information used to develop its foundation models:
- publicly available information on the internet;
- information accessed through third-party partnerships;
- information provided or generated by users, human trainers, and researchers.
OpenAI has also increasingly discussed the use of synthetic data during parts of model development.
Historically, GPT-3 provides one of the clearest public examples of what a training mixture can look like.
OpenAI’s Language Models are Few-Shot Learners disclosed that GPT-3’s training mixture included filtered Common Crawl, WebText2, Books1, Books2, and Wikipedia.
Common Crawl represented 60% of GPT-3’s weighted training mixture, while Wikipedia represented 3%.
That does not tell us what modern ChatGPT models are trained on. OpenAI does not publicly disclose the exact training mixture of its current frontier models.
But it illustrates something important.
The public web has historically been an important source through which language models learn relationships between entities, concepts, brands, products, attributes, and language.
Wikipedia is therefore significant as an information-dense source, but it should not be treated as the knowledge database behind ChatGPT.
It is one source within a much larger information ecosystem.
GPTBot Gives Brands a Training-Layer Entry Point
OpenAI also operates GPTBot.
Website owners can use robots.txt to control whether GPTBot may access their publicly available content. OpenAI describes GPTBot as a crawler associated with potential model training.
This gives us one possible pathway:
Brand Information → Public Web → Training Data → Model Parameters
But there is an important limitation.
Allowing GPTBot does not mean that a webpage will definitely become training data.
And even if information is included during training, there is no guarantee that the model will learn a particular fact, association, or product claim from it.
There is no equivalent of submitting a URL to an SEO index and knowing that the information has now been added to ChatGPT’s parametric knowledge.
Brands can make information available to the training ecosystem.
They cannot directly tell the model what it must learn.
Layer 2: Web Search
The second layer works very differently.
Instead of relying entirely on information already represented in model parameters, ChatGPT can retrieve information from the web while generating an answer.
OpenAI says ChatGPT may automatically search the web when a question would benefit from current information.
This is particularly important for product recommendation because parametric knowledge is inherently bounded by training.
Products change.
Prices change.
Reviews accumulate.
Specifications are updated.
New brands and products appear.
Web search gives ChatGPT another pathway to knowledge.
Instead of relying only on:
Question → Parametric Knowledge → Answer
ChatGPT can perform something closer to:
Question → Search → Retrieved Evidence → Answer
ChatGPT Can Rewrite the User’s Question Into Searches
This is where AI Product Recommendation starts becoming an information retrieval problem.
OpenAI explains that ChatGPT may rewrite a user’s request into one or more targeted search queries.
It may also perform additional searches after examining initial results.
Imagine someone asks:
“I need basketball shoes that are good for jumping, provide ankle protection, and are worn by professional players.”
ChatGPT does not necessarily submit that exact sentence to a search engine.
The information need could potentially lead to searches relating to:
- basketball shoes for jumping;
- basketball shoes with ankle protection;
- basketball shoes worn by professional players;
- reviews of particular candidate shoes;
- specifications or evidence concerning particular models.
That distinction matters enormously for brands.
A product does not merely need to exist online.
Information about that product needs to be discoverable within the kinds of searches ChatGPT might perform while attempting to answer a user’s decision narrative.
From WebGPT to ChatGPT Search
OpenAI has been researching this approach for years.
In 2021, OpenAI introduced WebGPT, a GPT-3-based system trained to answer questions using a text-based web browser.
WebGPT could search the web, follow links, inspect webpages, and cite sources while constructing an answer.
WebGPT should not be interpreted as evidence that modern ChatGPT uses exactly the same architecture.
Its importance is conceptual.
It demonstrated an early OpenAI approach to combining a language model’s parametric knowledge with external information retrieval.
Rather than forcing the model to answer everything from what it had already learned, the model could look for evidence.
Today, OpenAI exposes this concept much more abstractly through a built-in Web Search tool, which allows models to search the web and incorporate retrieved information into responses.
GPTBot and OAI-SearchBot Are Not the Same Thing
This distinction is especially important for brands.
GPTBot and OAI-SearchBot perform different functions.
GPTBot relates to potential model training.
OAI-SearchBot helps make webpages discoverable within ChatGPT search.
OpenAI recommends allowing OAI-SearchBot if publishers want their webpages to be eligible to appear, be cited, and receive links from ChatGPT search experiences.
OpenAI can also rely on external search infrastructure.
Its documentation explains that ChatGPT may send rewritten searches to third-party search providers and references Microsoft and Shopify among its search providers.
Therefore:
OAI-SearchBot = OpenAI’s search crawler
Microsoft/Bing = external search-provider infrastructure
They should not be treated as the same pathway.
For brands, however, both contribute to a larger objective:
web discoverability.
Layer 3: Commerce Infrastructure
The third layer is specifically designed around products and commerce.
Traditional web search can discover products from ordinary webpages.
But commerce contains structured information that ordinary webpages may represent inconsistently:
- product identifiers;
- variants;
- images;
- prices;
- inventory;
- availability;
- product descriptions;
- merchant information;
- promotions;
- purchasing information.
Commerce infrastructure gives ChatGPT a more structured way of accessing this information.
Shopify Catalog
For Shopify merchants, this layer is particularly important.
OpenAI says Shopify product data is integrated into ChatGPT through Shopify Catalog, helping product information appear accurately and completely in relevant shopping experiences.
Individual Shopify merchants currently do not need to create a separate OpenAI product feed to participate through this integration.
That can provide an important product-discovery and information-completeness advantage.
But it is critical to distinguish two concepts:
Product Discovery
and
Product Recommendation
Being represented within a commerce catalog can help ChatGPT know that a product exists and obtain structured information about it.
It does not prove that the product is the right recommendation.
The Agentic Commerce Protocol
OpenAI is expanding this commerce layer through the Agentic Commerce Protocol, or ACP.
ACP provides infrastructure through which merchants and commerce platforms can supply structured product and commerce information into the ChatGPT ecosystem.
OpenAI describes it as connective infrastructure between merchants and users throughout product discovery.
Merchants outside integrations such as Shopify can also pursue direct product-feed access. They must sign up and get approved as a merchant.
For approved merchants, this can provide an advantage.
Their products can potentially be represented more directly, completely, and accurately than products ChatGPT must discover independently from scattered webpages.
But that advantage should not be confused with preferential recommendation.
OpenAI states that product results are selected independently by ChatGPT and are not influenced by OpenAI partnerships.
The distinction is crucial.
Commerce infrastructure can help a product enter the candidate environment.
It does not give the merchant ownership of the recommendation.
Merchant Data Is Not the Same as Evidence
Imagine a merchant submits the following statement about its basketball shoe:
“The easiest basketball shoe to jump in.”
Or:
“The best basketball shoe for ankle protection.”
The merchant may be an authoritative source for factual information about its own product:
- its dimensions;
- materials;
- price;
- available sizes;
- variants;
- release date;
- product description.
But declaring its own product “the best” is different.
That statement does not independently establish that the product actually outperforms competing shoes.
OpenAI does not publish a simple rule stating that merchant marketing claims are automatically rejected.
What OpenAI does document, however, shows that its shopping systems can consider information beyond the merchant itself.
Shopping Research can use merchant product data, publicly available product information, and other relevant retail sources.
OpenAI also says Shopping Research can read public retail websites in real time, cite sources, and attempt to avoid low-quality or spammy information.
That gives us one of the most important distinctions in AI Product Recommendation:
Commerce data helps tell ChatGPT what the product is.
Evidence helps ChatGPT determine whether the product satisfies the user’s constraints.
Summary: The Three Layers of ChatGPT’s Product Knowledge Environment
We can therefore reduce the system to three major layers.
| Layer | How Information Enters | Primary Function |
|---|---|---|
| Parametric Knowledge | Training data, public internet information, partnerships, GPTBot-accessible content | Creates learned associations within model parameters |
| Web Search | OAI-SearchBot, search providers, publicly accessible webpages | Retrieves external information and evidence at inference time |
| Commerce Infrastructure | Shopify Catalog, ACP, direct merchant feeds | Provides structured and current product information |
These layers should not be thought of as three competing databases.
They are different information pathways available to ChatGPT.
And they can interact.
A product might already exist within the model’s parametric knowledge.
Commerce infrastructure could provide current structured information about the product.
Web search could then retrieve reviews, testing, specifications, professional usage, or other evidence needed to determine whether the product satisfies a user’s particular constraints.
That gives us a more complete model:
Parametric Knowledge → Prior Associations
Commerce Infrastructure → Product Representation
Web Search → Retrieved Evidence
Together, these systems can contribute to the product recommendation process.
What This Means for Brands
For brands, the goal should therefore not simply be:
“Get into ChatGPT.”
There is no single database called ChatGPT into which a company submits its products.
A better objective is to build presence across the different information environments ChatGPT can access.
A brand wants its products to be:
Known
Represented within the broader information environment from which language models can learn.
Discoverable
Accessible when ChatGPT searches the web for information or evidence relevant to a user’s question.
Structured
Represented accurately within commerce systems that help ChatGPT understand products, variants, pricing, availability, and other commercial information.
But even achieving all three does not guarantee recommendation.
Because the final problem is not simply whether ChatGPT can find the product.
The final problem is epistemic:
Does ChatGPT have sufficient reason to believe that this product is the right answer to this particular user’s decision narrative?
A product may be known.
It may be discoverable.
It may exist inside ChatGPT’s commerce infrastructure.
And yet another product may still have stronger evidence supporting the constraints that matter to the user.
That is why AI Product Recommendation cannot ultimately be reduced to feeds, crawlers, catalogs, or integrations.
Those systems help determine whether information about a brand can enter ChatGPT’s accessible knowledge environment.
The recommendation itself depends on whether ChatGPT can find sufficient evidence to conclude that the product is the right answer.


Leave a Reply