Overview

In ChatGPT AIPR, we need to know the baseline model’s capabilities and tendencies in relation to AIPR.

For that, we look at the following

  • Model Version (e.g X.0)
  • Token Nature
    • Transformer with Attention
    • Autoregressive
  • Inference-time Reasoning/Thinking
    • “Thinking” before responding
  • Sampling (ChatGPT) / Temperature (Gemini)
  • Internal Parametric Knowledge Cutoff
    • Case by case
  • Web Search Sequence
    • Mandatory vs. Selective Retrieval
  • SoA

ChatGPT Free Model Version – August 2026

On August 6th, 2026, OpenAI announced that Free and Go users will be using a new default model for everyday chats.

Additionally, OpenAI also announced on August 14th, that GPT 5.6 Luna is becoming the default model for Free and Go users.

Observationally, when using a controlled Free model, we observed this to be true.


Token Nature of GPT5.6 Luna

While OpenAI has not directly stated that GPT5.6 uses the Transformer architecture using Attention, it was published in the following documents by OpenAI researchers that its model use the Transformer architecture:


Reasoning Effort & Tokens

Currently, we do not know with certainty whether or not ChatGPT “thinks” before it starts resopnding with tokens.

One clue to whether or not ChatGPT would think before speaking would be the default reasoning level in the Instant mode – specifically ChatGPT 5.6 Luna at this point.

While the API Developer files say that reasoning is set to Medium by default, we cannot conclude from this that Instant also uses Medium.

We also know from another OpenAI source that GPT-5.6 is capable of operating with Reasoning Effort set to None.

However, that does not necessarily mean that GPT-5.6 Luna on Instant mode is using None.

Even if reasoning was set to None, ChatGPT-5.6 Luna would have to go through its normal computational interpretation of the user’s intent, and decide how to generate the first token

That being said, observationally, it seems that ChatGPT Free version does not have that great of reasoning – it is often seen “thinking out loud”, “talking too much”, and even contradicting its initial reasoning at the beginning of the response compared to the end of the response.

Narr Theory theorizes that ChatGPT’s Free version’s lack of reasoning capability and overtalking can attribute to its churn rate, transitioning users to other LLMs like Gemini.


Sampling Temperature Default by GPT-5.6 Luna

“Sampling” refers to the level of randomness in token generation by the LLM.

For instance, we can say…

“John drank _____ for breakfast”

Here, there are some possible probabilities to completing this sentence:

a) Coffee

b) Water

c) Orange Juice

d) Apple juice

e) Tea

f) Nothing

Having sampling temperature turned to 0 will prompt the LLM to answer with the highest probabilistic answer without much creativity or “randomness”.

Temperature closer to 2.0 will generate a more random response.

This is aligned with OpenAI’s own publication of sampling temperature.

Considering there is nothing specific published by OpenAI about GPT-5.6 Luna’s default sampling temperature, it may be reasonable to assume that 1.0 is set by default as it is the middle ground coherent with Gemini’s temperature at default.


Internal Knowledge Cutoff

GPT-5.6 Luna’s internal knowledge cutoff is February 2026 as per their developer publication.

The same goes for 5.6 Sol and Terra.

This internal knowledge cutoff point will become more relevant case by case when exploring ChatGPT’s internal knowledge regarding specific product categories.


Web Search Sequence

OpenAI has broadly published that ChatGPT will “automatically search the web if your question might benefit from information from the web.”

Web search for ChatGPT is a “tool” to be selected and used by the GPT model.

It is not publicly available information of exactly when and how this tool is decided to be used, nor does it make sense for ChatGPT to be rigid in its decision to use web search.

There can be many different reasons why web search may be beneficial to the user’s question, including:

  • Current information
  • Price
  • Availability / inventory
  • Current product offerings
  • Reviews / ratings
  • Merchant comparison
  • Niche / hard-to-find information
  • Local relevance
  • Market / competitor context
  • Verification / corroboration

Particularly for AIPR, we can broadly assume that the average LLM user has web search enabled and would prefer to benefit from the retrieved information.


Source of Authority Used

The Source of Authority (SoA) cited by ChatGPT in AIPR actually hints to GPT-5.6 Luna’s reasoning and web search sequence.

As seen in our Car Recommendation and Laptop Recommendation case studies, the sequence we observed from internal knowledge model to SoAs cited from online search seemed already heavily biased towards certain brands, citing sources like Toyota.com and Apple.com, in response to a “laptop recommendation”

In these case studies, ChatGPT had already “known” what determines a “easy-to-use” laptop and “reliable” car, largely based on the parametric knowledge from which the training text seemed to have Toyota and Lexus represented at a much larger scale than other car brands like Honda, Ford, Hyundai, Kia, Mazda, BMW, Tesla, etc.

In Phase 1, we will conduct more case studies to further examine what ChatGPT Free models’ cited SoA’s reveal about GPT-5.6 Luna’s product recommendation tendencies.

Leave a Reply

Discover more from Narr Theory

Subscribe now to keep reading and get access to the full archive.

Continue reading