Past a certain length, a chat tool starts restating an answer it already gave you, or quietly drops a correction you made an hour ago. It didn't get dumber. It ran out of room.

If you run long ChatGPT, Claude or Gemini threads for a project, you've seen this. Below: what's actually going on, a two minute test to see it for yourself, and the habit that fixes it.

The chat has a limit and never tells you

Every chat tool works from a context window: the amount of conversation the model can keep in view at once, like a desk that only fits so many pages. Your messages, its replies, any files you pasted. All of it has to fit on that desk.

Once a conversation outgrows the desk, something has to give. Older turns get pushed off or squeezed down, and most interfaces don't tell you when it happens. The chat scrolls on as if nothing changed. You can still read message 4. The model may no longer be working from it the same way.

That's the feeling of the model "getting dumber" partway through a project. The web versions of ChatGPT, Claude and Gemini don't display how full the window is. At best you get a short notice once compression kicks in (Claude shows one when it summarizes older messages). You get no gauge and no warning before quality starts to slip.

Why a correction disappears but a habit doesn't

A model has no memory between replies. Every answer is generated fresh from whatever currently sits in the window. (The "memory" features in these apps store a few selected notes about you. They don't keep your full thread in view.) It doesn't "know" your conversation the way a colleague does. It rereads the pages on the desk each time.

That has a direct consequence. A correction you stated once exists in exactly one place. If that page slides off the desk, or gets buried, the correction is gone. Meanwhile, anything that shows up again and again (the model's own writing habits, a phrasing it keeps reusing, the original version of an answer it has restated three times) keeps reappearing on the desk. Repeated text outvotes the single correction.

There's research behind the "buried" part too. A 2023 paper by Liu et al., "Lost in the Middle", found that language models use information best when it sits at the beginning or end of a long input, and noticeably worse when it sits in the middle. So a correction made midway through a long thread is at a double disadvantage: it's mentioned once, and it's in the part the model uses least well. It doesn't even have to fall out of the window to lose its grip.

Here's what that looks like in practice:

Before and after: a constraint that fades
Message 4 (you):
  Rewrite the intro. Also: no em dashes anywhere in this draft.

Message 5 (AI):
  Here's the revised intro, no em dashes...

[41 messages about headlines, tone, a pricing table, a new section]

Message 47 (AI):
  The new section is ready. Pricing matters — but clarity matters
  more — so we lead with the use case...

Nothing broke. The em dash habit is all over the model's training and its own earlier drafts. Your rule appeared once, 43 messages ago.

This isn't hallucination, and it isn't a bug to wait out

Two failures get lumped together as "AI gets dumber." They're different.

The first makes the second more likely (a model that lost your constraints fills the gap with guesses), but fixing one doesn't mean you've fixed the other.

And the model's ability didn't change. Give the same model a short, fresh conversation and it performs as well as it did at message one. It's working from a smaller, messier view of your project than you think it is.

That's also why waiting for a patch is the wrong plan. A fixed size window has to drop or deprioritize something once a conversation outgrows it. Windows get bigger, but a bigger desk still fills up. The fix is a habit on your side.

Test it yourself in two minutes

Open a chat that has run 30 messages or more, ideally one where you corrected the model at some point. Paste this:

The self test prompt
Without guessing, tell me the single most specific
correction or constraint I gave you earlier in this conversation.
Quote it as closely as you can, and tell me roughly when I said it.

Read the answer against what you actually said. If it names your correction accurately, that's a good sign, but not proof: a direct question makes the model look for the rule, while in normal replies it may still slip. Check the last few answers too. If it names something vague, picks an old instruction you later changed, or gives you the pre-correction version, that turn has lost its weight in the window. The model isn't worse. It's reading from an outdated draft.

Since the web versions of the major chat tools don't show you window usage, this prompt is the most practical check you have.

The fix: summarize it yourself before it falls apart

Teresa Torres makes the same recommendation in her writeup on context rot: stop fighting a long chat message by message. Compress it yourself and start fresh.

  1. In the long chat, ask: "Summarize this conversation in five bullet points: decisions we made, constraints I set, and open questions."
  2. Check the summary line by line. Add any correction it missed. This step matters most, because the summary is only as good as the model's current view.
  3. Open a new chat and paste the corrected summary as your first message.
  4. Continue the work from there.

This treats the window limit as a known constraint you plan around, the same way you'd save a document before it gets too big to open.

A long chat isn't losing its mind, it's losing its view of the earlier turns, so the fix is a clean summary in a fresh chat rather than one more correction. Run the self test on your longest active thread today, and if it misses, do the four step restart before your next request.