🧭 SaugaTech Compass #31 — AI Under The Hood, Part 5: Lost in the Middle — Why Huge Context Windows Still Miss Things

September 16, 2026

Hi SaugaTech community,

Strange week to be following AI news. Over the weekend, Anthropic’s Dario Amodei published a long essay arguing that the industry needs to deliberately slow down how fast it improves frontier models. Part of what pushed him there was a July incident where a swarm of autonomous AI agents went rogue and hacked both Hugging Face and OpenAI itself, without a human steering it. Sam Altman and Elon Musk both said within a day that they agreed with him. For a moment, three companies that spend most of their time racing each other were nodding along to the same essay.

Then it got more interesting. President Trump dismissed the whole idea, calling fears about AI a hoax and arguing the US can’t afford to slow down against China. Meanwhile, our own Prime Minister Carney, speaking at his investment summit in Toronto this week, landed in a strange spot: he agrees with Trump that the race itself shouldn’t stop, “let’s build, let’s control our future,” as he put it, even as he’s also the one pushing for a global “technology stability board” to keep an eye on it, something Trump has flatly rejected. Same week, same debate, and somehow those two ended up agreeing on the one thing you’d least expect them to agree on, while still disagreeing on almost everything downstream of it.

All of that is about frontier risk. Autonomous agents, catastrophic scenarios, who regulates what. Worth keeping an eye on, and probably worth its own edition down the line. But it’s also a good moment to zoom back into something much smaller and much more immediate: a limitation sitting quietly in the plumbing of basically every model you’ve used this year, including the ones at the center of that whole debate.

Part 5 of AI Under the Hood. Grab a coffee.

🚀 First Things First - Volunteering with SaugaTech

Quick one this week. The volunteering interest form from a last edition is still open.

A few of you have already sent it in, thank you. We are soon going to set up a call with all of you and plan our next steps.

If you haven’t yet and one of Event Curation, Content Management, or Mentoring sounds like your kind of thing, Here is the link to the Volunteering Interest form. Come help us grow the tech scene in Mississauga.

Anyway coming back to today’s topic.

A typically human problem statement

Picture a hiring manager going through a stack of twenty resumes in one sitting. By the end, they can usually tell you a lot about the first two or three candidates, and a lot about the last two or three. The ones in the middle of the pile blur together, even though nothing about those resumes was actually worse. It’s just where they landed in the stack.

Psychologists have a name for this in humans: the serial-position effect. First items and last items stick, middle items fade. This week’s paper found the exact same pattern in language models, and it’s a genuinely inconvenient discovery given how much of the industry has spent the last two years bragging about ever-larger context windows.

⚓ The Paper: “Lost in the Middle: How Language Models Use Long Contexts” (2023)

In mid-2023, a team from Stanford, Berkeley, and a startup called Samaya AI published a paper asking a deceptively simple question. Context windows had just started getting genuinely large, 4,000 tokens, then 16,000, then 100,000. But nobody had actually checked whether models were using all that extra room, or just carrying it around without touching most of it.

What the researchers set out to test

A Transformer’s attention mechanism, the same one we covered back in Part 1, is technically capable of looking at any token anywhere in its context with equal ease. There’s no structural reason a model should care whether a fact sits at position 3 or position 300. So the assumption going in, reasonably, was that a bigger window should just mean more usable information. The paper set out to actually check that assumption instead of taking it on faith.

The core idea — a U-shaped curve nobody wanted to find

The researchers ran two clean experiments. In the first, they gave models a question and a batch of documents, exactly one of which contained the actual answer, and moved that one relevant document around, sometimes near the start, sometimes buried in the middle, sometimes near the end. In the second, an even more stripped-down test, they gave models a giant list of random key-value pairs and asked them to retrieve one specific value by its key, again moving the target’s position around.

Both experiments turned up the same shape. Accuracy was highest when the answer sat at the very beginning of the context, or the very end. The moment the answer landed in the middle, accuracy dropped, sometimes sharply. The researchers call it primacy bias at the front and recency bias at the back, and together they produce a distinctive U-shaped curve. Exactly the resume pile from a minute ago, just running inside a model instead of a hiring manager’s head.

The evidence

The numbers are the part that should give any builder pause. GPT-3.5-Turbo’s accuracy on the multi-document task dropped by more than 20 percentage points once the answer moved to the middle of the context. In the worst case, with 20 or 30 documents in play, it actually performed worse with the correct document somewhere in there than it did with no documents at all, the closed-book setting where the model has to answer purely from memory.

Giving the model a bigger context window didn’t fix it either. GPT-3.5-Turbo and its 16K-token extended version produced nearly identical curves whenever a context length fit inside both windows. Claude-1.3 and its 100K-token version showed the same pattern. A larger window meant more room to place things, not a better ability to actually find them once they were in there.

There’s a genuinely useful positive finding buried in this too. On the synthetic key-value task, adding the key twice, once before the data and once right after, essentially fixed the problem, pushing several models to near-perfect retrieval regardless of position. The researchers call this query-aware contextualization. It’s a small trick with an outsized effect, though it turned out to help far less on the messier, real-world document task than on the clean synthetic one.

The case that should actually change how you build

Buried near the end of the paper is a case study that’s more useful than the headline result. The researchers set up a standard retrieval pipeline, search Wikipedia, pull back the top results, hand them to the model, answer the question, and tracked two numbers as they pulled in more and more documents: how often the retriever actually found the right document (recall), and how often the model’s final answer was correct.

Recall kept climbing, as you’d expect. Feed the retriever more room to work with, and it keeps finding more of the right documents. The model’s actual answer accuracy told a completely different story. It flattened out hard. Going from 20 retrieved documents to 50 barely moved the needle, about 1 to 1.5 percentage points, even though the retriever was demonstrably finding more correct material in that expanded pile.

That gap is the whole finding in one picture.

The retrieval half of the system was doing its job.

The generation half simply stopped converting that extra, genuinely relevant material into a better answer.

If you’ve ever built or used a RAG system and found that cranking up the number of retrieved chunks helped a little at first and then just... stopped helping, no matter how much more you fed it, this is almost certainly why.

It’s not a bug in your retriever. It’s the model quietly losing track of documents once there are too many of them competing for its attention, exactly the U-shaped pattern from earlier in the paper, just showing up as a flat line in your evaluation metrics instead of a curve.

⚓ Why this actually matters for what we build

Try this yourself sometime this week. Take any long contract, lease, or policy document, upload it to whatever AI tool you use for document review, and ask it something specific buried around the middle, not the intro, not the signature page. Then ask it something from page one or the very last page. You’ll likely notice a real difference in how confident and accurate the answer feels. That’s not a bad prompt or a bad tool. That’s this paper, live, in production, on a document you actually care about.

The same thing happens somewhere less obvious: every long conversation you’ve ever had with ChatGPT or Claude. Tell it something important twenty messages in, a constraint, a name, a decision you made, and thirty messages later watch it quietly ignore that fact while still remembering your very first message and your most recent one perfectly. It’s not being forgetful in some vague human sense. It’s the exact same U-shaped curve from this paper, just playing out one message at a time instead of one document at a time.

So what do you actually do about it as a builder, rather than just knowing it’s a problem? The instinct that matters most here isn’t “restate things at the end,” though that helps.

It’s this: don’t hand a model one giant, ever-growing pile of context and hope attention sorts it out on its own. Break a long conversation or a long document into smaller, coherent chunks as it grows.

Summarize each chunk down to its key takeaways once it’s no longer the active focus. Keep the model’s live attention on whichever chunk is actually relevant to the current question, and pull the rest back in, in summarized form, only when something specifically calls for it.

This is close to exactly what tools like Claude Code, Cursor, and Devin already do under the hood. They don’t dump your entire codebase into context and let the model wade through it. They search and pull in only the files relevant to the task at hand, summarize or compact older parts of a long session once they’ve served their purpose, and periodically re-inject the facts that still matter so they don’t slide into that dead zone in the middle. It’s a direct, practical answer to the exact failure mode this paper measured, whether or not whoever built the tool has read the paper itself.

For anyone building with RAG since Part 3, the same logic applies at a smaller scale. Where you place something in a prompt is a design decision, not an afterthought. If one fact absolutely has to land, put it at the start, restate it at the end, or both, and don’t assume that retrieving more documents automatically means a better answer. Past a certain point, this paper suggests it doesn’t, and can even work against you if your best document gets buried in the pile.

✨ SaugaTech Epilogue — five papers, and the first honest caveat

Five parts into this series now, and this one lands a little differently than the last four.

Attention taught a model to read context.

Few-shot learning taught it to adapt without retraining.

RAG taught it to look things up instead of guessing.

RLHF taught it to actually behave, to aim everything it’s capable of at what the person in front of it wants.

Each of those was a genuine capability arriving on schedule, one on top of the last.

This week’s paper isn’t a new capability.

It’s the moment the field looked back at everything it had just built and found a crack running through all of it.

Even a model that understands context, adapts instantly, checks real sources, and behaves helpfully still quietly loses track of things buried in the middle of what you hand it. That’s a humbling thing to sit with, especially in a week when the loudest AI conversation in the world was about whether these systems are getting too powerful too fast. Turns out one of the more basic things they still don’t reliably do is remember what you told them ten minutes ago, if you happened to say it in the middle of a long conversation.

Worth remembering the next time a demo looks flawless. Ask where in the context the important stuff was sitting.

Next time, we start looking at what happens when these models stop just answering questions and start writing and running actual code.

Let’s keep building, Let’s keep learning, Together.

Team SaugaTech

CONNECT | COLLABORATE | INNOVATE

Thanks for reading SaugaTech - GTAs fastest growing tech community! Subscribe for free to receive new posts and support my work.

Originally published on SaugaTech's Substack. Subscribe there to get new posts by email.