{"slug":"research-research-series","name":"Agent Context Research: The Evidence So Far","rules":{"start":"open"},"nodes":{"open":{"short":"the note","pose":{"base":"centre"},"beat":{"layout":"center","marker":"evidence map · interactive","title":["Agent Context Research:","The Evidence So Far"],"glass":1,"lede":"A guided map of five agent-context studies, including the 84-run replication that did not reproduce the pilot's cheaper-context headline, and the question each follow-up was built to answer.","caption":"8 beats · 4 min to read in full"}},"premise":{"short":"the premise","pose":{"base":"flooded"},"beat":{"layout":"left","marker":"01 · evidence map","title":["This page is a reading map","for the research program"],"lede":"This page is a reading map for the research program, not a new experiment. It shows how the question changed as the evidence accumulated: from whether a richer workspace helps, to whether agents consume it, to which information is actually…","caption":"Condensed from the note."}},"the-evidence-path":{"short":"the evidence path","pose":{"base":"drained"},"beat":{"layout":"right","marker":"02 · the evidence path","title":["The evidence","path"],"lede":"The sequence below follows how the questions developed. Some later studies branch from the same result rather than forming a strict publication chronology.","panelRows":[{"title":"1","note":"Study How the Environment Affects Agent Performance and Token Cost · Question Does progressively richer project context change success or token cost? · What changed In a small pilot, every environment completed the task and richer conditions sometimes used fewer effective tokens. The full scaffold's input was heavily cached. The paper correctly labels this as a pilot."},{"title":"2","note":"Study The Incurious Agent · Question Does the cost result repeat, and what context do agents actually read? · What changed Across 84 runs, the cheaper-context headline did not repeat. Both agents largely skipped the deepest curated material; context that went unread could not improve the outcome. This revised the pilot's main cost claim and shifted attention from authoring to consumption."},{"title":"3","note":"Study The Self-Sufficient Agent · Question If the tasks and repository get harder, does a bare agent finally need the documentation? · What changed On this repository and task family, Codex and Claude solved all nine functional landmines without docs. They failed mainly on arbitrary project-specific formats. The useful surface for context narrowed from general capability to local convention."},{"title":"4a","note":"Study Relevance, Not Volume · Question Is the amount of context the lever, or whether it contains the needed rule? · What changed A generic contract and a same-length rule-bearing contract produced radically different conformance. On these tasks, useful context was information the agent needed and could not derive, not simply more documentation."}],"caption":"Condensed from the note, which carries 5 rows."}},"a-related-evaluation-track":{"short":"a related","pose":{"base":"centre"},"beat":{"layout":"left","marker":"03 · a related","title":["A related","evaluation track"],"lede":"A sixth study, Fable 5 vs Opus 4.8: A Coding-Agent Evaluation, uses the same broader discipline (pinned tasks, objective gates, captured runs, and explicit limitations) but asks a different question: how two models compare on real…","caption":"Condensed from the note."}},"the-synthesis":{"short":"the synthesis","pose":{"base":"flooded"},"beat":{"layout":"right","marker":"04 · the synthesis","title":["The synthesis","evidence map"],"lede":"Across the context studies, four conditions now matter more than document count:","rows":[{"title":"Headroom","note":"the agent must have a gap that information can close. Context cannot improve a task the bare agent already solves at the ceiling."},{"title":"Relevance","note":"the information must supply something needed and underivable, such as a local wire-format convention."},{"title":"Consumption","note":"the agent must actually encounter the information. A perfect memory tree that stays unopened has no treatment effect."},{"title":"Agent behavior","note":"reading habits differ. Results from one agent cannot be assumed to describe every agent."}],"caption":"Condensed from the note."}},"choose-a-reading-path":{"short":"choose a reading","pose":{"base":"drained"},"beat":{"layout":"left","marker":"05 · choose a reading","title":["Choose","a reading path"],"rows":[{"title":"Start with the correction","note":"read the pilot, How the Environment Affects Agent Performance and Token Cost, then The Incurious Agent."},{"title":"Understand what agents already know","note":"continue to The Self-Sufficient Agent."},{"title":"Turn the result into documentation practice","note":"read Relevance, Not Volume."},{"title":"Compare reading behavior across agents","note":"read Curiosity Comparison Between Agents."}],"caption":"Condensed from the note, which carries 5 points."}},"boundaries-that-still-matter":{"short":"boundaries","pose":{"base":"centre"},"beat":{"layout":"right","marker":"06 · boundaries","title":["Boundaries","that still matter"],"lede":"The studies cover a small number of repositories, tasks, agents, and runs. Several results are directional, and the experiments deliberately use different task families to answer different questions.","caption":"Condensed from the note."}},"close":{"short":"the end","pose":{"base":"finale"},"beat":{"layout":"center","title":["That is the note,","in one pass."],"glass":1,"lede":"This is Agent Context Research: The Evidence So Far condensed to its beats. The full note carries every paragraph, the tables, the code, and the source it was adapted from.","links":[{"href":"/research/research-series","label":"Read the full note"},{"href":"/research","label":"All research","tone":"ghost"}]}}}}