Computer Science editorial
Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding
The core problem
Agentic coding workflows increasingly rely on repository-level instruction files, such as `CLAUDE.md`, that are read by an agent before it acts on a codebase. In real repositories these files grow without bound, stopping only when the repository is retired or when someone rewrites the file wholesale. The paper traces this behavior to imperfect recall: appending an instruction is always cheap, but once an instruction's rationale is gone, deleting it without risking a correctness regression becomes prohibitively expensive.
The authors formalize this cost. For a prompt containing instructions, deleting an instruction whose rationale is no longer recoverable costs in the worst case, because the agent must consider the combinatorial space of behaviors that may depend on it. This asymmetry between cheap appends and expensive deletions produces a divergence the authors call **catastrophic remembering**, the inverse of catastrophic forgetting around which continual learning is organized. The central question is framed bluntly: if English is the new code, why don't we have comments yet?
Innovation
The observational results confirm unbounded growth. Across 247,694 instruction lifetimes in 1,867 repositories, agentic prompts more than triple over their lifetime, gaining **+226%** in size. On average, each commit adds **+4.9 net instructions**. Deletion becomes less likely as instructions age: the log-hazard of deletion is **-0.032 per commit**, meaning older instructions are progressively more entrenched.
The controlled IFEval inversion shows that comments can halt growth. In verifiable worlds where optimal prompts are known, comments encoding latent reasoning remove **99.3%** of excess instructions, reducing excess growth from **+211.3% to +1.4%**.
The WildIFEval transfer shows real-world benefit. Applying the same inversion to WildIFEval, prompt comments improve real-world agentic instruction-following by up to **23.1%**. Together, these results establish both the mechanism of catastrophic remembering and a practical mitigation.
Why it matters
The findings reframe prompt maintenance as a software engineering problem rather than a prompting trick. The core asymmetry is economic: appending is while deleting without rationale is , so rational maintainers append and never prune. Over time, the prompt becomes a sediment of historical decisions, and its growth is bounded only by repository retirement or wholesale rewrite.
The proposed remedy is conceptually simple but structurally significant: treat English instructions as code and give them comments. A prompt comment that encodes latent reasoning preserves the information needed to evaluate deletion, collapsing the effective deletion cost. The IFEval result, reducing excess instructions by 99.3%, demonstrates that this is not a marginal effect but a near-complete correction in controlled settings. The WildIFEval result, up to 23.1% improvement in instruction-following, shows the benefit survives contact with real-world tasks.
The taxonomy candidates for this work span **Architecture**, **Cybersecurity**, **Network**, and **Cryptography**, because agentic prompt files increasingly encode operational, security-relevant, and protocol-level instructions whose silent accumulation has correctness and safety implications. The paper's closing question is therefore both technical and cultural: if English is the new code, why don't we have comments yet?
Who should read this
Opening member contentโฆ