When Context Changes: Understanding Update Failures in LLMs
Quick summary
arXiv:2609.38866v1 Announce Type: new Abstract: As preferences, goals, and facts change, LLM agents must use the current state while earlier versions remain in context. Yet they can answer with an old value of the same variable, a failure that we call stale binding. To study when models use outdated information and why, we introduce Controlled In-Context Memory (CICM), a benchmark for tracking and using updated information in conversations and agent logs. We observe that even frontier reasoning models can fail to recover the current state. We find that in open-source models probes can still re
Key takeaways
- arXiv:2609.38866v1 Announce Type: new Abstract: As preferences, goals, and facts change, LLM agents must use the current state while earlier versions remain in context.
- Yet they can answer with an old value of the same variable, a failure that we call stale binding.
- To study when models use outdated information and why, we introduce Controlled In-Context Memory (CICM), a benchmark for tracking and using updated information in conversations and agent logs.
Why it matters
This development shows AI moving deeper into everyday software. Productivity potential should be weighed against price, data permissions, exportability and the preservation of human control.

Member comments