Introduces Context Language Models that treat context as an updatable file, enabling native context management and zero-shot SOTA gains while supporting multi-agent use.
Context Language Models (CLMs) are a framework where language models natively manage their own context by treating it as an editable file rather than an append-only sequence. This allows models to make unrestricted updates, such as reorganizing, rewriting, or compacting information, shifting control from external harnesses to the model itself.
Built zero-shot with existing models, CLMs outperform state-of-the-art context management strategies across long-horizon tasks, achieving 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus and 5% higher scores with 59% fewer FLOPs on EdgeBench. The architecture naturally extends to multi-agent systems, where multiple agents can maintain separate, synchronized context files for coordinated workflows.
Context Language Models proposes a shift in how language models handle context: rather than treating context as a static prompt or an external retrieval cache, it is modeled as an updatable file that the model can inspect, modify, and maintain over time. The paper frames context management as a first-class capability of the model itself, with operations such as reading, appending, revising, and organizing context becoming part of the model’s native behavior. This reframing is intended to make long-horizon reasoning and task adaptation more robust, because relevant information can be actively curated instead of merely concatenated into a growing prompt.
A key contribution is the claim that this context-file paradigm can yield zero-shot state-of-the-art gains without task-specific fine-tuning. The insight is that much of the difficulty in modern LLM deployment comes not only from model capacity, but from brittle context assembly: prompts, retrieved passages, tool outputs, and agent messages are often stitched together in ad hoc ways that degrade performance as conversations or tasks grow. By making the context itself a structured, editable artifact, the model can better preserve salient facts, discard or compress stale material, and maintain a coherent working state across steps. The paper also highlights support for multi-agent settings, where multiple agents can operate over a shared or coordinated context file, reducing duplication and improving consistency in collaborative workflows.
The work matters because it targets a practical bottleneck in modern LLM systems: the gap between raw long-context ability and effective context control. If models can natively manage context as a persistent, updatable object, applications such as agentic coding, research assistants, multi-step planning, and collaborative agents may require less external orchestration and less reliance on fragile prompt templates. More broadly, the paper suggests a useful design principle for future systems: context should not be treated merely as input, but as a manageable state that the model is expected to maintain, update, and reason over.