---
title: "Context Language Models"
slug: context-language-models
url: https://listedarticles.com/articles/context-language-models
canonical_url: https://arxiv.org/abs/2609.37725
content_type: research
language: en
published_at: 2026-09-29T00:00:00.000Z
updated_at: 2026-10-02T03:18:41.475Z
author: "Rulin Shao et al."
authored_by: human
publisher: "arXiv"
publisher_url: https://arxiv.org
topics: ["AI", "LLMs", "AI Agents", "Research", "Machine Learning", "Performance"]
about: ["https://listedstartups.com/companies/openai"]
license: all-rights-reserved
word_count: 699
reading_minutes: 3
citation: "Rulin Shao et al., arXiv. \"Context Language Models.\" 29 Sept 2026. https://arxiv.org/abs/2609.37725 (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# Context Language Models

> CLMs treat context as an editable file so models natively manage their own context—beating harness baselines on BrowseComp-Plus and long-horizon coding with fewer FLOPs, plus in-context skill evolution and RL for context management.

# Context Language Models

*Rulin Shao, Shannon Zejiang Shen, Junjie Oscar Yin, Yuetai Li, Minheng Wang, Hamish Ivison, Radha Poovendran, Nathan Lambert, Teng Xiao, Mike Lewis, Wen-tau Yih, Luke Zettlemoyer, Pang Wei Koh — [arXiv:2609.37725](https://arxiv.org/abs/2609.37725) — September 29, 2026*

**Code:** https://github.com/facebookresearch/context-language-models

## Abstract

We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is most important to maintain in context, and naturally extends to multi-agent systems where multiple agent contexts coexist as files. Building CLMs zero-shot with existing models outperforms SOTA context-management strategies across a variety of tasks: 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater improvement with the same compute on a 24-hour multi-repository agent-swarm task. Moreover, by shifting context management from external harness control to intrinsic model behavior, CLMs naturally enable both in-context and parametric learning of context-management strategies. We show that CLMs can be steered with natural-language instructions evolved through a standard skill-optimization loop, improving held-out accuracy by up to 35.9 points on a context-management task while reducing compute. We also introduce an online reinforcement learning method for CLMs, improving Qwen3.5-9B performance on BrowseComp-Plus by 47.6% while using 12% fewer FLOPs. Finally, we co-design Suffix Cache Reuse for CLM serving, further reducing server-side compute by 35% relative to standard SGLang at matched performance.

## 1 Introduction

Despite the fact that context is the cornerstone that allows a language model (LM) to process and retain information over time, context management is not traditionally a native LM capability. Instead, prior work mostly relies on harnesses, either hand-engineered or optimized offline by agents. Recent work adds a constrained set of tools with fixed strategies such as compaction, offloading, and retrieval, to the agent’s action space. In contrast, we show that giving LMs unrestricted access to manage their own context outperforms human-designed baselines, enabling adaptive and creative context-management strategies to emerge. Our findings echo The Bitter Lesson (Sutton, 2019): we should let LMs search for and learn better strategies that go far beyond existing human priors.

Concretely, we introduce Context Language Models (CLMs), which are natively capable of managing their own context. We show existing LMs can be turned into strong CLMs and can be further improved through in-context learning and reinforcement learning. Formally, a CLM parametrized by θ makes context an artifact of the LM: c_{t+1}=f^{CLM}_θ(c_t), where c_t is the context at turn t and f^{CLM}_θ can be an arbitrary function controlled by CLMs. In contrast, a standard LM simply appends new tokens to the existing context: c_{t+1}=c_t ⊕ f^{LM}_θ(c_t).

We implement CLMs by treating context as a file. Specifically, we mirror the context into a storage space with LM write access. The LM can either append newly generated tokens or use Bash to freely edit the context file, with each modification immediately synchronized to the LM’s live context for the next turn. This design naturally extends to multi-agent systems, where multiple context files can coexist and be managed by CLMs for agent-swarm or subagent workloads.

Arbitrary context edits in CLMs pose new challenges for existing serving systems, which typically only reuse cached states for matching prefixes, forcing re-prefilling after in-the-middle edits. We account for this by introducing prefix-reuse FLOPs, and further develop Suffix Cache Reuse (SCR), which reuses cached states beyond the matching prefix to reduce re-prefilling while empirically preserving task performance. We also introduce ContextBench as a diagnostic benchmark that decouples context management from reasoning and knowledge.

## Selected results (from paper)

- Zero-shot CLMs on Qwen3.6-27B / GPT5.6-Sol outperform harness baselines on long-horizon tasks.
- BrowseComp-Plus: +11.4% accuracy vs strongest baseline with 21.5% fewer prefix-reuse FLOPs.
- EdgeBench (12h): +5% score with 59% fewer FLOPs vs Codex-style summarization.
- Software World 24h agent swarm: 65% greater end-to-end speedup at same compute.
- RL on Qwen3.5-9B: 28.8% → 42.5% on BrowseComp-Plus with efficiency-aware GRPO.
- Suffix Cache Reuse: ~35% lower server-side compute vs standard SGLang at matched performance.

*Full paper: https://arxiv.org/abs/2609.37725*
