Meta, MIT, and UW create models that edit their own context to cut compute costs
A study by researchers from Meta, MIT, and the University of Washington, published in October 2026, proposes Context Language Models (CLMs): models that rewrite their own conversation history as if it were a file, without fixed rules for summarizing or discarding.
The problem CLMs try to solve
Every AI agent that runs for a long time hits the same limit: the context window fills up. The solutions so far are external to the model and come with known trade-offs. Summarizing can discard important details or introduce errors, compacting follows fixed rules defined by whoever built the agent, and searching external memory requires the agent itself to decide what to bring back.
A group of researchers from Meta, MIT, and the University of Washington proposed, in October 2026, an alternative described in the article Context Language Models: Self-Managing Context to Improve Performance and Reduce Compute Costs: instead of external rules, the model itself starts controlling what stays and what leaves its context.
Context as a file the model edits
In practice, the CLM treats the conversation history as a text file it can rewrite freely, without predefined constraints. This means the model can rewrite old messages, preserve relevant facts, delete irrelevant information, and keep progress notes about an ongoing task.
According to the researchers, this autonomy is the central point of the proposal: by taking context management control away from the external "harness" (the code that orchestrates the agent) and making it an intrinsic behavior of the model, space opens up for it to learn, in context or through training, its own management strategies, which may outperform "existing human priors" (original, en: "existing human priors").
An interesting side effect: since context management becomes a behavior of the model, the user can simply tell the agent, in natural language, how they prefer the context to be handled. The CLM can also evolve a skills document in context, storing useful management procedures to reuse in future tasks.
Three paths to learning context management
The researchers tested three different approaches to teach the CLM to manage its own context:
- Zero-shot: the model is given the ability to edit context without any specific training for it.
- In-context learning: the model is instructed in natural language on how to manage context and refines the strategy in an iterative skill-optimization loop.
- Reinforcement learning: the model learns context management strategies using task success as the main objective and computational efficiency as an additional criterion.
New observed behaviors include creating internal notes, removing irrelevant intermediate results while preserving useful ones, and tracking failed experiments alongside ideas still to be explored.
The numbers reported by the researchers
The gains reported in the article are substantial in both performance and computational cost. In the zero-shot configuration, CLMs achieved 11.4% higher accuracy with 21.5% fewer FLOPs on the BrowseComp-Plus benchmark, 5% higher scores with 59% fewer FLOPs on the 12-hour EdgeBench, and 65% greater improvement with the same compute on a 24-hour multi-repository agent swarm task.
With in-context learning, CLMs improved accuracy on ContextBench tasks by up to 35.9 percentage points, with lower computational cost. With reinforcement learning, Qwen3.5-9B's performance on BrowseComp-Plus rose from 28.8% to 42.5%, a relative improvement of 47.6%, using 12% fewer FLOPs.
What remains open
The authors themselves acknowledge important limitations. A CLM can discard relevant information it later cannot recover, and giving the model full control over its context creates a new risk surface, since editable context "can become another channel through which prompt injections or self-generated instructions persist across turns" (original, en: "can become another channel through which prompt injections or self-generated instructions persist across turns") (autor: the study's researchers). More control over context, according to them, does not by itself guarantee better behavior, since the model can still make poor decisions about what to keep, modify, or discard.
Reception outside the paper also brought caveats. On X, user @Rennix7t noted:
the accuracy and computational power results in the paper come from specified tasks
@Rennix7t, on X
@omarsar0, meanwhile, described the approach as an interesting research direction, but said he had reservations about trusting a model to manage its own context end-to-end, arguing that better solutions are still needed. On Reddit, user Combinatorilliance argued that CLMs are not yet ready for production, highlighting a concrete technical effect:
[the] caching trick [used by the researchers] isn't ideal either and has some drawbacks that need to be taken into consideration
Combinatorilliance, on Reddit
Combinatorilliance's technical point is relevant for anyone operating inference infrastructure: rewriting the context invalidates the model's cache (the mechanism that reuses computation from prefixes already processed across calls). This means part of the FLOP savings reported in the article can, in practice, be partially offset by the cost of recomputing lost cache, something the paper does not detail in depth.
What changes for those running AI in production here
For teams in Brazil that pay per token and per GPU hour to keep agents running, the practical takeaway from CLMs is that the cost bottleneck isn't just in model size, it's in how context is handled over a long session. Today, most production agents use manual compaction or RAG (retrieval-augmented generation) with hand-written rules, exactly the kind of solution this study tries to replace.
The research is still academic, and the very critics cited by InfoQ point out that the gains come from specific benchmarks, not general use. But the point worth watching for anyone architecting long-running agents, such as repository-analysis pipelines or automations that run for hours, is whether agent frameworks (LangChain, AutoGen, and similar) start incorporating context management as a behavior of the model itself rather than fixed logic in the orchestrator. That would shift where the complexity lives: less harness code, more responsibility (and less predictability) inside the model itself.
Translated from the Brazilian Portuguese original · Read the original
Nvidia in talks to buy or expand investment in Reflection AI, FT reports
According to the Financial Times, Nvidia is weighing options ranging from putting in more capital to acquiring Reflection AI, maker of the open model Beam, aimed at code and agents.