NEWS

ByteDance explains long-context retrieval flaw in DeepSeek models

Research from ByteDance's Seed team identifies the technical cause of inconsistent long-context retrieval in DeepSeek models, with variations of up to 40 percentage points depending on where information sits in the compressed cache.

Uma equipe do time Seed da ByteDance identificou a causa técnica por trás de resultados inconsistentes na recuperação de informações em contextos longos nos modelos DeepSeek. A diferença de acurácia pode chegar a 40 pontos percentuais dependendo de onde o dado cai dentro da janela de compressão do cache.

What the research found

Researchers from the Seed team, ByteDance's AI research arm, identified a technical mechanism behind a known issue in DeepSeek models: information retrieval in long contexts is inconsistent, even when the model performs well on aggregate benchmarks. According to the study, reported by TechNode, the cause lies in key-value cache (KV-cache) compression into blocks (chunks), a technique used to reduce memory and attention cost in large context windows.

The central finding: the same information can be easy to retrieve at one position in the text and difficult at another, depending on where it falls within the compression window. The researchers named this pattern "phase sensitivity" and measured differences of up to 40 percentage points in retrieval accuracy between different positions.

How cache compression creates the problem

KV-cache is the memory a language model keeps during text generation so it doesn't have to recompute attention over tokens it has already processed. In long contexts, this cache grows quickly and becomes the main memory and latency bottleneck in production. To deal with this, model families like DeepSeek compress groups of consecutive tokens into a smaller number of cache entries, trading granularity for memory savings.

The problem is that this compression isn't neutral with respect to position. Different attention components (heads) end up specializing in retrieving information from specific positions within the compressed block. When relevant data falls in a position that no head covers well, retrieval fails, even though the same information, shifted slightly, would be retrieved without issue.

In summary: this isn't an isolated bug in a specific DeepSeek checkpoint, but a pattern tied to how block compression interacts with attention head specialization. The researchers reproduced the behavior in models trained from scratch to test the hypothesis, which reinforces that the cause lies in the design of the compression technique, not in an isolated training effect of a specific model.

Why this disappears in benchmarks

Most "needle-in-a-haystack" tests measure retrieval at various positions in the text and then report an average. A model can have a respectable average and still have systematic blind spots: positions where retrieval is consistently poor. ByteDance's research shows this mechanism in practice, with variation of up to 40 percentage points between positions within the compression window.

This changes how long-context benchmark reports should be read. An aggregate retrieval score doesn't guarantee the same performance on any real document: if the structure of your use case tends to place critical information at the same relative position within a context block, the practical result can fall well short of the average announced by the model's provider.

What changes for those using DeepSeek in production

DeepSeek is one of the open-weight model families used in RAG scenarios, long document analysis, and agents that maintain extensive conversation history, situations where long context is central to the use case. In practice, many of these deployments go through inference engines that apply their own paging and cache compression techniques to control memory in production.

For those in this scenario, the research offers three practical points:

  • Test with the actual structure of your own documents, not just public benchmarks, since the position of relevant information in your use case may coincide with a weak retrieval zone of the model.
  • Evaluate whether the inference engine you use allows adjusting the aggressiveness of KV-cache compression in tasks where retrieval reliability matters more than memory savings.
  • Be wary of model comparisons based solely on long-context benchmark averages, and look for reports that detail variation by position, not just the aggregate number.

None of these measures solve the problem at its root: they are ways to reduce risk until a structural solution emerges, something ByteDance's study doesn't yet announce.

What remains open

The material released so far identifies the cause of the problem, but doesn't include a published fix for DeepSeek models already in production. It's also unclear, from what TechNode reported, whether the phase sensitivity mechanism affects all variants of the DeepSeek family equally or is concentrated in specific compression configurations.

It also remains open whether other model families that use block-based KV-cache compression, an increasingly common practice for cutting long-context costs, suffer from the same pattern. If the cause is structural to the compression technique, as the reproduction in models trained from scratch suggests, it's reasonable to expect that the problem isn't exclusive to DeepSeek.

Translated from the Brazilian Portuguese original · Read the original