ByteDance Paper Finds Periodic Weak Spots in DeepSeek's Chunked KV-Cache Compression
A ByteDance paper reports that DeepSeek V4 models can change answers when small token shifts place key information in weak slots created by chunked KV-cache compression. The flaw is described as structural and affects other models using similar compression.
The finding began with a code-completion test. The task was to complete the last token of an FP8 quantization function. The correct output was 8, but after researchers added a string of decorative equals signs at the beginning of the code, DeepSeek-V4-Flash-Base returned 32. The error was not random: it was tied to the number of equals signs. When the number was divided by 4 and the remainder was 0 or 1, the error rate reached 71.3 percent; when the remainder was 2 or 3, the correct rate was 91.5 percent.
Chunked KV-cache compression is one of DeepSeek's tools for reducing the memory cost of long contexts. It cuts continuous tokens into fixed-length chunks and compresses each chunk through a learnable gating layer into a small summary. In a standard Transformer, attention depends mainly on the relative distance between tokens. With chunked compression, a token's position inside its compressed chunk, called its phase, becomes important. The gating layer assigns different weights, or slot gain factors, to positions. Tokens in strong slots are retained and can help the model answer correctly; tokens in weak slots are largely ignored, leading to errors. The paper also reports that attention heads specialize in earlier or later positions within a chunk, making some positions especially fragile.
In a needle-in-a-haystack test, the researchers built a 128K-token context containing about 16,000 key-value pairs. They kept the key-value relationships, the question and the total context length unchanged, and moved one key-value pair to different positions. DeepSeek-V4-Flash-Base showed a maximum accuracy gap of 40.2 percentage points across positions. DeepSeek-V4-Pro-Base showed a gap of 34.8 percentage points. After post-training, the gaps were still 19.1 and 14.8 percentage points. DeepSeek-V4.1-Flash halved the compression step to 2, reducing the gap to 6.1 percent, but did not eliminate it.
To test whether the problem came from DeepSeek's implementation, ByteDance researchers trained Qwen3-0.6B models from scratch with identical configurations except for the compression mechanism. All models using chunked compression showed periodic accuracy fluctuations. A full-attention baseline did not. The fluctuation period matched the compression step. Removing rotary position embeddings did not remove the fluctuations, and replacing the learnable gating weights with simple average allocation did not remove them either. When the compression step was set to 4, 6, 8 or 12, the accuracy period was also 4, 6, 8 or 12. The team concluded that fixed-length chunked compression creates a structural defect that cannot be fully fixed by parameter tuning, changing position encodings or post-training alone.
One proposed direction is dynamic chunking, which abandons fixed token counts and sets boundaries according to semantics and importance. Low-density text can be merged into larger blocks, while dense or critical information can be split more finely. This would restore a form of shift invariance and let the model read some parts quickly and others more carefully. The report says dynamic chunking requires changes across attention mathematics, memory management and hardware alignment. Memory management would move from static contiguous matrices toward topologically sparse matrices with dynamic indexing, and boundaries across Transformer layers are highly similar and can be reused to reduce addressing overhead. The hardware challenge is that GPUs prefer matrices divisible by 8, 16 or 32, so irregular chunks can leave Tensor Cores underused. Zero-padding algorithms can pack variable semantic blocks tightly in physical memory while preserving logical boundaries.
Several research efforts are pursuing this area. ChunkKV, proposed by a Hong Kong University of Science and Technology (Guangzhou) team, uses continuous semantic blocks as compression units, either keeping or discarding a whole block. In the NIAH long-context retrieval benchmark with a KV-cache limit of 128, a LLaMA-3-based ChunkKV reached 73.8 percent accuracy, compared with 58.9 percent for SnapKV. Jina AI's team, led by Xiao Han, has proposed delayed chunking for retrieval, letting the model read the full text before splitting it at semantic boundaries. Microsoft Research Asia's MInference framework targets million-token inference by matching sparse computation patterns to different attention heads and reports up to 10x speedup in the prefill stage without reducing retrieval accuracy. The report argues that as long-context models move into real tasks, usability and information retention are becoming the new measures, not context length alone.
Editor's Summary
A ByteDance paper reports that DeepSeek V4 and other models using chunked KV-cache compression can show periodic weak spots, where a small change in token position changes the answer. The paper attributes the problem to fixed-length compression and says dynamic chunking, semantic compression and sparse attention are among the approaches being explored to reduce it.