ByteDance Seed Finds DeepSeek-V4 Performance Shifts With Token Position, Tied to KV Cache Compression
ByteDance Seed researchers found DeepSeek-V4 models can flip answers based on token position, with retrieval accuracy gaps up to 40.2 percentage points in 128K tests. The pattern is tied to chunked KV cache compression.
The researchers began with a code-completion test on DeepSeek-V4-Flash-Base. They took an FP8 quantization function from DeepSeek-V4’s official inference code and asked the model to complete the final token. The correct answer was 8, because the code needed to finish an FP8-related type conversion, but the model sometimes produced 32 instead. To isolate the cause, the researchers placed a decorative docstring containing repeated equals signs before the code and varied the number of equals signs. The code, the completion position, and the correct answer stayed the same, yet the model’s answer changed repeatedly.
When the padding length modulo 4 was 0 or 1, the model tended to answer 32. When the remainder was 2 or 3, it tended to answer 8. The pattern repeated every 4 tokens. In one set of positions, the average probability of the wrong answer 32 reached 71.3%, while the correct answer 8 had 26.4%. In another set, the correct answer 8 rose to 91.5% on average, and the wrong answer 32 fell to 7.2%. Changing a few irrelevant characters at the front was enough to produce a sharply different judgment on the same question.
The team then tested a classic long-context task, needle-in-a-haystack retrieval. Researchers built a 128K-token context containing about 16,000 key-value pairs and asked the model to find the value for a specified key. The key-value relationships, the question, and the total context length were held constant, while the target information’s position relative to the compression window boundary was adjusted. DeepSeek-V4 series models showed clear periodic swings in accuracy. DeepSeek-V4-Flash-Base had a maximum accuracy gap of 40.2 percentage points across positions, and DeepSeek-V4-Pro-Base had a gap of 34.8 points.
Post-training reduced the gaps. DeepSeek-V4-Flash-0731 narrowed the difference to 19.1 percentage points, DeepSeek-V4-Pro-0813 to 14.8 points, and the newer DeepSeek-V4.1-Flash-0910 to 6.1 points. The periodic differences remained. DeepSeek-V4’s fluctuation period was 4 tokens, while DeepSeek-V4.1’s became 2 tokens, matching the KV cache compression step used by each generation.
Chunked KV cache compression is designed to make long-context processing more memory-efficient. Large models must store key and value information for many historical tokens for later attention calculations, and this cache can become a bottleneck as context grows. The technique groups consecutive tokens into windows and compresses information within each window into fewer cache entries. The researchers found that the periodic weakness may lie in this chunking process. The position of a piece of information relative to the compression window boundary is called its Phase. Models showed systematic differences in retrieval ability across phases, a phenomenon the team named Phase Sensitivity.
The problem could not be reduced simply to information being split between two windows. Even when a key and its value both fell inside the same compression window, retrieval accuracy varied greatly by position. That indicated the issue also involved how the model writes information into the compressed cache and how it reads from that cache later.
To confirm the mechanism, the Seed team trained models from scratch based on the Qwen3-0.6B architecture. They built several KV cache compression schemes and used a full-attention model without chunked compression as a control. All tested chunked compression models showed periodic changes corresponding to the compression step. The full-attention baseline did not show the same degree of periodicity. When window size and compression step were varied, the period mainly followed the compression step: a step of 4 produced roughly a 4-token period, a step of 6 produced roughly a 6-token period, and a step of 8 behaved similarly. The phenomenon persisted even without RoPE position encoding or when learnable compression weights were replaced with a simple average. The design of chunked compression itself, the researchers concluded, may introduce periodic retrieval weaknesses.
Further intervention on attention heads showed that different heads contribute differently to different phases. Some heads are better at handling information in certain positions, while others are more useful in other positions. The researchers call this Phase Specialization. The model appears to form an internal division of labor in which different attention components develop preferences for different positions inside the compression window. That division can help retrieval, but it may also leave some positions relatively weak. A simplified theoretical analysis of training suggested that gradient flow may push the compression module toward stable position preferences, which would explain why the periodicity is not simple random noise and may be a natural result of learning to compress information.
Such problems may not be visible in ordinary benchmarks, because standard evaluations usually aggregate many test results into an average score. A model could perform well at some positions and poorly at others while still receiving a decent average. The Seed team proposed that evaluations of models using chunked KV cache compression should place the same information at different compression phases and measure each case separately. Compression still has practical value given the memory and compute costs of long-context inference, but saving cache may mean the model no longer treats information at all positions equally. The experimental results also showed that post-training and architecture iteration can significantly narrow the gap. The paper is available at https://arxiv.org/pdf/2609.36322.
Editor's Summary ByteDance Seed researchers found that DeepSeek-V4 models can answer the same question differently depending on the token position of relevant information, with accuracy gaps up to 40.2 percentage points in 128K retrieval tests. The effect is tied to chunked KV cache compression and appears as periodic phase sensitivity, while later models and post-training reduce but do not eliminate it. The team recommends phase-aware evaluation for compressed long-context models.