GLM-5.3-Flash Configuration Puts Kimi and DeepSeek Attention Routes in One Model
A Sept. 30 report by Lei Feng Network says a GLM-5.3-Flash configuration file shows 34 of 45 layers use Kimi-linked KDA linear attention and 11 use DeepSeek-linked KPool-DSA sparse attention, as Chinese model makers increasingly mix linear and sparse attention instead of relying on global attention.
Attention mechanisms generally fall into three groups, the report said. Full or global attention reads the complete token history, preserving the most information at the highest compute and cache cost. Linear attention compresses history into a more compact state to cut the cost of long sequences. Sparse attention keeps token-level history but selects only part of it for computation. Efficient attention research in recent years has mainly evolved along the latter two paths.
The shift is visible in other models too. Qwen3.8-Flash-Next, released in August, kept its earlier structure of three Gated DeltaNet layers plus one attention layer, but changed the layer responsible for precise global reading into Qwen Sparse Attention. In February, MiniCPM-SALA directly combined Lightning Attention and Sparse Attention. In GLM-5.3-Flash, the same broad pattern appears with KDA and KPool-DSA.
MiniMax's experience shows how difficult the trade-offs have been. In late 2023, Zhong Yiran, who later led MiniMax-01 network architecture design, was looking for a company willing to scale up Linear Attention. He had researched the route since 2021, and his largest experimental model had reached 15 billion parameters. After talks with other companies, a conversation with MiniMax founder Yan Junjie at the end of 2023 led MiniMax to take on the unproven route. Zhong recalled that he saw the probability of success as close to 99 percent, while Yan estimated it at about half. More than 80 percent of MiniMax's research and development resources would be involved if the company made Linear Attention the architecture of its next main model.
Early experiments made the bet look less risky. A pure Linear model of about 15 billion parameters approached Transformer performance, suggesting the route could be scaled further. As the model grew, however, retrieval tasks exposed a weakness. Linear Attention compresses historical information into a limited state, which is efficient for remembering themes and general meaning but can fail when the model must recover a specific detail from a long text. MiniMax eventually added Full Attention back. The resulting MiniMax-01 architecture used a 7:1 hybrid: seven Lightning Attention layers followed by one traditional SoftMax Attention layer, giving the model a regular chance to read the full context directly. Zhong later described hybrid attention as a fallback.
The 15-billion-parameter signal was not enough to justify a full training run at a much larger scale. Before training MiniMax-01, the team ran about 3,700 pretraining experiments, changing model size, architecture and parameter combinations to test whether scaling laws still held. In mid-2024, MiniMax began training the 456-billion-parameter MiniMax-01. When the model was released in January 2025, the seven-layer Lightning Attention plus one-layer Full Attention design remained, and the context length reached 4 million tokens.
A second problem appeared as models continued to scale. In multi-hop reasoning, where a model must connect conditions, intermediate steps and conclusions, hybrid attention showed a clearer capability loss than Full Attention. MiniMax redesigned tests and training methods, and at small scale hybrid attention caught up. The company's concern then shifted from the known defect to unknown ones: the Text-01 benchmarks had not exposed the later multi-hop gap, and new tests could not prove that larger models would not reveal further capability losses. In M2, MiniMax returned to Full Attention, accepting higher cost for more predictable capability boundaries and engineering behavior.
Agent workloads then changed the calculation again. In the Text-01 period, long context often meant reading a book, a report or a document of hundreds of thousands of tokens at once. In the Agent stage, context became a growing work history. Search results, web pages, code, tool returns, errors and intermediate results accumulate over dozens of rounds or more, and earlier events cannot easily be discarded. That amplified the cost of Full Attention. MiniMax listed Context Scaling as a core problem for complex Agent tasks when it released M3 in June 2026 and turned to its own MiniMax Sparse Attention, or MSA. Instead of processing the full history at every step, MSA first screens the long context for the parts that need attention computation. MiniMax said official data showed that at a 1-million-token context, M3's per-token compute is about one-twentieth that of the previous generation.
Across these cases, Linear and Sparse attention are being placed in the same model because they address different problems. Linear attention is suited to maintaining a long history at lower cost but is weaker at recovering a specific detail on demand. Sparse attention preserves token-level history while reading only a small part of it at each step. One is closer to remembering an overall state, the other to looking up specific content when needed. The three models differ in implementation, but they point to a common change: attention architecture is moving from the search for one better mechanism toward assigning different mechanisms to different roles.