Qianwen Office First Launches Qwen3.8-Flash, Claiming 100% Faster Generation and 75% Less Token Consumption
Qianwen Office released Qwen3.8-Flash with a standard mode, claiming a 100% generation speed increase and a 75% drop in token consumption.
The future model supply of Qianwen Office will only consist of standard and advanced modes, according to the reports. Ninety-five percent of daily office tasks can be completed using the standard mode, while only 5% of complex tasks require the advanced mode.
The improvement is attributed to model upgrades and agent collaborative optimization. The new-architecture Qwen3.8-Flash, with a total parameter count of hundreds of billions, delivers performance surpassing Claude Opus 4.6. The Qianwen large-model team and the Qianwen Office team jointly introduced an office-specific version of Qwen3.8-Flash, trained and tuned for multi-step planning, tool selection, and context compression. With inference optimization and a customized Harness architecture, throughput efficiency is further improved. In real office scenario tests, the standard mode's single-task generation speed increased by about 100%, and average token consumption decreased by 75%.
In real AI application scenarios, high performance often means high cost and latency, while low cost tends to sacrifice intelligence. The deep collaboration between agents and models is breaking this "impossible triangle" of performance, cost, and speed. As model intelligence density rises and the two-way optimization between Qianwen Office and models continues, agents are expected to move past token anxiety into an era of abundant and cheap tokens.