Lenovo Tianxi's TianxiCode Tops SWE-bench-Live Lite With 71% Solve Rate
Lenovo's TianxiCode, paired with DeepSeek-v4.1-Flash, topped SWE-bench-Live Lite with a 71% solve rate and passed verification.
SWE-bench-Live is described as an authoritative dynamic benchmark in software engineering, not a simple syntax or single-function code test. It uses issues from real GitHub projects to evaluate whether models and agents can generate code patches and fix problems, and it provides a reproducible execution environment. The report said the evaluation is closer to real development work because it tests a system's ability to understand code, locate problems, and complete repairs.
To ensure fairness, SWE-bench-Live operates a Verified review mechanism. Entrants must submit full agent trajectories, and the official team reviews inputs and the information-isolation environment and checks for possible leakage of answers, test cases, or results to agents. TianxiCode and DeepSeek-v4.1-Flash passed that review and received Verified certification, which the report said demonstrates that the code-repair result is reproducible and credible.
TianxiCode is a code-intelligence sub-capability within Lenovo Tianxi AI, focused on code generation and engineering development. Lenovo said the framework will be applied to its AI hardware products in the future.
The report compares the combined system to a software engineer. DeepSeek-v4.1-Flash is the brain that reads code, understands semantics, and forms solutions. TianxiCode is the eyes, hands, toolbox, and workflow discipline. According to the report, a capable brain without hands that can write code and inspect logs, and without strict working habits, cannot solve bugs in complex real-world engineering.
Lenovo pointed to benchmark comparisons showing that without top-level agent architecture and system coordination, even top closed-source models often struggle with real engineering problems. TianxiCode's engineering abilities are described in three areas. Cross-file multi-hop retrieval and precise context management allow it to follow clues through large codebases and locate root causes. Autonomous planning and multi-turn tool calling let it act like a human programmer: listing a troubleshooting plan, opening a terminal to run tests, reading logs, comparing changes, and trying another approach when one fails. Closed-loop patch generation and test-driven self-correction allow it to run code in an isolated environment, detect new errors or regressions, reflect and revise, and continue until the patch passes project acceptance.
Lenovo said TianxiCode is a key research and development placement for the Tianxi AI ecosystem in professional technology. Its entry into the top tier on an international benchmark validates technical strength and clears a path for agents to be used in real development scenarios, according to the report. The report said the final value of code agents is not a single ranking but reducing the cost of debugging and engineering coordination for frontline developers. Lenovo Tianxi AI plans to move TianxiCode's technical results into developers' actual toolchains and, with its agent architecture, apply safe and efficient code tools that can independently solve problems to Lenovo hardware devices.