Alibaba Cloud Unveils Agentic Cloud Stack as Qianwen Expands Personal Agent
Alibaba Cloud introduced AgentCore, Agent Sandbox and new storage products at the 2026 Apsara Conference, while Qianwen detailed Personal Agent progress for health, finance and work. Kantata launched Agent Studio, and Jiyuan Ludong outlined model-routing and RSI work.
Li said enterprise agents need four capabilities to scale: a comprehensive, unified and continuously updated context to form broad awareness; support for autonomous operation and continuous evolution; embedded security, control, trust and identity for trusted collaboration between humans and agents and among agents; and a shift from static rules to real-time perception, dynamic reasoning and autonomous action. Alibaba Cloud rebuilt its technology stack around these pillars, spanning AI Native Cloud, Agent Native Cloud and Context Engine, according to Leiphone.
For AI Native Cloud, Alibaba Cloud showed a full-stack self-developed computing foundation. Its Lingjun Zhenwu M890 super node, described as China's first 100,000-card super node cluster, has supported commercial services for models with more than 2 trillion parameters, including Qwen3.8 Max and KIMI K3. The company plans to launch the Lingjun Zhenwu V900 super node next year, with three times the performance, 200Pbps communications bandwidth, a 6-microsecond latency cluster architecture, self-developed SNPO for all-optical interconnect, support for 1,000-card Scale-Up interconnect and expansion to 500,000 cards in a single cluster. The new CPFS storage system offers up to 100TB/s throughput and billions of IOPS, with a single file system scaled to 100PiB. In model training, it can cut model startup time by 50%, raise peak computing utilization by 30% and lower AI storage costs by 69%, according to Leiphone.
On inference, Alibaba Cloud introduced Tair KVCM, a KVCache management and scheduling system that achieves a 99% effective hit rate and reduces cost per token by 50%. Its KV CacheStore storage acceleration engine expanded the cache coverage time window by 900%, raised throughput by 20% and cut first-token latency by 54% in customer production environments. The new TPN intelligent computing center network uses a two-layer design, increasing access bandwidth by 2.5 times, network scale by 10 times and reducing latency by one-third. Alibaba Cloud's PAI platform added support for an asynchronous AgenticRL training framework; in Qwen post-training, it completed one SOTA model post-training in five days, according to Leiphone.
For Agent Native Cloud, Alibaba Cloud launched AgentCore, an enterprise agent building and governance platform. It standardizes and hosts agent infrastructure, supports long tasks, failure retries, checkpoint recovery and asynchronous execution, and provides secure execution environments, resource isolation, identity permissions and full-link auditing. It manages models, MCP Servers and Skills as assets, and can improve task completion rate by 99% and reduce TCO by 70%, according to Leiphone. Agent Sandbox offers an out-of-the-box, highly elastic and secure execution environment with creation throughput of 100,000 per minute and deep sleep wake-up in under 600 milliseconds, compatible with E2B and K8s. Agentic OS saves more than 30% of token consumption and triples sandbox deployment density. Agentic Computer provides agents with a unified, long-term online dedicated cloud computer workstation, compatible with mainstream agent frameworks, with Computer Use and GUI operation capabilities, controllable permissions and auditable operations.
Alibaba Cloud also proposed Context Engine, a three-in-one engine covering storage, computing processing and context engineering that turns enterprise-wide data into real-time context. The lower layer uses multimodal data storage for dialogues, video frames, point clouds, trajectories and agent feedback; the middle layer covers data integration, transformation, governance, quality and modeling; the upper layer uses a global knowledge base, multimodal semantic integration, continuous memory management and real-time context assembly to organize scattered data into context that agents can call on demand and update dynamically. Agent Context raises accuracy by 50 percentage points and improves token efficiency by more than three times. OpenLake provides a unified all-modal data lakehouse and reduces total cost of ownership by 38%. Apsara Lakebase supports sub-second, zero-copy data branch creation, MC-MaxFrame improves computing performance by 12%, and Flink Streaming Agent provides more than 60 all-modal operators. Li said Alibaba Cloud has formed a complete layout from AI chips and server CPUs to interconnect, network cards, storage controllers and AI software stacks, allowing full-stack coordination and joint tuning across computing, networking and storage.
Qianwen product head Zheng Sishou said the goal of giving every person their own AI assistant has remained unchanged since Qianwen launched last year, according to Leiphone. Relying on the agentic capabilities of the Qwen 3.8 series models, Qianwen is building a new Personal Agent form to provide sustained, personalized intelligent services to 300 million users. With user authorization, Qianwen connects personal data such as health and exercise data to offer more targeted analysis for health consultations and exercise planning. It also connects to large volumes of professional wealth management data and institutional agents, and supports some users in linking their personal holdings. Its open platform now covers more than 20 fields, with more than 1,000 partners applying to join.
In wealth management, users can tell Qianwen their financial goals in a conversation, and Qianwen combines real-time market data, professional institutional information and user-authorized holdings to provide analysis and suggestions closer to their personal situation. Qianwen has packaged financial institutions' professional research capabilities as Skills so ordinary users can access professional decision support. It also carried out special Harness optimization for complex wealth management tasks, improving response performance by 20%. In sports and health, for a goal such as running a half marathon in under two hours, Qianwen can build a long-term, dynamic personal context from the user's continuous exercise status and authorized data from sports health devices and continuous glucose monitors, giving daily training and diet suggestions. When the user reaches the target pace, Qianwen may suggest advancing the goal to one hour and 40 minutes. In work scenarios, teachers can use Skills such as homework grading and test paper generation to quickly grade assignments, identify weak points and generate a math practice sheet. More than 60% of users who handle complex needs in wealth management, health and daily life turn to the work assistant, and Qianwen says dynamically adjusting Harness can save each user 40% of token consumption, according to Leiphone.
At the enterprise Agent practice summit at the same conference, Han Kai, co-founder and CTO of Jiyuan Ludong, gave a speech titled 'From Harness to RSI Flywheel,' according to QuantumBit. Han said multi-model heterogeneity and diversified computing supply will be long-term structural features of the large-model industry. As agents take on more complex tasks, organizing different models and turning application feedback into model improvement is becoming a new technical issue. Data from the National Data Administration showed that China's average daily token calls exceeded 140 trillion in March 2026, more than a thousand times the level in early 2024. Han said a complex agent task often involves more than ten or even dozens of model calls, and performance depends both on the model itself and on system capabilities such as model selection, tool calling and context management. Harness is the runtime framework that supports agent task execution, and model routing is a key capability. Different models have different strengths in capability, cost and response speed, so complex tasks need the right model for each step. Heterogeneous model adaptation and heterogeneous computing scheduling also require specialized capabilities.
Jiyuan Ludong is exploring this direction with the open-source project OpenSquilla. According to the company's end-to-end agent evaluation results, under a specific test configuration, OpenSquilla retained 99.96% of the task quality of a fixed flagship model baseline while reducing cost by 88.9%. In another DRACO deep research evaluation, a multi-model coordinated configuration scored higher than the strongest single-model baseline in that experiment, at about one-third of the cost. Han said the value of the routing layer is to let each step use the right model and an affordable one. Beyond routing, Han described a feedback loop for RSI, or recursive self-improvement: capability needs in scenarios drive evaluation that identifies directions for improving models and scheduling strategies, targeted optimization follows, and upgraded capabilities are then applied back in scenarios for validation. NeoHorse-1, released in September, is an initial validation. The series was post-trained on the Tongyi Qwen3.5 base. According to its initial technical report, in ten benchmark evaluations, the 4B version's macro-average score rose from 58.94 to 64.87, and the 9B version rose from 65.60 to 69.04. The 4B post-trained version surpassed the Qwen3.5-9B base model in five evaluations—VitaBench, τ²-Bench, PinchBench, QwenClawBench and HumanEval—and its overall average score moved closer to the 9B base. Jiyuan Ludong and Alibaba Cloud are conducting scenario testing, evaluation and application validation around the Qwen series, using models in their own Harness and agent products, accumulating feedback and pushing new versions onto the TokenRhythm routing platform. Alibaba Cloud provides cloud service support, while Wuwen Xinqiong provides computing support and infrastructure optimization for NeoHorse-1 training and inference, according to QuantumBit.
According to SiliconANGLE, Kantata Inc. unveiled Agent Studio, a tool that lets services firms build custom AI agents by describing the job they want done in plain language. Kantata says the professional services automation market is producing dozens of narrow agents, each scoped to a single job such as resourcing or risk reporting, and every new one arrives as another silo to maintain. Agent Studio sits inside the Kantata Expertise Engine and works through conversation. A user describes the job, and the Expertise Agent assembles the new agent and refines it the same way, with every build and edit checked for correctly structured inputs and outputs before it is published. No one has to fill in a form or write code. Once live, an agent can be triggered by a person or picked up automatically by a workflow on a single command. Under Agent Studio, a standard is written down once, covering the format, escalation rules and outputs a recurring deliverable needs, then applied automatically rather than rebuilt by whoever happens to be staffed that week. Kantata calls the overhead it removes a 'prompt tax.' Chief Executive Michael Speranza said, 'When it comes to professional services, AI isn't the differentiator — expertise is. Any vendor can give a firm a tool to build AI agents. But if all you're doing is automating processes, you're just racing to the bottom and simply reducing costs, rather than developing differentiation.' Chief Technology Officer Vikas Nehru called an agent that improvises 'a liability in a business where the same deliverable has to hold up in front of a client every time.' Finnish invoice automation company Basware Oyj uses Kantata's software in its own professional services operation. Its chief customer officer, Mark Johnston, said he wants his services team with customers rather than 'doing documentation, figuring out resourcing and entering their billable time,' and that Kantata lets the firm build agents to absorb that work. Kantata said Agent Studio is available now and in use by customers. The company launched in 2022 following the December 2021 merger of Mavenlink and Kimble Applications, a deal Accel-KKR led, and more than 1,500 organizations now run its software.