AI News Feed
Market watch
Products & Applications

openJiuwen releases X-Router self-evolving model routing, reports over 50% token savings

openJiuwen has released X-Router, a self-evolving routing engine that decides which model handles each agent request. The team reports token consumption down more than 50% in tests and describes the release as Ascend-friendly.

The release, described as Ascend-friendly, targets what the team calls the default state of most AI applications: one model doing all the work. In openJiuwen's account, a user asking an assistant to translate a single sentence triggers the same thousand-billion-parameter cloud model as a request to write fifty lines of code and debug it over three rounds.

An agent increasingly holds local models, cloud models and services from several vendors at once. That spread creates its own problems. Expensive models are spent on simple tasks, results turn unstable when light models take on hard tasks, and fixed rules cannot keep pace with changing workloads and model updates.

openJiuwen is built by teams from Huawei's 2012 Lab, Huawei Cloud, and its terminal, computing and computing advance units, together with universities, companies and outside developers. The routing engine it proposes sits between the agent and the model pool as a request-level decision layer: it picks a model and a strategy for each turn based on task complexity, and can also organize several models to work together. The team says the system has to be configurable, so businesses can bring their own preference for cost or speed; evolvable, so policies track a moving model ecosystem; and observable, so every decision can be traced back to its basis.

The first of three layers profiles the models. Capability profiles are drawn offline from historical data and refreshed online when a model version or its performance changes, with an update latency of better than one minute, the team says. The profiles record performance by task type and difficulty, along with latency, power characteristics and cost. Routing and execution are kept separate, so the routing layer only judges while the host platform makes the actual call.

The second layer makes the choice. The router reads the recent dialogue window together with tool-call progress; because tool output often fills the tail of an agent loop and pushes the task out of the window, the team preserves the most recent user request so the scheduler still knows what is being asked. It then weighs system state that pure algorithm design tends to miss: KV cache affinity, or whether a request's context matches caches already held by a given model; real-time load, since two similarly priced models can differ twofold in latency; and target availability, which rules out a model that has just timed out or been rate-limited. The decision itself is framed as a multi-objective problem over quality, cost, latency, context and tool support and user preference, solved with a heuristic algorithm that selects the optimal candidate rather than the apparently strongest one.

The third layer makes the system improve, through a design the team calls state decoupling. The algorithm is a pure function that must return the same decision for the same request and state snapshot, with no hidden cross-request memory or randomness. All memory, including past results, exclusions, cache affinity and experience data, sits in an external state layer that the algorithm reads once per turn. After each call the host reports back success or failure, duration, approximate cost and a quality rating, and those reports feed the next round. If state is lost, the system degrades to cold routing rather than failing requests, and the algorithm can be replaced without touching the state layer or the host. When samples are insufficient or the advantage is unclear, the default is to keep the original decision.

That produces two independent paths of evolution: runtime self-learning that draws on real feedback through contextual bandits and lightweight reinforcement learning, and module-level iteration in which profiles, algorithms and policies are upgraded separately.

Above those layers, the platform supports multi-model collaboration. When a problem is hard enough that a single model is not trustworthy, routing can dispatch several models, whose answers pass through semantic deduplication, quality filtering, conflict resolution and compression before being merged. The team says this capability is implemented in WorkSwarm's MoA and can be plugged in through the routing framework, with routing deciding whether and which models to summon and MoA deciding how their outputs combine. Strategy self-orchestration goes further, placing models, skills, tools and sub-agents under one scheduling framework, with feedback hooks and fallback mechanisms used to verify results and the root causes of errors.

The team describes the central difficulty as weighing cost, latency, quality and context or tool compatibility in real time. Pushing cost to the limit can break quality, maximizing quality can break the budget, and optimizing each request in isolation can raise overall cost by ignoring cache reuse. Its answer is not a single universal optimum but a system in which the trade-off itself is configurable, evolvable and observable.

In implementation, the Algorithm layer is pluggable. The test results cited by the team were obtained through the x-router path, which uses five complexity tiers, device-cloud tiering, judgment of local capability limits, an in-process small-model classifier and failure downgrade, with contextual bandits correcting decisions from experience. A separate lightweight router written in Rust predicts quality in a high-dimensional feature space and forecasts task trajectories, is statically compiled, has no runtime dependencies and can be embedded in the host process. X-Router is one algorithm module inside the engine rather than the engine itself, so replacing an algorithm does not change the framework and replacing the framework does not change the algorithm.