Jiushi Builds First L4 10,000-Card Cluster to Support Multimodal Foundation Model
Jiushi disclosed an industry-first L4 10,000-card compute cluster, nearly 15,000 accelerator cards, backing its APEX multimodal foundation model and city-level physical AI strategy.
Jiushi said the cluster, made up of more than 10,000 accelerator cards, is a high-performance computing system used to train foundation models. The company described a single AI accelerator card as a high-speed computing unit and a 10,000-card cluster as an 'AI super factory' that can break a complex model-training task into parts and process them in parallel. Jiushi said the level of computing power determines the intelligence ceiling of an autonomous driving company.
The company was founded in August 2021. Its founding team previously led Robotaxi and Robotruck development at Baidu's Silicon Valley research institute and contributed to the first generation of the Apollo open-source platform, according to the report. Jiushi chose the RoboVan sector. Over five years it said it has invented the world's first RoboVan, expanded L4 technology from logistics to other industries, and built a commercial model based on selling vehicles, services and transport capacity.
Since selling its first vehicle in 2023, Jiushi said its global fleet has exceeded 30,000 vehicles, covering more than 20 countries and 300 cities, with cumulative real-operation mileage of more than 270 million kilometers. As scale grew, the company faced long-tail scenarios in real cities that cannot be exhausted by rules, such as unmarked rural roads, temporary construction barriers, irregular obstacles, mixed and disorderly traffic, and sudden light changes in bad weather. Jiushi said it is using large models to learn the physical world and that a 10,000-card cluster is a prerequisite for training those models.
Jiushi said its APEX multimodal foundation model combines large amounts of real L4 driving data with internet language and video corpora and real operational human decision data. After APEX base training, the model feeds back into data, models and evaluation to form an iterative closed loop. The company said the cluster gives its vehicles 'self-evolution' capability: 30,000 vehicles, more than 300 cities and 270 million kilometers of real L4 operational data feed the computing system every day. Stronger computing leads to faster model iteration, which can lower operating costs, attract more customers, put more vehicles on the road and generate more scarce data.
Kong said, according to the report, 'Complex open urban roads can generate high-value data, and the business model itself needs to make money to continuously acquire data and drive technological progress.' Jiushi also said that 'computing power is not a cost, it is the starting point of the flywheel.'
The core of Jiushi's strategic upgrade is the 'Jiushi Brain,' which the company describes as the base for its unmanned vehicle services and city governance ecosystem. At its center is APEX. Jiushi said APEX improves four areas: scene reasoning and decision-making, with predictive simulation and long-horizon game prediction to improve driving smoothness and safety; knowledge fusion, combining internet videos, traffic text and human driving experience to understand traffic lights, police gestures, road signs and local rules across cities; simulation, with offline dynamics and human and vehicle behavior restored at high fidelity, million-level operating conditions run in parallel, and edge dangerous scenarios verified in cloud simulation; and model distillation and vehicle deployment, with one end-to-end driving brain that unifies perception, prediction and planning weights and adapts to multiple RoboVan models.
Jiushi also described a vehicle-cloud collaboration mechanism. It said current mainstream vehicle-side end-to-end models can cover general road driving above 30 km/h but have capability boundaries. Jiushi's architecture deploys models by road conditions and speed. In complex road scenarios at 5-30 km/h, the vehicle-side VLA model performs local real-time inference while calling a cloud VLA model for enhanced understanding. In terminal scenarios at 0-5 km/h, it adds a safety officer Agent that can autonomously complete rescue and escape in extreme scenarios and make higher-level decisions such as detouring, rerouting or pulling over. Jiushi said the difference between its system and L2 is not speed but that the vehicle knows when it 'can't' and actively slows or stops, calls the cloud for analysis and instructions, rather than passively waiting for takeover.
The loop is continuous: real data flows back to the APEX base, the improved base supports cloud model upgrades, and the cloud model is distilled into better vehicle-side models. APEX builds a model matrix covering vehicle-side VLA, cloud VLA and safety officer Agent through knowledge distillation and physical reasoning. Jiushi said its VLA does not force translation into human language before controlling the vehicle; it is more like a 'thought' before action, such as seeing a car ahead and deciding to brake and yield, which is faster and more abstract than saying 'I will brake.' The company said vehicle-side models ensure real-time safety, the cloud provides the cognitive ceiling, and the Agent covers extreme scenarios. It described the 10,000-card cluster as a hardware factory, APEX as a software brain and 270 million kilometers of real L4 operation as a scenario foundation, forming barriers in data, computing, models and engineering.
Jiushi's commercial scale covers more than 20 countries and 300 cities, with more than 270 million kilometers of real L4 operation, a fleet of more than 30,000 vehicles and 100-fold capacity growth in three years. The company said the value of those numbers lies in scarce physical-world data that general large model companies cannot obtain, and that autonomous driving companies without enough real operational data cannot match. Each vehicle encounters police gestures, construction barriers, temporary controls and mixed traffic every day, feeding APEX. Jiushi said such scenarios cannot be exhaustively listed in a laboratory or fully reproduced in simulation, and can only come from real city operations. The company said this is the underlying logic for its move from an unmanned vehicle company to a city-level physical AI service provider, extending from hauling goods to city governance. After large-scale commercialization, Jiushi said it is essentially a large model company with scarce data resources.
Editor's Summary
Jiushi disclosed an industry-first L4 10,000-card compute cluster, nearly 15,000 accelerator cards, to train its APEX multimodal foundation model as it shifts toward city-level physical AI. The company is using more than 30,000 vehicles and 270 million kilometers of real L4 operation data to feed a vehicle-cloud model loop. The move places compute, foundation models and real-world data at the center of its autonomous driving strategy.