Naive AI Ties Model Training to 'AI for AI' Push With Open-Source Naive-N0.5-Flash
Leiphone reports that AI-for-AI research is accelerating, with OpenAI, Anthropic and Google expanding automation. Chinese startup Naive AI, valued at $1.42 billion in seven months, released Naive-N0.5-Flash, an open-source model developed with heavy AI involvement.
The report said AI for AI has become a publicly stated goal at leading overseas labs. In September, OpenAI said its AI had reached the capability of an automated research intern and set up a recursive self-improvement team, targeting a fully automated AI researcher by March 2028 that can generate and audit high-quality training data and take over its own R&D. Anthropic disclosed in a June report, When AI Builds Itself, that more than 80% of runnable internal code was written independently by Claude and that each engineer's output efficiency rose eightfold; in September, the report said, 30,000 AI agents were collaborating on its core internal R&D platform in a digital sandbox for code tests and algorithm experiments.
Google is also increasing investment in AI for AI, including AlphaEvolve, a self-evolving coding agent intended to accelerate Gemini iteration. The report said Demis Hassabis stepped down as Google DeepMind CEO to focus on long-term AGI strategy, global affairs and Isomorphic Labs' AI drug discovery, while Jeff Dean left to found Discovery Loop, which focuses on automated machine learning and reached a $50 billion valuation within weeks, with Google investing.
In China, the report said, AI is entering different parts of model development: Mianbi uses AI to write pretraining frameworks, while Zhipu has AI build inference infrastructure on domestic chips. Some teams start from training frameworks, others from inference systems or evolving agents. Naive AI combines architecture exploration, training systems and inference optimization, and also trains its model specifically to carry out R&D work.
The report places mid-training and post-training at the center of AI for AI. Pretraining builds the model using trillions of public data tokens, while mid-training and post-training use tens of billions to trillions of high-quality tokens to strengthen capabilities. The report cited Zhipu's GLM-5, which after 27 trillion tokens of pretraining added about 1.55 trillion tokens of mid-training across 32K, 128K and 200K context stages, with extra long-context and Agent data. GLM-5.3, released in August, kept the same base as GLM-5.2 and did not retrain from scratch; its gains came from post-training, mainly reinforcement learning, and its internal coding evaluation score rose about 50%.
Naive AI does not pretrain from zero. It starts from open-source bases and concentrates on mid-training and post-training. According to the report, it rebuilt the base attention architecture by retaining sliding window attention for local information and replacing all global attention layers with sparse attention, forming a hybrid of five SWA layers to one DSA layer with no global attention. It replaced MLA in the original DSA with GQA using four KV groups and a lightweight indexer with 16 query heads. Naive-N0.5-Flash maintained a 1M context and trained for 3.25 trillion tokens in three stages: a 50 billion-token indexer warmup that froze other parameters, a 3 trillion-token sparse-attention training phase described as continued pretraining or mid-training, and a 200 billion-token learning-rate decay phase entering SFT.
The report said the training system itself was optimized by AI. When communication was slow, AI proposed different parallel strategies for DSA layers, which read full context, and SWA layers, which use adjacent shards. It also managed memory down to individual operators, locked selected key positions to prevent recalculation from disrupting them, and maximized memory efficiency for 1M long sequences. It discovered, located, fixed and verified bugs including insufficient positional encoding precision and sequence index out-of-bounds errors. Naive AI founder Dai Jifeng's background includes the deformable convolution operator in PyTorch, BEVFormer for autonomous driving perception and the multimodal base InternVL.
Under Naive AI's research system, the report said, human researchers set goals, constraints, evaluation criteria and final decisions, while AI handles execution. The company built AI-centered infrastructure with isolated runtimes, GPU compute and a unified platform for compute, environments, tools, permissions and security. It can run nearly 10 million sandboxes per week and up to 100,000 concurrently. Naive-N0.5-Flash was trained for three research abilities: reproducing frontier papers, running ablation experiments autonomously, and managing research repositories and versions.
In benchmark results cited by the report, Naive-N0.5-Flash surpassed Claude Opus 4.7 and GPT-5.5 on PaperBench, which tests reproduction of frontier papers, and beat Claude Sonnet 5 and Gemini 3.6 Flash on MLE-bench-30, which tests machine learning experimentation. On SWE-bench Pro it scored 73.6, above Qwen 3.8 Max, a 2.4-trillion-parameter model, while using only 15.5 billion activated parameters, less than 6% of its total parameters. The model weights and inference code are released under the MIT license, with API pricing at 0.6 yuan per million input tokens, 2.6 yuan per million output tokens and 0.07 yuan per million cache-hit read tokens.
The report also described two real R&D tasks. In one, AI optimized Naive AI's inference system and produced NaiveRT; in extreme mode, output throughput reached 2,122 tokens per second, 30 to 50 times traditional single-stream decoding. The work involved 151 experiments over six days. In the other, AI explored a field unfamiliar to the Naive AI team from scratch and produced AutoWM, a world model that entered the industry's first tier. Naive AI's long-term goal, the report said, is to have AI deeply participate in developing the next generation of large models.