AI News Feed
Market watch
Companies

Former ByteDance RL Expert Sun Peng Joins Xingchen Intelligence to Boost Physical AI Stack

Sun Peng, a former ByteDance reinforcement learning expert, joined Xingchen Intelligence on Sept. 2 to lead robot RL R&D, completing the company's full-stack Physical AI strategy.

Over the past two years, robot foundation models have mainly focused on whether a robot can perform a task. As robots move from demonstrations to real deployment and commercialization, the question has shifted to whether they can do so stably and reliably in real environments. How a robot can keep revising its policy based on actual execution results has put reinforcement learning back at the center of industry attention.

Sun has long worked on reinforcement learning (RL), intelligent agents and multi-agent reinforcement learning (MARL), with research and industry experience spanning robot RL, large-scale distributed RL systems and LLM post-training. He will also help expand Xingchen's robot RL team and push deployment of these technologies, according to the report.

Sun received a doctorate from Tsinghua University and conducted postdoctoral research at Cornell University and Rutgers University. During his time at Tencent AI Lab and Robotics X, he served as head of the intelligent agent center and conducted research on deep RL and robot control. He trained a wheeled robot with deep RL and adversarial games for end-to-end active target following, and developed StarCraft AI agents TStarBots and TStarBotX, making progress in MARL research.

After joining ByteDance, Sun broadened his practice to large-scale RL system engineering. He led development of ByteRL, a core RL infrastructure used by multiple self-developed game AI systems. His team won the IEEE CoG 2023 strategy card game AI competition, and the Hearthstone AI they developed showed an ability to beat top competitive players. As LLMs emerged, Sun extended RL to model post-training: at ByteDance AI Lab/ByteResearch, he was project lead for the reinforced fine-tuning method ReFT and the mathematical reasoning agent DeltaProver, and at the Seed team he participated in RLHF and pretraining, gaining experience in alignment, reasoning and agents.

For Xingchen, Sun's addition will strengthen the "post-training" link in its push to build a full-stack technology system combining AI models, an embodied OS and tendon-driven robot hardware. The company has adhered to a Design for AI principle since its founding, arguing that robot body, data and AI models should not be designed in isolation. Its Lumo foundation model series includes Lumo-1, which helps robots understand why an action happens, and Lumo-2, which the company describes as the first home-oriented implicit world action model, predicting future world states in latent space before generating actions. The front-end agent Philia is aimed at long-term memory, task management, multi-robot collaboration and natural interaction. Reinforcement learning is, in Xingchen's view, an essential part of scaling foundation models to real deployment, as it enables robots to optimize policies through continuous interaction and feedback.

Sun's future role will focus on connecting the "training after training" approach proven in large-model development with real robot execution, data loops and long-term deployment. After joining, Sun said: "For more than a decade, I have been working on reinforcement learning and intelligent agents, witnessing how RL moved from robot control and complex games to LLM post-training. Now I hope to bring these achievements back to the physical world. Xingchen has already established a complete technical foundation from robot hardware, embodied OS to AI models. What excites me most is working with the team to embed RL deeper into real robot training and deployment, so that robots will not only 'learn a task' but will keep improving through real execution and feedback."