AI News Feed
Market watch
Companies

Graphics researcher Tong Xin joins Meshy as chief scientist

QbitAI reports Tong Xin, a 25-year Microsoft Research Asia veteran, has joined AI 3D company Meshy as chief scientist. He will lead long-term research strategy alongside founder Hu Yuanming.

Tong graduated from Tsinghua University with a doctorate in 1999 and joined Microsoft Research China, later Microsoft Research Asia, in its first cohort of researchers, according to QbitAI. He stayed for 25 years, rising to global research partner and head of the Internet Graphics group. His research covers computer graphics and 3D computer vision, including material capture and modeling, texture synthesis, 3D geometry processing and modeling, light transport analysis and simulation, realistic rendering, and 3D facial animation. His Google Scholar citations exceed 21,000. His work contributed to Microsoft technologies including Xbox game development APIs, Xbox compatibility software, Windows 3D printing drivers, and the Direct3D graphics development toolkit, the report said. During his time at Microsoft Research Asia, he collaborated with and mentored researchers including Zhou Kun of Zhejiang University, Xu Kun of Tsinghua University, and Wang Pengshuai of Peking University. Colleagues described him as approachable, hands-on, and able to discuss technical details rigorously while explaining complex problems in plain language, QbitAI reported.

Meshy was founded by Hu Yuanming after he received a doctorate in computer graphics from MIT. The company turns text and images into 3D models. It has 12 million users, and its customers include half of the top 10 global companies by market capitalization or valuation, according to the report. In July, Meshy completed a nearly $400 million Series B round at a post-money valuation of more than 10 billion yuan, setting records for both single-round financing size and valuation in the global AI 3D sector.

Hu has described Meshy's goal as 'AI for Fun.' In his account, after AGI solves many productivity problems, humans will face a central question: how to create, express, connect, and find meaning. Reading, travel, games, and short dramas can generate happiness, but in an era of abundant free time people will seek newer and more effective ways to generate happiness and meaning, and AI is expected to drive those methods, according to QbitAI. Hu said Meshy aims to 'reshape graphics based on generative AI and find the global optimum for materializing imagination.' Tong's views align with that direction. In 2016, he proposed solving graphics content production for 'everyone' and 'everywhere,' allowing anyone anywhere to create visual media. A decade later, that idea has become the goal of enabling any person to turn a sentence into an immersive, interactive world, the report said.

Meshy is conducting basic technical research on three levels, according to QbitAI. The first and most important is to reinvent graphics, moving beyond the local optimum of traditional graphics rendering pipelines and using AI to build a global optimum that does not depend on triangles, the smallest unit used by GPUs to render 3D objects. The second is to build a director system that defines a world's operating mechanisms and 3D skeleton and enables real-time scriptwriting. The third is infrastructure: low-latency, high-quality, low-cost 3D generation and rendering, because even momentary stutters would be magnified hundreds of times in AI for Fun experiences. This infrastructure work connects to Hu's doctoral research and to Taichi, the programming language and compiler he created and that Meshy used at its founding. Beyond these basic research areas, Meshy's new models are addressing specific bottlenecks in AI 3D, such as controllability: whether generated results faithfully match input images in overall proportions, spatial distribution, and surface details. On Aug. 10, Meshy released Meshy-7, which made a significant step in geometry alignment. For organic characters, the model can restore facial expressions, anatomy, and skin folds; for hard surfaces and mechanical parts, components land in image-specified positions without sticking to adjacent parts; for text and patterns, engraved characters and relief patterns are clean enough for industrial uses such as laser engraving.

Tong's recent research lies at the intersection of 3D and video generation. In 2024, he posed a classic question: is 3D merely a special case of video generation? If a video can already 'see' an object from any angle, is it still necessary to explicitly build its 3D model? The answer is not yet clear. QbitAI reported that Tong's frontier research and engineering experience in AI multimodal training could take Meshy's exploration further.

While the article was being written, Hu published new work from the Meshy team, Mora, on his public account. Mora stands for Multimodal Open-world Real-time Architecture. It consists of three parts: a coding agent that generates a game world's skeleton and running code; 3D generation, Meshy's main piece, which enriches the skeleton and outputs control signals to a video model; and a video model that receives 3D scene signals and outputs final visuals and sound effects. Hu called it a technology that goes 'beyond world models,' according to QbitAI. In January, Google's Genie 3 launched a playable demo and claimed to be a general world model that lets users generate explorable, realistic environments from text. Unity fell 24% and Roblox fell 27% in a day as the market reacted. After playing all available world model demos, Hu listed eight shortcomings in current interactive video models: weak consistency, simple interaction, weak physics, short experience time, inability to support plots and complex logic, mainly single-player play, inability to use coding generation mechanisms, and high latency. Mora takes a different route: let coding agents build the skeleton, let Meshy generate 3D, and let video models render the images. Mora 1 is still at an early stage and is intended to validate a framework before scaling toward the goal, the report said. It offers a preliminary answer to Tong's question: at least for now, 3D and video generation need not replace each other and can work together under a coding agent. Over two to three decades, Tong has witnessed successive technology shocks, from hand-written 3D rules to neural rendering, generative AI, and video models that can produce dynamic worlds. As new technologies press the boundaries of traditional graphics, he is setting out again, with young researchers, to create a new graphics, QbitAI reported.