Robot Brain Developers Face ‘GPT-2 Era’ Data Hurdles as Unitree IPO Value Crashes
Unitree's $66B valuation halved as physical AI startups struggle with data scarcity, prompting developers to seek new training methods.
At last week’s Actuate conference, a gathering of developers building AI brains for robots, enthusiasm was evident. The event tripled in size since its 2023 launch and drew 1,500 attendees, according to organizer Foxglove, a company that helps physical AI model builders manage and visualize data. The risk was also on display: a booth for Avala, another physical AI infrastructure player, promised to solve “the robotics data crisis.”
That crisis is the lack of high-quality training data for AI models. Attempts to build generalized robots capable of any task remain far off, and end-to-end learning for specific tasks has yet to deliver products with reliable commercial performance. Developers say the answer is to better mimic advances at frontier AI labs—find or create more diverse datasets, experiment with different training regimes, and design better reinforcement learning scenarios.
Harry Mellsop, a founder of Antioch, a startup building simulation tools for model builders, suggested physical AI is in its “GPT-2 era,” referring to the OpenAI model that predated ChatGPT. More data and compute, particularly GPUs optimized for ray tracing to create high-fidelity simulations, will be needed to get over the hump.
The furthest ahead are autonomous vehicles, partly because relevant data can be collected from human-driven cars and partly because the main task is avoiding contact, not manipulating the physical environment. Much of the tooling for model-building comes from autonomous vehicle companies; Foxglove, for example, was founded by former employees at Cruise, General Motors’ erstwhile self-driving effort. Car companies are now betting their machine-learning tooling investments will let them compete with dedicated humanoid makers. Tesla is already trying this with its Optimus robot, and AV-focused Wayve and ride-share giant Uber have each launched robotics labs focused on humanoid form factors.
“I think you need to start in vehicles…manipulation robotics is like self-driving five years ago,” Wayve CEO Alex Kendall told TechCrunch. “The data infrastructure, the simulation, ML ops infrastructure, will probably be shared, but the specific world model for the simulator will be a different post-training. There’s going to be a lot more commonality than not, but then there’s going to need to be some differences for different embodiments.” Kendall argued it is too early to commit to any one hardware platform, as sensor and component advances are coming quickly and a truly general model should be more agnostic.
Théophile Gervet, CEO of Genesis AI, a vertically integrated humanoid robotics company that raised a $105 million seed round this year, disagreed. “We’re too early in this wave for a brain strategy to work; our take is there’s lots of opportunities to co-design hardware and AI,” he told TechCrunch. Gervet also addressed how specifically to focus a physical AI business. Robotics companies targeting specific tasks are getting robots into the field—Gritt is building solar farms, Agility is deploying robots in industrial settings, and Bedrock is operating excavators autonomously. Meanwhile, general-purpose humanoids are not leaving labs.
“No customer cares about the general purpose robot that works at 80% success rate,” Gervet said. “We see a lot of other players go general, but there is no value provided because there’s no vertical focus…but then, if you’re building [for a narrow] vertical on top of GPT-2, you’re going to get crushed by the company building on GPT-4.”
Investing in a specific vertical is tempting because it provides both revenue and real-world deployment data. While task-specific data may lack the diversity to push general-purpose models forward, it is important for making a robot that adds value. Bedrock CTO Kevin Peterson noted his company was starting with excavation to understand “manipulation in the wild,” but plans to develop an intelligence layer that stretches across a series of construction machines.
Managing all that data is a challenge, especially because of the density of visual information and the need for precise labeling, a challenge that remains unresolved.