Xspark AI’s Ding Wenbo Says Robots Need a Tactile ‘Spinal Cord,’ Not Just Better Sensors
Xspark AI co-founder and chief scientist Ding Wenbo says tactile sensing will not replace vision but is needed for contact, slip and safety. He proposes a low-latency spinal-cord layer and a three-tier architecture, while flagging hardware, data and cross-embodiment challenges.
Ding, a Tsinghua University professor whose background spans communications, materials science and machine touch, became co-founder and chief scientist of Xspark AI, also known as Wujie Zhihang, in 2025. According to the interview, he has not defined the company as a sensor company. His judgment is that a better sensor does not automatically produce a smarter robot, and robotic touch may need its own model.
Ding does not reject the dominant vision-language-action path in embodied AI. He called VLA a solid logic, comparable to early Transformer and diffusion policy, and said it will leave a significant mark in the history of embodied intelligence. But robots differ from language models because they eventually extend a hand and touch the world. In Ding’s view, vision has already solved 99% of the problem of how robots perceive the world, and touch does not need to compete with vision. Touch is more like a supplement after contact, when vision fails or is uncertain: whether a cup has begun to slip, whether fingers are gripping too tightly, whether a robotic arm has hit a person, and whether a dexterous action should continue applying force or correct immediately.
The question then becomes how touch should become intelligence. Ding is not fully satisfied with the most direct approach, which is to insert touch as another modality into existing models. From first principles, he said, tactile information is far less rich than visual information but demands higher feedback efficiency. Fusing it directly with an entire visual model can produce large redundancy, or allow a small amount of critical tactile information to be drowned out by vision. This led him to propose tactile-native intelligence. Such intelligence may not need the rich semantic understanding of a visual large model. It may need to be lighter, faster and closer to the body.
Ding uses the human spinal cord as an analogy. When a person is about to fall, the body does not first identify the fact of falling, analyze the cause and then decide to reach out; the body completes a reflex first. Tactile intelligence, he said, can process less information but must be low-latency enough to react at the instant of slipping, collision or imbalance.
Several obstacles remain between recognizing the importance of touch and building a tactile intelligence. On hardware, consistency among different tactile sensors is still poor. Even devices from the same production batch can show clear performance differences, making it hard for models to transfer directly across hardware as visual models do. On data, touch lacks the massive images and videos accumulated during the internet era. The same action can have a different data distribution when performed by another person, another sensor or another robot. More fundamental questions follow: How should touch be represented and embedded? Can force, temperature, shear and vibration collected by different sensors be converted into a common representation independent of specific hardware? Ding summarized the problems for the next two to five years as three crossings: across sensors, across modalities and across embodiments.
Ding does not claim that a mature tactile large model already has an answer. He acknowledged that a more realistic path is a compromise. In the short term, the priority is to build one kind of multimodal tactile sensor well and fuse it with existing models. In pretraining, the focus is on building general capabilities, while tactile information is introduced and fused during post-training for dexterous manipulation and action correction. Only in the long term does he lean toward the view that touch will form its own model and logic.
Xspark AI’s proposed three-tier architecture, described as slow brain VLM, fast brain VTLA plus TWAM, and spinal cord VTA, can be seen as an engineering expression of that current-stage judgment. In this architecture, a vision-language model forms the slow brain, responsible for semantic understanding, high-level cognition and task planning. Touch first enters the fast brain, participating with vision, language and action in long-horizon tasks and dexterous manipulation. One layer down, touch becomes core information for the spinal cord, completing faster local reactions in collision, slipping and safety scenarios. The same touch therefore plays two roles: in the fast brain it helps the robot do better, and in the spinal cord it ensures the robot can react in time.
The unresolved question is more specific than whether robots need touch. Vision will remain the most important source of information for robots to understand the world, and VLA will remain a mainstream path. What has not been clearly defined is whether touch is merely another input added to existing visual models, or whether its special requirements for contact, motion and real-time feedback will eventually grow a model and computing logic different from visual intelligence. Ding said he is now trying to find that answer.
Ding’s academic path, as described in the interview, moved from a communications PhD at Tsinghua to postdoctoral work at Georgia Tech on electronic skin, after he turned down an offer from MIT. He returned to China in 2019 and built the SSR lab in his first year back, according to the interview, while he and colleagues had begun thinking about robotics around 2017 and 2018. At Xspark AI, his current work focuses on turning tactile sensing into a practical layer of robotic intelligence rather than a standalone sensor business.