AI News Feed
Market watch
Robotics

After a Tiny Duck, a Bigger Question for On-Device AI: Running Large Models Locally

Lei Feng Network reports that Hugging Face-backed Pollen Robotics’ 25-cm Microduck spotlights edge AI, but the real challenge is whether large language models can run efficiently on end devices.

The little duck has drawn attention not only because it is cute and affordable, nor solely because robots are starting to reach ordinary developers’ desktops. The deeper signal, the report says, is that AI is moving from models inside screens to physical devices that can perceive, act, and interact. For years, discussions about AI centered on cloud-based large models, parameter scale, training clusters, and computing power. But when AI enters the real world, concrete questions emerge: Can a device respond locally? Can it run offline? Can private data stay on the device? Can power consumption and heat be acceptable? Can the cost support mass adoption? These are exactly the questions edge AI must answer.

Microduck offers one vision built on small models, small chips, and small robots. A low-cost, open-source, trainable robot shows that intelligent devices do not have to start from phones, PCs, or cars. They can also grow out of desktops, education, developer communities, home companionship, and lightweight robotics.

But after the little duck became popular, the report highlights a much larger issue for edge AI: when devices need to handle complex conversations, long-term memory, knowledge retrieval, multimodal understanding, and agent tasks, can large models also run on edge devices? This is the dividing line between edge AI that is merely usable and edge AI that is actually good. Many current on-device AI applications in China still rely on small models or fixed tasks, such as voice wake-up, image recognition, simple Q&A, gesture detection, and low-power sensing. These capabilities make devices move, hear, and see. For AI to become a true agent, however, it must go beyond single-point abilities. Future AI PCs, personal AI hosts, home smart terminals, robots, and industrial edge devices may all need to run more complex large-model capabilities locally: understanding context, processing private data, executing continuous tasks, and working stably without a network.

Porting large models to terminals is not simply copying cloud models. Cloud systems can stack servers, power, and cooling, but terminal devices cannot. A robot, laptop, home edge device, or conference terminal must complete inference under limited power, space, and cost. The real difficulty, the report says, is balancing low power, low latency, data security, offline availability, and mass-production cost.

This is the problem that Chinese chip company Houmo Smart has been focused on. Houmo believes cloud remains important because it makes AI intelligent, but part of that intelligence must happen on edge devices, especially for robots, AI PCs, smart meetings, government and enterprise offices, industrial sites, and home terminals that demand real-time response, privacy protection, offline operation, and continuous running. Houmo’s M50 chip, based on compute-in-memory architecture, is designed for edge and endpoint large-model inference. According to public information cited by Lei Feng Network, the M50 has a typical power consumption of about 10W and delivers 160 TOPS at INT8 and 100 TFLOPS at bFP16. It is available as an M.2 card, PCIe accelerator card, and computing box, covering mobile terminals and edge scenarios.

The chip has already appeared in real products. In May, Lenovo launched the AI host P7, a personal home edge device for agents that uses the M50. Within a body of about 300 grams, it supports local running of models with up to 122 billion parameters and achieves local inference speeds of 50 tokens per second offline. It acts like a pocket-sized local AI workstation. In the same month, Great Wall’s N90 Pro full-stack AI laptop also adopted the M50, using 160 TOPS of end-side computing power to run a 35-billion-parameter large model offline for document writing, offline Q&A, meeting notes, quick memos, and fuzzy search. Houmo has also demonstrated on-edge large-model deployment in China Mobile’s ecosystem, including AI official-document writing, AI coding and office tools, AI meeting assistants, AI knowledge Q&A, general security plus large models, and multi-channel video super-resolution on video-conferencing displays.

These cases show that on-edge large models are not a distant concept. They are entering real deployment through AI PCs, personal edge devices, government and enterprise local solutions, smart meetings, security, and robots. The value of Microduck, the report argues, is not just that small robots have become cheaper. It reminds people that the carriers of edge AI are multiplying: cute robot ducks, desktop AI hosts, local large models in laptops, and edge devices in meeting rooms, factories, parks, and homes. Small models let smart devices enter the physical world first; on-edge large models determine whether those devices can handle more complex tasks and become genuine agents.

Underneath all of this lies the chip architecture. The scarcest resource for edge devices is not absolute computing power but usable computing power: high enough performance, but also low power, stable, low latency, and able to fit inside a real product. In traditional chip architectures, much energy is wasted on data movement, and larger models make the “memory wall” and “power wall” worse. Compute-in-memory reduces repeated data transfers between storage and computing units, putting more energy into actual AI computation. For on-edge large models, energy efficiency is not an afterthought but a precondition for running inside terminals, operating continuously, and scaling broadly.

A $399 robot duck will not transform the robotics industry overnight. But it serves as a light-weight window, letting more people see the direction of edge AI: AI will not stay in the cloud forever, nor will it only live on phone and computer screens. It will enter devices that move, perceive, and collaborate, and it will appear on desks, in meeting rooms, homes, factories, and urban edges. The next phase needs both small models that make devices lighter, cheaper, and more flexible, and on-edge large models that make devices smarter, more autonomous, and closer to true agents. The real big question after the little duck became popular, the report says, is whether the industry is ready to run large models on edge devices.