FreeToken Runs 753B GLM-5.2 on a Single Workstation GPU, Researchers Say
FreeToken, an edge-native MoE serving engine from UC Berkeley and UT Austin, runs a 753B GLM-5.2 on a single workstation GPU.
Latest news, developments and reporting on Large Language Models.
FreeToken, an edge-native MoE serving engine from UC Berkeley and UT Austin, runs a 753B GLM-5.2 on a single workstation GPU.
Shakespeare debuted as the world's first emotionally intelligent jewelry brand in Shenzhen, selling smart rings for 1,499-2,399 yuan to urban women.
According to ifanr, Elon Musk endorsed a Cloudflare report predicting AI will generate most web traffic. Meanwhile, BMW, Roku, Microsoft, Google and Perplexity are embedding ads into devices, systems and AI answers, sparking user backlash.
MarkTechPost's new tutorial shows developers how to build layered guardrails for LLM-based financial assistants using NVIDIA's NeMo Guardrails framework.
AI systems have become strikingly skilled at professional-level mathematics yet still fail at basic counting, prompting both promise and anxiety, according to a Slashdot report.
Inherent, a London lab founded by DeepMind alumni, says its Faraday agent beat Anthropic and OpenAI models at replicating scientific research using a much smaller system.
Canva cut its 2026 revenue-growth forecast as AI costs soared, prompting SiliconANGLE to argue that enterprises must control AI economics rather than vendor token metrics.
According to TechRadar, Google's $10 million bid to buy Spirit Airlines' old business data for AI training has encountered unspecified obstacles.
At WRC 2026 in Beijing, embodied AI companies shifted focus from flashy demos to stable long-hour operation and viable business models, with logistics sorting and service scenarios leading the charge.
A new open-source course reveals that changing the agent harness, not the model, can dramatically boost performance, and details three run modes with distinct cost trade-offs.
Ordnance Survey CEO Nick Bolton discusses transforming the 225-year-old institution into an AI powerhouse, emphasizing location data as a universal connector.
At the 2026 World Robot Conference, Xinghai Robotics unveiled model upgrades and new robot platforms, while aerial robot startup Guiyu Technology revealed funding and plans to commercialize autonomous flying agents.
A TechRadar journalist tested the BodyPark Atom, an AI-powered home fitness coach, and was impressed by its movement mapping despite initial skepticism.
Geekom has deployed DeepSeek V4 Flash across a $16,000 cluster of four A9 Mega mini PCs linked by USB4, delivering 512GB of memory for local AI inference. The vendor-reported performance of 14.61 tokens per second awaits independent verification.
DeepSeek has introduced V4 Flash Vision Exp, a multimodal LLM that outperformed Anthropic's Opus 4.8 on two image benchmarks, with a paid platform launch and possible open-source release later.
Nvidia research shows a custom harness boosted Claude Opus 5's ARC-AGI-3 score from 30% to 100%.
Linus Torvalds says AI enormously helped him debug a stubborn Intel graphics driver bug, after 24 patches and 18 boots.
SenseTime has open-sourced SenseNova U1.5 Lite, a lightweight multimodal model supporting 3-4k character instructions and native 4K output for stable visual creation.
At the 2026 World Robot Conference, Superdynamic KAI showcased the world's first complete autonomous humanoid robot table tennis match and its full-stack embodied AI systems.
An MIT Technology Review story entwines a child’s question about dead words with an escalating global crisis over AI system Tingsu.
According to Leiphone, Wang Xintao, a core technical researcher of Kuaishou's Kling AI, has left the company, raising uncertainty about the model's R&D pace amid fierce competition.
New embodied AI startup Lingxi Zhichuang launches; WRC 2026 debates world models and 'physical token' business models.
OpenAI CEO Sam Altman said there is no way to reach superintelligent AI without a breakthrough, according to TechRadar. He delivered a blunt reality check on the challenges ahead.
Superwhisper introduced S1-mini, a compact open-weights model that cleans raw ASR transcripts into readable text, runs on a laptop CPU, and achieves 94.8% token accuracy.