AI News Feed
Market watch
AI Chips & Compute

Lenovo, AMD and Huawei Put Local Large-Model AI at Center of New Devices

Lenovo and AMD showed PCs that run large AI models locally; Huawei unveiled a new chip and phones on the same day.

Lenovo's machine combines Nvidia's 20-core Grace CPU with a Blackwell RTX GPU through NVLink-C2C, offers up to 128GB of unified memory and reaches 1 PFLOPS of AI performance at FP4 precision. It weighs 1.65 kilograms and is 16.7 millimeters at its thinnest point. Because the CPU and GPU share the same memory pool, the laptop avoids the traditional split between system memory and video memory that often limits consumer PCs to much smaller models. Nvidia's reference configuration for the platform lists support for 120 billion parameters and a 1 million-token context window.

The same design philosophy was visible in AMD's IFA opening keynote. AMD introduced the Ryzen AI Max 400 series under the code name Gorgon Halo, raising unified memory to 192GB and claiming it can run local models with up to 300 billion parameters. During the keynote, AMD showed Zhipu's 320-billion-parameter GLM-5.3-Flash running locally on the platform and presented local open-weight models outperforming cloud models on selected benchmarks. AMD also unveiled Threadripper Halo Station, a liquid-cooled desktop workstation with a 96-core Threadripper PRO processor, system memory up to 2TB and Instinct MI350P accelerators; AMD said it can run models with more than 1 trillion parameters.

The emphasis on memory and local execution extends a rivalry led by Apple's Mac lineup. Apple recently updated Mac mini and Mac Studio with up to 512GB of unified memory on the M5 Ultra, and has demonstrated multiple Mac Studio machines linked over Thunderbolt running trillion-parameter models. Lenovo and Nvidia's laptop brings similar capabilities to Windows and the CUDA developer ecosystem for the first time in a shipping notebook, making Windows a more credible host for long-running AI agents, according to ifanr. Microsoft, which is working with AMD on Windows support for unified memory and agent execution containers, described the PC as the home of 'unmetered intelligence.'

Huawei's autumn launch event on Sept. 7 added a mobile front to the same local-AI push. Huawei introduced the Kirin 9050 Pro in the Mate XT 2 tri-fold phone, which has a 6.5-inch external screen, opens to a 10.2-inch display, measures 3.5mm at its thinnest point and weighs about 290 grams. Huawei said the phone performs 42 percent better than the Mate XTs and can run a 30-billion-parameter Mixture-of-Experts model on-device. Pricing starts at 19,999 yuan for the 16GB+256GB version. The company also unveiled the Pura X View, an unusually wide 16:9.5 non-folding phone with a 7,000mAh battery and a 200-megapixel main camera, priced from 5,999 yuan.

The new hardware uses large unified memory and on-device model execution as primary selling points, and open-weight models are playing a central role in proving those capabilities. At AMD's event, models including DeepSeek, Qwen, GLM, gpt-oss and Laguna were used as benchmarks for what local hardware can run. The product launches indicate that competition among PC and phone makers is moving from adding AI-assisted features to determining which devices can carry long-context, agent-style workloads without depending on the cloud.