AI News Feed
Market watch
AI Chips & Compute

NVIDIA Announces 64GB DGX Spark Desktop for Local AI Agents

NVIDIA announced a 64GB DGX Spark desktop for local AI agents, fine-tuning and inference, available Oct. 23, 2026.

NVIDIA positions the 64GB model as enough for today's most capable 30-35B class open models, according to the report. It keeps the same GB10 Grace Blackwell superchip, NVIDIA CUDA accelerated AI software stack and ConnectX-7 networking as the original DGX Spark, but replaces the original's 128GB of unified LPDDR5x with 64GB. NVIDIA continues to offer the 128GB DGX Spark for larger single-box workloads; clustering 64GB units adds memory and compute together.

The GB10 pairs a Blackwell GPU with fifth-generation Tensor Cores and a 20-core Grace Arm CPU, with 10 Cortex-X925 cores and 10 Cortex-A725 cores. The GPU delivers up to 1 petaFLOP of FP4 AI compute with sparsity. The 64GB model has 64GB of coherent unified LPDDR5x memory at 273 GB/s, 1TB, 2TB or 4TB of self-encrypting NVMe M.2 storage, a ConnectX-7 NIC with 200GbE, Wi-Fi 7 and Bluetooth 5.3, one HDMI 2.1a port and NVIDIA DGX OS, which is Ubuntu-based. It measures 150 by 150 by 50.5 millimeters and weighs 1.2 kilograms. The listed maximum local model size is up to 100 billion parameters.

Unified memory is the key design choice in the report: the CPU and GPU share one pool over NVLink-C2C at five times the bandwidth of PCIe Gen 5, so weights are not copied between system RAM and VRAM. For agents, several models, their KV caches and tool processes can live in one address space. DGX OS ships with the NVIDIA AI stack, including PyTorch, Jupyter and Ollama. NVIDIA NemoClaw installs with a single command and adds privacy and security controls to OpenClaw agents. NVIDIA OpenShell, part of the NVIDIA Agent Toolkit, adds policy-based guardrails, and NVIDIA Nemotron models are optimized for the box. The system runs on a standard wall outlet, with no server room or special cooling.

ConnectX-7 networking is built in for clustering. NVIDIA says two DGX Spark 64GB systems can be clustered for 128GB of memory and more compute, and NVIDIA Sync Cluster Assistant simplifies setup.

The report also frames the 64GB configuration around 30B-class open models that it says have become good enough for agent work. Meta's Muse Glimmer is a 29.6B dense, text and image model under Apache 2.0, with an estimated quantized footprint of about 17GB, and is positioned as the main agent model. NVIDIA's Nemotron 3.5 Lightning is a 30B MoE model with a 30B-A3B configuration and an NVFP4 checkpoint, positioned as a fast executor for long-running agents. Alibaba Qwen's Qwen3.8-27B is a 27B dense model with about 13.5GB of weights at 4-bit, positioned for general agent and coding work. The reported footprints are weight-only estimates; KV cache and runtime overhead come on top.

Meta says Muse Glimmer was distilled from Muse Spark, the model family behind the Meta AI assistant, and targets local agents with reliable tool calls, long multi-step tasks and recovery from failures, according to the report. Meta reports 51.2 on SWE-Bench Pro and 75.5 on MCP Atlas, with context running to 131K tokens. It is also available as an NVIDIA NIM. Full BF16 Glimmer needs 55GB or more, leaving almost nothing for context, so the roughly 17GB quantized build is the practical choice on 64GB, the report says. Larger models such as DeepSeek V4 Flash need more than one box; NVIDIA's benchmarks run it on four clustered 64GB Sparks.

MarkTechPost described several uses for one box. An always-on personal agent could run Hermes with Muse Glimmer or Nemotron 3.5 Lightning, with tools connected to GitHub repositories, a test runner and an RSS feed of arXiv categories. Running overnight, it could triage new issues, reproduce a failing test, draft a pull request for review and summarize papers. Private notes, code and email would not leave the machine. A coding model could be fine-tuned on a user's repository: QLoRA on a 70B model fits in 64GB, according to the report, and training on a codebase, internal documents and past pull request reviews could produce a local coding assistant. NVIDIA measured about 18,400 tokens per second on a single node for nanochat distributed fine-tuning. A day-1 model evaluation bench could pull a new open-weight model from Hugging Face through Ollama or vLLM the same day, run a question set, score accuracy and measure time to first token and tokens per second.

Token consumption is part of the pitch. Agents burn tokens continuously through tool calls, retries, long context and multi-step plans, and token consumption has grown 14 times since early 2026, the report says. On a cloud API, every token is billed; on owned hardware, there is no per-token fee. The message NVIDIA is making, according to the report, is to run open models and always-on agents on a desk and cluster DGX Spark systems as work grows, instead of relying on a metered API.