Perplexity and Nvidia launch local AI agent; OpenAI touts Broadcom chip
Perplexity, Nvidia launch no-fee local AI agent; OpenAI reports custom chip gains
Perplexity packages the local model, inference engine, agent harness, tool sandbox and app connectors as one system, according to MarkTechPost. Users can choose between Qwen 3.8 27B and Perplexity's post-trained PPLX 27B; Nvidia's open Nemotron 3.5 Lightning is listed as coming soon. The software is available to Pro, Max, Enterprise Pro and Enterprise Max subscribers on Linux, with Windows support expected in September. It runs on a DGX Spark with the GB10 superchip, 128GB of memory and at least 1TB of storage, or on an RTX GPU with at least 24GB of VRAM, roughly a GeForce RTX 3090 or newer.
Every task begins on the device, and work handled by local models carries no per-token charge. If a step needs live web access or frontier reasoning, the orchestrator stops and asks before sending that single step to one of more than 15 cloud models. Before a cloud call, the harness selects the relevant context, runs a PII classifier and shows the user exactly what would leave the machine. The remote adviser returns text guidance and never receives direct access to local files, tools or the conversation. Tool execution runs inside an OS-enforced sandbox; if the sandbox is unavailable, tool execution is disabled rather than silently downgraded.
In Perplexity's tests, Portable Computer scored 82.6 percent on the company's 53-task Local Knowledge Work Bench when running Qwen 3.8 27B on a DGX Spark, and 85.4 percent with PPLX 27B. Perplexity said the open-source Pi harness scored 77.6 percent and Hermes 74.0 percent on the same model. On BrowseComp, Portable Computer scored 66.7 percent versus 50.2 percent for Pi and 43.9 percent for Hermes, using 51 percent less wall time and 70 percent fewer tokens than Pi. On Terminal Bench 2.1, a fully local run scored 59.6 percent at zero marginal cost; escalating one step to a cloud adviser lifted the score to 73.0 percent at about $0.415 per rollout, while Claude Opus 5 alone scored 82.4 percent at about $0.65. Perplexity said it plans to open-source the Local Knowledge Work Bench.
In a separate announcement Tuesday, OpenAI said testing of its custom inference chip, Jalapeño, showed "a significant performance advance," according to CNBC. The chip, developed with Broadcom and announced in June, "can process more AI workloads per unit of power while also delivering faster responses," OpenAI said, and it plans to begin deployment in its compute infrastructure by the end of the year. CNBC said the results received third-party validation from SemiAnalysis and add credibility to Broadcom's fiscal 2027 AI revenue outlook of more than $100 billion.
CNBC said the OpenAI chip news should not be seen as the beginning of the end of its partnership with Nvidia. Last week, Nvidia said it would provide up to $105 billion in credit support for OpenAI's Ohio data center project, which will exclusively use Nvidia compute. OpenAI also said it expects to "widely deploy accelerators from Nvidia and other partners for both training and inference workloads." Custom-chip competition is expected to be a topic on Nvidia's earnings call Wednesday night, CNBC said. CNBC's Investing Club said the news did not change its concerns about political pressure facing data center construction and that it cut its Broadcom position in half on Monday.