Broadcom launches VMware Private AI Cloud to integrate AI workloads with private infrastructure
Broadcom launched VMware Private AI Cloud for AI and traditional workloads with data governance.
“It brings the private cloud infrastructure with the private AI services into one, enabling production inference and agentic AI with the data sovereignty, the compliance, and the cost predictability required for our customers,” said Prashanth Shenoy, chief marketing officer and vice president of marketing for VMware’s cloud platform, infrastructure and solutions organization, as quoted by SiliconANGLE.
The foundation of the new offering is VCF 9, which supports virtual machines, containers, inference workloads and agentic applications on one platform. According to Broadcom, VCF 9 has more than 3,000 customer deployments representing over 19 million allocated processor cores. Its nonvolatile memory express memory-tiering technology can replace lower-cost flash storage for some dynamic random-access memory and has reduced per-host costs by up to 42% in customer deployments while maintaining performance, Shenoy said. Cluster-wide storage deduplication is intended to further reduce capacity requirements, while token monitoring, graphics processing unit tracking and an AI metrics dashboard provide operators with visibility into resource consumption.
VMware AI Factory packages VCF with hardware, accelerators, AI software and models that Broadcom and its partners have tested together. Initial configurations include Advanced Micro Devices Inc.’s Instinct MI350-series graphics processors and servers from Cisco Systems Inc., Lenovo Group Ltd. and Super Micro Computer Inc. Integration with MetalSoft Cloud Inc.’s orchestration platform automates provisioning and lifecycle management across heterogeneous bare-metal systems through the VCF operations console.
The platform supports accelerators from AMD, Intel Corp. and Nvidia Corp. and uses vLLM as its default model for inference and LLM serving. Broadcom said customers can run more than 150 open-source and commercial models. Newly validated models include Nvidia’s Nemotron 3, Google LLC DeepMind’s Gemma 4, NEC Corp.’s Japanese-language cotomi, Alibaba Group Holding Ltd.’s Qwen 3.7-Max and GLM 5.2 from Jingsheng Hengxing Technology Pte. Ltd., better known as z.ai.
“Not every AI use case requires a frontier AI model and a large language model to be deployed,” Shenoy said. Model choice increasingly depends on the job, cost and governance requirements. New multitenant model sharing will let an organization deploy a model once and make it available to multiple business groups while isolating each tenant’s data.
Broadcom is designating Tanzu Platform as the agent layer of VMware Private AI Cloud. The Cloud Foundry-based platform-as-a-service that serves a pre-engineered AI application development platform will provide deny-by-default sandboxes. Agents cannot access application programming interfaces, networks, Model Context Protocol servers or the internet unless permission is explicitly granted. Credentials are kept in a separate store so agents can’t view or disclose them. An out-of-the-box development harness will include approved skills, workflow buildpacks, human-review controls, and memory services. A curated marketplace will provide access to vetted models, tools, skills and data products, while an AI gateway will monitor, rate-limit and log agent actions.
The new AI-ready data foundations also address the other side of agent governance: controlling the information agents use. Tanzu will ingest and parse structured, unstructured and multimodal information inside the customer’s environment, add metadata and a semantic layer and turn it into governed data products. Those can then be published in the Tanzu marketplace with role-based access controls and lineage information. “You simply give access to that curated data product that is consistently kept in sync,” said Purnima Padmanabhan, vice president and general manager of Broadcom’s Tanzu Division, as quoted by SiliconANGLE. The Tanzu capabilities are scheduled to become generally available in fall 2026.
AgentMinder, which is generally available immediately, provides an additional control plane independent of the agent runtime. It assigns agents identities and evaluates each attempted action against the agent’s owner, declared mission, intent, approved tools and authorized resources. A gateway can allow, deny, redirect or log actions accordingly, according to the report.