Compute Exchange launches GPU inventory manager, Lightbits ships Inferra
Compute Exchange launched a GPU inventory manager and Lightbits made Inferra GA, both targeting neocloud customers.
Compute Exchange's new inventory capability gives enterprise buyers a consolidated view of available and upcoming AI computing capacity across more than 100 providers, SiliconANGLE reported. The feature extends the company's request-for-quotation marketplace by letting buyers search current, reserved and forward inventory and compare prices, specifications, locations, service-level agreements and commercial terms. It currently covers major Nvidia Corp. and Advanced Micro Devices Inc. accelerators, including the H100, H200, B200, A100 and MI series, as well as complete servers and commitments for inference tokens.
GPU procurement has become increasingly complicated as the number of specialized GPU cloud operators expands. Compute Exchange Chief Executive Carmen Li told SiliconANGLE the market now includes hundreds of such providers, making direct comparison impractical. "You don't want to call 400 salespeople and compare specs and negotiate," she said. "It's just not efficient."
Inventory data flows into the marketplace through application programming interfaces or spreadsheet uploads. Li said the company has worked with more than 100 vendors over the past 18 months and is adding nearly one provider a day. Compute Exchange says it verifies providers through know-your-business and know-your-customer checks; for forward capacity, it reviews purchase orders and colocation arrangements, and for installed systems, it verifies chip identifiers, configurations, network health and performance metrics. "We try to verify and give as much information as we can to the users, not just the specs," Li said.
Compute Exchange does not own GPUs or operate data centers, which Li said removes the potential conflict created when a marketplace sells its own capacity. Buyers initially see technical details, benchmark results, location, price and service commitments but not the provider's identity. Names are disclosed after a buyer selects one or more candidates for further evaluation. "Neutrality is the most critical thing," Li said. "We don't take positions." Providers pay a standard platform fee only when a transaction is completed; the new inventory feature carries no additional charge. Li said the company has generated more than seven figures of revenue this year and is profitable. Requests on the marketplace have ranged from a single node to 20,000 nodes, though limited inventory has prevented the company from fulfilling more than about 40% of request-for-quotation volume.
The marketplace supports reserved and forward contracts for GPUs and tokens, as well as secondary sales of physical hardware. Li said supply is expected to increase next year, but demand could rise just as quickly. Prices vary with chip type, geography and contract structure, limiting the usefulness of simple comparisons. "The price is only one indicator," she said. "You get the specs, get the [service-level agreements], locations and prices. It's up to you to choose." Compute Exchange ultimately aims to evolve from a marketplace into a fuller exchange with automated matching and spot trading.
In the second announcement reported by SiliconANGLE, Lightbits Labs released Inferra, a software engine designed to improve the economics and performance of AI inference by moving key-value cache data beyond the limited high-bandwidth memory attached to GPUs. The company, known as the inventor of the NVMe over TCP storage protocol, is positioning Inferra for neocloud providers and enterprises running large language models with long context windows or many simultaneous sessions. Announced in March, the product manages KV cache across GPU high-bandwidth memory, dynamic random-access memory and NVMe storage, using predictive prefetching to place data near the GPU before it is needed.
KV caches hold intermediate attention data that a model generates while processing a prompt. Their size increases with conversation or document length, consuming scarce GPU memory. When cache data is unavailable, a system must retrieve it from slower memory or recompute it, leaving GPU resources idle and increasing response times. "The data is always available for the compute to operate on," said Ramesh Chettuvetty, senior vice president of product and business for AI solutions at Lightbits. "We do predictive prefetch, which essentially prevents this stall." Lightbits claims Inferra can cut time to first token by more than 100-fold in some long-context workloads, support context windows of more than 10 million tokens on commodity hardware and increase the density of concurrent sessions by more than 16 times. Those figures come from company benchmarks and extrapolations and have not been independently verified.
Chettuvetty said Lightbits has achieved cache hit rates of nearly 99.9% in most tested scenarios. Inferra includes quality-of-service controls, encryption and isolation between tenants, and the ability to move cached data when a session migrates to another GPU cluster. Arthur Rasmusson, director of AI architecture at Lightbits, compared the approach to techniques developed for CPUs as their speeds outran memory performance. The company breaks a context into smaller blocks and feeds the GPU what it needs just in time, allowing operators to maintain larger caches and fit more users on the same infrastructure. Rasmusson said the largest benefits should come with retrieval-augmented generation, AI agents, lengthy prompts and heavily shared GPU services, while value would be limited for organizations with oversized private GPU clusters, few users and contexts that already fit in available memory.
Lightbits said it has production pilots underway and is initially targeting neocloud providers, whose need to raise utilization and margins makes them faster adopters than hyperscalers. The company is demonstrating Inferra at the AI Infra Summit in Santa Clara next week.