Positron raises $875M for HBM-free AI inference systems using smartphone memory
Positron AI raised $875 million in a funding round led by NEA and other investors, valuing the HBM-free inference appliance maker at $5 billion. The company plans to tape out its Asimov chip and begin mass production of Titan systems in the second half of 2027.
The funding comes as large language models frequently move data between the computing and memory circuits of graphics cards. The speed at which they do so is measured by memory bandwidth. The higher a chip's memory bandwidth, the faster it can perform inference. Server-grade graphics cards ship with HBM memory, the RAM variety with the highest memory bandwidth on the market. HBM supply is currently well below demand, and there is also a shortage of the interconnects used to integrate HBM memory with graphics cards' computing circuits.
Positron says it has sidestepped the supply chain crunch. Its inference appliances substitute HBM with LPDDR5X, a memory variety mainly used in smartphones. LPDDR5X is more readily available than HBM and costs less. The major catch is that it has significantly less memory bandwidth, which usually translates into slower inference. Positron says it has found a way to bridge the performance gap.
According to the company, most AI accelerators use less than 30% of their HBM modules' memory bandwidth. Positron's inference appliances unlock more than 90% of LPDDR5X's throughput. The company says the increased hardware utilization makes up for the difference in theoretical peak performance.
Positron's flagship system is called Titan. It ships with up to 18.4 terabytes of LPDDR5X that can provide 23.68 terabits per second of memory bandwidth. According to Positron, that RAM pool enables a single Titan appliance to run an LLM with 32 trillion parameters and a context window of 10 billion tokens. Titan performs inference calculations using a custom chip called Asimov. The processor is built around a systolic array, a set of identical computing modules that each include co-located memory. Asimov uses its co-located memory to store LLM weights, numerical values that play an important role in LLM output generation.
The chip's systolic array is supported by modules optimized to run activation functions. Those are code snippets that determine which of an LLM's neural networks should participate in an inference task. Asimov also includes central processing unit cores that function as a programmable escape hatch. They can take over tasks that the chip's other modules are not optimized to perform. Each Titan appliance includes up to eight Asimov chips. Customers can link multiple systems into clusters with up to 16,384 accelerators.
Titan and Asimov are not yet in production. According to Positron, simulation data indicates that a server rack powered by its silicon can process up to 26 times as many tokens per dollar than Nvidia Corp.'s Blackwell GB300 NVL72 appliance.
"Our focus now is to tape out Asimov, bring Titan to production, and scale manufacturing to meet the demand in front of us," Positron Chief Executive Officer Mitesh Agrawal said. "This financing gives us the resources to do exactly that."
Positron expects to tape out Asimov at the end of the year using Taiwan Semiconductor Manufacturing Co.'s three-nanometer node. Mass production is set to follow in the second half of 2027. In conjunction, Positron will ramp up manufacturing of its Titan appliances. That effort will place an emphasis on securing LPDDR5X supply commitments from partners.