Volantis raises $88M to develop photonic inference systems
Volantis raised $88M for photonic AI inference systems and plans to ship an A-1 data center appliance next year.
The Series A deal drew more than a half-dozen other participants, including Kleiner Perkins chair John Doerr and Naveen Rao, the former head of Intel’s artificial intelligence products group. Volantis’ founding team has similar industry credentials: the company employs engineers who previously worked at major chipmakers such as Nvidia Corp. and Broadcom Inc. Their technical achievements include the first commercial implementation of CoWoS, an interconnect widely used in graphics processing units.
Memory bandwidth is a major contributor to AI model performance. It measures the speed at which data travels between a GPU’s processing cores and HBM memory. The more memory bandwidth a GPU has, the faster it can perform inference.
Volantis’ chip architecture uses interconnect wires with a range of more than 200 millimeters. According to the company, that extended range makes it possible to equip an AI accelerator with more than 220 memory chiplets. The design provides more than 30 times the memory bandwidth of current accelerators, Volantis says. The simplest way to increase a GPU’s memory bandwidth is to add more memory modules, but today the number of modules that can be placed on a GPU is constrained by the wires that link them to the processing cores. Those wires are up to 5 millimeters long, and memory modules must sit within that range, limiting the total number of modules that can fit on a chip.
Volantis’ memory interconnects are based on an optical design, transmitting data in the form of light. The light is generated by microscopic devices called VCSELs. A VCSEL comprises three main components: a so-called quantum well and two mirrors. The quantum well turns some of the electricity that runs through the host chip into light, while the mirrors amplify it. VCSELs are easier to manufacture than the lasers that optical networking devices typically use to generate laser light, and they often cost less. They can also be made from gallium arsenide, which is more readily available than the materials most commonly used to produce miniature lasers.
Volantis will ship its chips with the A-1, a data center inference appliance about a third the size of a standard server rack. According to the company, the system features 10 terabytes of memory with 250 terabits per second of memory bandwidth. Volantis estimates that the A-1 will be capable of processing up to 10,000 tokens per second when running a model with 20 trillion parameters. “This will enable real-time frontier inference, restart scaling laws & enable entire code bases in context windows,” Volantis co-founder and Chief Executive Officer Tapa Ghosh wrote in a blog post. “As a starting point, imagine a coding agent that completes a task in 30 seconds rather than 30 minutes.” Volantis plans to start shipping the A-1 system next year.