CoreWeave expands full-stack AI services to target inference bottlenecks
CoreWeave is adding managed services, the Forge platform and RL Rollouts to speed AI inference and agentic model iteration as inference workloads grow, an executive says.
Inference demand is climbing, according to a survey of CoreWeave customers and prospects by theCUBE Research. The survey found one healthcare customer's inference workload share rose from about 10% in the first year to 40% in the second, with a roughly 50% share expected within 12 months.
Chowdhary said CoreWeave has focused on tuning every layer above the hardware, from the vLLM engine to quantized models and custom speculative decoders. She said the company uses open-source tools and technologies, contributes back to open systems so customers have flexibility, and builds its services on top of one another.
Reinforcement learning adds pressure, she said. When customers train agentic models with rewards and verifiers, inference becomes the bottleneck during rollouts. CoreWeave RL Rollouts, a preview capability built on Nvidia Corp.'s Dynamo framework, loads new checkpoints into a live deployment. In testing, it improved model reload latency by 15x compared with a baseline configuration.
"When you're doing RL rollouts, you're continually creating new model checkpoints and versions and you want those to roll out into your inference setup so you can scale it independently," she said. "And we were able to speed that up by 15x, which means your training runs fast and your inference is scaling while you're continuing to train your model quickly."
Those capabilities now sit inside CoreWeave Forge, a platform launched at the event that connects serving, observability, post-training and evaluation. Forge is free to start, with paid tiers offering additional capabilities, extending access to AI development tools and services to individual developers, Chowdhary said.
"We want to create accessibility for these leading technologies," she said. "Even if you're an individual developer signing up today on your own, you still get the best performance, you still get the best reliability and you don't have to compromise."
The move reflects a shift among specialized cloud providers. Training built the first wave of GPU clouds, according to SiliconANGLE, but serving models faster and cheaper will define the next. Chowdhary said AI developers want to solve a problem as quickly as possible with the best performance and cost to scale. "So, we've been really focused on building up the layers of our stack, building on top of the reliable infrastructure to build more managed services, whether it's for training, post-training, inference."