Iterate.ai Launches Lifeboat Inference Engine, Claims Up to 6x More AI Agent Sessions per GPU
Iterate.ai launched Lifeboat, an LLM inference engine with confidential computing that it says can run two to six times more concurrent AI agent sessions on each GPU.
The company is targeting a memory problem that emerges when AI agents run at scale. A chatbot question usually triggers one model call, while an agent working through a single task can make dozens, and its context window grows with each call. Hundreds of agents on shared hardware can quickly fill the key-value cache. Iterate.ai said standard inference engines can stall at four or five long-context requests running at once.
Teams that hit that limit often buy more GPUs or move the work to per-token cloud services, handing data to a third party. Brian Sathianathan, Iterate.ai's co-founder and chief technology officer, said banks, insurers and health systems want to run agents on their own data "inside their own walls" without doubling GPU spending.
Lifeboat's approach begins with fair scheduling and admission control, which give each session its share of the GPU so a heavy document-processing agent cannot crowd out interactive ones, according to the company. Optimizations to the key-value cache double its effective capacity while model weights remain at full precision. On mixture-of-experts models, only the experts a given agent needs are loaded. Every session also runs inside its own security capsule with filtering, token budgets and sandboxed execution.
Iterate.ai's own testing used a single Nvidia Corp. RTX PRO 6000 Blackwell GPU running a Qwen 30B-A3B model. Lifeboat held 2,048 concurrent sessions on it, with every request completing, the company said. The same engine with Lifeboat's optimizations switched off topped out at half that number and managed 4,965 tokens per second against 8,714 with them on. In a memory-pressure run of 128 sessions sending 18,000-token requests, Lifeboat kept 99th-percentile time to first token at 1.5 seconds. The baseline needed 189 seconds.
"Before any enterprise buys more GPUs for its agents, it should find out what the ones it already owns can do," said Jon Nordmark, co-founder and chief executive of Iterate.ai. "A data center or neo-cloud that doubles concurrent sessions per card gets that capacity back without adding racks or power."
The top-tier Confidential Computing edition will not serve a request until hardware attestation passes. That check covers trusted execution features in Advanced Micro Devices Inc. and Intel Corp. processors, plus Nvidia's confidential computing mode on the H100, B200, GB300 and other supported GPUs. Model weights are sealed inside the trusted execution environment and stay encrypted in use, whether the engine sits in a cloud confidential virtual machine or on customer-owned hardware.
Early access to Lifeboat drew thousands of downloads within days, the company said, and the engine is generally available now. A free Developer License covers noncommercial and evaluation use on up to two inference servers on a single node. Customers who want professional support can move to the Standard License at $49.99 per month. The attestation features sit in the Confidential Computing edition, which costs $499.99 per month and, like the Standard tier, can be tried free for seven days without a credit card.
Nordmark and Sathianathan also discussed Iterate.ai's Generate platform and its partnership with NetApp Inc. on theCUBE, SiliconANGLE Media's livestreaming studio, at NetApp's Insight conference last week.