AI News Feed
Market watch
AI Chips & Compute

Prime Intellect Launches Prime Inference With Serverless and Reserved Serving for Open Models

Prime Intellect has launched Prime Inference, a platform serving frontier open-source models through serverless endpoints and reserved GPU capacity, with GLM-5.3 running on GB200 NVL72 systems.

Before the public release, the platform processed close to a trillion tokens per day internally. That traffic came from reinforcement learning rollouts, synthetic data generation, evaluations and long-running coding agents, the company said.

Prime Inference is the serving layer of Prime Intellect's open training stack. The company already ships post-training tools including prime-rl, verifiers and sandboxes, and serving closes that loop by allowing deployed models to generate production traces that can feed back into training.

Prime Intellect says its GLM-5.3 endpoint ranks among the fastest on OpenRouter and cites a near-zero tool-call error rate and 100% uptime since launch. The API is OpenAI compatible, so any OpenAI SDK can be pointed at https://api.pinference.ai/api/v1. Automatic failover across data centers routes traffic to healthy deployments. The hardware is NVIDIA Blackwell today, with Vera Rubin listed as coming soon. Billing is unified, with team-level usage tracking, though per-model pricing is not yet fully published in the documentation.

The serving stack combines NVIDIA Dynamo, vLLM, Mooncake and FlashInfer. It was built with Inferact and NVIDIA, and fixes are contributed upstream. The target workload is agentic: a typical agent turn adds about 6,000 tokens to a 140,000-token prompt. Prime benchmarks this mix with SemiAnalysis AgentX plus injected cold arrivals.

Prefill and decode run on separate GPU groups. Dynamo handles routing, vLLM runs the model on each group, and decoders pull computed KV through NIXL. Prime reports nearly 40% lower p90 inter-token latency in its tests.

For cache-aware routing, Dynamo's KV-aware router weighs cached prefix overlap against queued work, and sessions stay on the same decoder between turns. Mooncake adds a second KV tier in host DRAM.

On GLM-5.3 running on GB200 NVL72, the interactivity target was 100 end-to-end tokens per second per user. At that bar, a 1:4 prefill-to-decode ratio served the most users, reaching 66 sessions per prefill group at 101 tokens per second per user and 100 output tokens per second per GPU. A DEP8 prefill topology delivered roughly five times more usable prefix-cache capacity than TEP8 on the same hardware. Halving tokens per step from 8K to 4K per GPU cut median queue wait from 550 milliseconds to 110 milliseconds, while median time to first token fell about 20%. NVFP4 KV compression shrank each MLA cache row from 576 to 352 bytes and raised cached tokens per decoder from 1.09 million to 1.63 million. A native sparse-MLA kernel measured about 12.0 microseconds at 15 query tokens, against 17.7 microseconds staged and 13.7 microseconds in FP8, which Prime notes is workload specific. With a BLHNC KV layout, transfer descriptors fell from 19,559 to about 1,940 and mean transfer time dropped from 146 milliseconds to 78 milliseconds.

On tool calls, agents fail when calls carry wrong names or broken arguments. Prime Intellect's team contributed a structural-tag builder to Dynamo for GLM's tool format, and vLLM then uses xgrammar to mask tokens that violate the tool schema. The team also fixed parsing bugs, including a less-than sign being decoded into an HTML entity inside code.

Together AI, Fireworks AI and Baseten also offer serverless GLM-5.3, according to the report. Those providers list input and output pricing of $1.40 and $4.40 per million tokens, with the figures verified on October 2, 2026, and tracker numbers coming from ComputePrices, a third-party price tracker. Prime has not published comparable pricing. Reserved capacity is available now, while one-click dedicated deploys and batch inference are listed on the roadmap.