AI News Feed
Market watch
AI Chips & Compute

Ai2 Replaces Priority-Based GPU Scheduler With Time Budgets and Fair-Share Allocation

Ai2’s AI Infrastructure team has replaced a priority-based scheduler for its Nvidia GPU clusters with GPU time budgets, hierarchical fair-share allocation and a time-slicing contract, aiming to make allocation more transparent and improve how often high-value workloads receive resources.

Ai2 manages thousands of Nvidia H100, B200 and B300 GPUs in clusters ranging from 88 to 1,024 GPUs. The clusters are built for large-scale distributed training of AI models and serve about 150 internal researchers whose work covers LLM and VLM training, robotics reinforcement learning simulation, and post-training for scientific agentic use cases.

Demand for GPU time far exceeds supply. Based on submitted workloads, Ai2 said it has outstanding requests for two to three times the number of GPUs that are available at any moment. In practice, every available GPU hour on the cluster has two to three different research workloads competing for it.

The previous priority-based scheduler allowed workloads to opt out of preemptability. Each team had a limit on concurrent GPUs that could be used by workloads protected from preemption, while preemptible workloads could exceed that limit on idle GPUs. The team said this produced predictable pathologies. Users parked no-op workloads, or 'squatting', so they could connect to them when needed; this happened because researchers could not launch debugging workloads with low enough latency to solve problems in real time. Priority inflation followed, with 100% of scheduled workloads eventually using HIGH priority, leaving lower priority levels starved of GPU time. Because preemptability was optional, on-call engineers spent a majority of their ticket response time negotiating the organized shutdown of non-preemptable workloads running on hosts with known maintenance problems.

Ai2 said it was slow to identify the root causes when these problems emerged. Its initial attempts to ensure the most important work received GPU time focused on tighter control of how priorities were set and, ultimately, working around the priority-based scheduler by explicitly assigning GPU monopolies to important projects. The team later described the situation as a laboratory for observing the 'tragedy of the commons': individuals competing over a scarce, shared resource and, by seeking to maximize individual outcomes, achieving a non-optimal global result and abusing the underlying resource.

The team noted that resource allocation is a research domain combining algorithm development, economics and system management. A central problem is that users often know the value of their own jobs better than the organization does, but they may have incentives to hide that value or hold on to resources even when doing so hurts total performance. It cited a 2011 paper introducing Dominant Resource Fairness by Ghodsi et al., which recounted an anecdote in which a search company provided dedicated machines to jobs only if users could guarantee high utilization. The company soon discovered that 'users would sprinkle their code with infinite loops to artificially inflate utilization levels', according to the post.

The classic solution to a tragedy of the commons is to privatize the shared resource, because owners are incentivized to maximize the value of their property. Ai2 was already doing a version of this when it assigned teams monopolies over sets of GPUs, but the team said the approach was too coarse. It caused GPUs to sit idle due to the seasonality of research. Teams are ready to run experiments and training at different times, so assigning a monopoly would ensure there would be times when no jobs were ready to execute while another team was left waiting for capacity. The team said it was manually solving a knapsack problem, trying to fit dynamically changing research needs into a static schedule.

The new system is intended to retain the ownership incentive while moving allocation decisions into a more transparent process. Ai2 said the shift changed the debate about how much GPU time each research project deserves from a case-by-case operational task to an administrative budgeting process. The team frames its work as a pyramid of four metrics that build on each other: availability, or how often hardware is healthy and ready; occupancy, or the fraction of available time assigned to a workload; impact, or how often the most valuable workloads are chosen to receive resources; and utilization, or the fraction of GPU capacity used over a workload’s lifetime. The blog post focuses on improving the impact layer of that pyramid.