groundy
Infrastructure & Runtime

Priority Queues or Budgets: How Ai2 Divides an Oversubscribed GPU Cluster

Ai2 reports that replacing priority queues with GPU budgets and fair-share allocation delivered 98% of owed hours, though evidence is self-reported and limited to one 30-day.

Published 8 references
Two jointed resin dinosaurs flank a yellow-topped stool on an ivory background. The larger green dinosaur walks away with its tail wrapped around a stool leg; the smaller black dinosaur raises a foot toward the seat.
On this page12 sections

If your shared training cluster has two to three times more GPU demand than supply, Ai2’s infrastructure team reports a concrete alternative to the priority queue: GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract, with a 30-day audit showing 98% of owed GPU hours actually delivered to teams. The result is worth taking seriously, with one condition stated up front: every outcome number in this article comes from Ai2’s own account of its scheduler change, a single-organization, self-reported blog post. No one has independently replicated it, Ai2 did not measure research output, and the evidence says nothing about clusters smaller than Ai2’s smallest or fleets mixing training with inference serving. What the post does provide, and what almost nobody else publishes, is an organization-level description of the administrative change itself, with fairness-delivery numbers, queue-wait numbers, and the operational price attached.

The oversubscription math

The premise is scarcity that a queue cannot argue away. Ai2 manages thousands of NVIDIA H100, B200, and B300 GPUs in clusters from 88 to 1024 GPUs, serving about 150 internal researchers, according to its infrastructure blog post. Demand outruns that fleet by a wide margin: “Based on submitted workloads, at any moment in time we have outstanding requests for 2-3x more GPUs than are available. One way to think about this is that every available GPU hour on our cluster has 2-3 different research workloads competing for it.”

When every GPU hour has two or three claimants, scheduling stops being a throughput problem and becomes a legitimacy problem. Someone has to decide which of the three waiting workloads gets the hour, and the question is who makes that call, with what information, and whether the losers accept the answer. Ai2 frames the whole task as a pyramid of four metrics that build on each other, per the same post: availability (healthy hardware) at the base, occupancy (available time assigned to a workload) above it, then impact (the most valuable workloads chosen), with utilization (capacity used over a workload’s lifetime) as the capstone. The scheduler change targets the third layer: choosing well, given that the hardware is already up and already busy.

What Ai2 actually replaced

The old system was a priority-based scheduler, the familiar arrangement where projects carry priority levels and contested cases get resolved by operations judgment. In its place, per Ai2’s post: “We recently replaced a priority-based scheduler with a system including GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract. As a result, we shifted the debate about how much GPU time each research project deserves from a case-by-case operational task to a transparent administrative budgeting process.”

The three mechanisms divide the work:

  • GPU time budgets give each team an allocation of GPU hours, set by managers rather than by whoever argues loudest in a ticket. Allocation decisions sit with the person who has context: a lead researcher within a project, a principal investigator within a program, a lead program manager or the CEO across programs. Unallocated (free, preemptible) time still exists, but any trick to grab GPU time now draws from the benefiting user’s own budget, which makes squatting expensive.
  • Hierarchical fair-share turns budgets into scheduling order. The scheduler tracks occupancy over a sliding lookback window (Ai2 defaults to 7 days) and sorts workloads from under-utilized allocations above over-utilized ones, so teams that fell behind their budget get priority until they catch up. Ai2 is explicit that the algorithm is not new: hierarchical fair-share descends from the 2009 Hadoop Fair Scheduler and is in active use today in Slurm’s Fair Tree and YARN’s Fair Scheduler. What changed is the input: the tree mirrors the research program structure and the weights are manager-set budgets rather than static quotas.
  • The time-slicing contract is the enforcement piece. Each workload declares a minimum runtime, the shortest occupancy needed to bank progress, and is protected from preemption during that window. After it, the scheduler may preempt and automatically re-queue resumable workloads. A minimum runtime of zero marks the workload as unallocated: free of budget charge, preemptible at any time.

That third mechanism is where the model stops being paperwork and starts moving running jobs, which is also where it can hurt people.

The 30-day evidence

Ai2 audited the new system over a 30-day window, and the fairness claim is the headline, per the post: “Over the 30-day test period, teams were delivered 98% of the GPU hours they were owed, and 13 of 15 team allocations received 95% or more with the worst case receiving 90%.” (the same post). For a system whose core promise is that you get what the budget says, that worst-case floor is a real result.

The responsiveness gain is sharper, and needs a caveat attached to it: “Debug workload p90 queue wait time fell from 2 hours to 30 seconds under the new scheduler, against a simulated prediction of 6 hours to 5 minutes from hand-crafted test scenarios.” Ai2 itself flags that “the smaller sample size of debug workloads in the baseline meant there was higher variance in those measurements.” Read the 2-hour baseline as noisy; the direction (interactive work no longer waits behind long training runs) is more reliable than the exact ratio.

The number to read most carefully is occupancy, per the same post: “Occupancy on the cluster held steady at 98% before and after the change, with demand exceeding capacity by 2-3x in both periods.” And, per the same post, “18% of delivered GPU time was unallocated, which is how we maintained high occupancy during periods when funded use cases were not ready to run.” Two things follow. First, as the post reports, the old priority queue was already keeping the cluster 98% occupied, so the demonstrated win is equity and wait times, not added capacity. Anyone piloting this model should set success criteria on fairness delivery and queue waits, not on throughput. Second, occupancy is not usefulness: nearly one GPU hour in five was unallocated fill, work charged to no budget, preemptible from the start. That fill may be valuable (per the post, it “prevents teams from ever declining free GPU cycles”), but a dashboard showing 98% occupancy, per the same post, does not mean 98% funded research.

Where time-slicing bites

The clearest failure mode in Ai2’s report is self-inflicted and worth quoting at length: “In the old system, a researcher could hold such a session for up to a week. With time-slicing, they were subject to the 8-hour cap on protected runtime, after which a session becomes preemptible if it exceeds its allocation. We did not appreciate the extent to which researchers were dependent on the volatile state of these sessions. Being preempted meant waiting to secure a new session and also rebuilding their state by hand.”

Interactive sessions, notebooks, debug environments, half-configured experiments, accumulate state that lives nowhere but the running process. A scheduler that treats every workload as a resumable batch job will destroy that state on a schedule. Ai2’s own fix direction is implicit in the complaint: checkpoint and restore for sessions has to exist before the 8-hour cap is humane.

There is also a subtler effect Ai2 is still investigating: minimum-runtime protection may be making it harder to place the largest workloads. The scheduler now has fewer opportunities to interrupt many jobs at once to free a contiguous block for a big pending training run, and Ai2 is using its simulator to reproduce the problem while measuring ground truth in production. The protection that makes small jobs safe can starve the biggest ones.

The literature says this pain is structural, not incidental. Distributed training jobs are gang-scheduled, meaning all assigned GPUs must appear simultaneously and stay dedicated for the job’s duration; per a paper on dynamic multi-objective GPU cluster scheduling, “This requirement can cause significant queuing delays for large multi-GPU workloads and increases the risk of starvation under static policies.” And preemption is not free: interrupting a distributed training job, saving and loading model state to host memory and potentially reallocating GPUs, “incurs large overhead on the order of seconds to minutes,” per work on prediction-assisted DDL scheduling, which is why some recent scheduler research has gone deliberately non-preemptive. Seconds to minutes per preemption event, multiplied across a fleet rebalancing every few hours, is a tax that must be weighed against the fairness gain.

What the scheduler literature says

The strongest intellectual counterweight to Ai2’s model is Themis, which argues that the fair-allocation schemes GPU clusters inherited from big-data schedulers “are far from effective” for ML. The reason, per the Themis paper: “First, unlike batch analytics workloads, ML jobs have long running tasks that need to be scheduled together, i.e., gang-scheduled. Second, each task in a job often runs for a number of iterations while synchronizing model updates at the end of each iteration,” making jobs placement-sensitive, so co-locating tasks on the same machine or rack yields significant speedups. A fair-share system that treats GPU hours as fungible ignores both properties. Ai2’s answer is the time-slicing contract: fair-share decides order, but minimum runtimes decide interruption, and large gang-scheduled runs get long protected windows. Whether that answer suffices is exactly what Ai2’s large-job placement investigation is testing.

The baseline everyone is escaping is genuinely bad. Real GPU cluster deployments report average utilization near 50%, attributed to fragmentation, heterogeneous workloads, and static policies, per the same paper. Fragmentation alone is a quiet drain: the standard practice of allocating same-type, co-located GPUs to a training job leaves “1-2 unallocated GPUs scattered on different machines, causing substantial wastage of expensive AI resources,” per topology-aware placement research. That idle-capacity failure mode is not the one Ai2 started from: per its post, the cluster was already 98% occupied under the old scheduler, and what the priority queue produced was GPU squatting and priority inflation, not wasted capacity. The new system never had idle hardware to reclaim; its job was deciding who got the busy hours, which is what the fairness-delivery and wait-time numbers measure.

There are also alternatives to the budget model that compete on a different axis. Pollux optimizes goodput rather than fairness, adaptively reallocating GPUs and tuning batch sizes as jobs progress; in its HPO experiments, per the Pollux paper, “Pollux completes HPO 30% faster due to adaptive (re-)allocation of resources as trials progress and adaptive batch sizes” versus a static 4-GPU baseline. And preemption is not the only lever for reclaiming time: the ONES scheduler’s elastic batch-size scaling “does not need to stop the original training process” while adjusting effective resource allocation at “almost invisible cost,” per the ONES paper. These systems answer a different question (how to make each job finish faster) than Ai2’s (how to make each team whole), but on a cluster where the binding constraint is researcher time rather than allocation equity, they are the comparison class.

Mapping this onto your stack

If your cluster runs Slurm, which, per a paper on bridging HPC clusters with big-data technologies citing Slurm’s lead architect, handled resource management for approximately 60% of the Top 500 HPC systems in 2023, the lineage is encouraging: per Ai2’s post, the same approach is “in active use today in SLURM’s Fair Tree and YARN’s Fair Scheduler.” Lineage is not a configuration, though. No cited source documents how to express Ai2’s model in any specific scheduler, and I flag any such mapping as inference, not measured fact: Ai2’s post names the fair-share lineage but does not discuss implementation equivalence in Slurm, Kubernetes, or any other system, and no cited source measures any of them against Ai2’s outcomes. What the post does establish are the two properties its numbers depend on: hierarchical weights over a sliding lookback window, and per-workload protected minimum runtimes. Before committing, verify in current vendor and project docs that your queueing layer supports both; those two properties, not the brand of scheduler, are what Ai2’s results rest on. None of that means configuration is trivial: Ai2’s contribution is the budget governance layer (who sets weights, how often, with what appeals process), not the sorting algorithm.

Decision guide

Your situationRecommendationWhy, and what breaks
Research training cluster, demand 2-3x supply, 88 GPUs and up, priority queue plus ops favors todayPilot budgets + hierarchical fair-share + time-slicing contractAi2’s demand profile matches yours; the reported 98% delivery of owed hours and flat 98% occupancy are the evidence. Success criteria: fairness delivery and wait times, not throughput.
Fleet dominated by a few large gang-scheduled runsKeep long protected, effectively non-preemptible windows; adopt budgets for the remainderPreemption costs seconds-to-minutes per event (arXiv:2501.05563), and minimum-runtime protection may already be impeding large-job placement in Ai2’s own fleet.
Heavy interactive/debug usageAdopt, but build session checkpoint/restore firstThe 8-hour cap with manual state rebuild was Ai2’s clearest self-reported failure.
Cluster smaller than Ai2’s smallest, or mixed training-plus-inference fleetDo not assume transfer; the evidence is silentPer its post, Ai2’s smallest cluster is 88 GPUs and the fleet is training-focused. No outcome claim covers your case.
Binding constraint is job completion time, not allocation equityEvaluate goodput-oriented schedulers (Pollux class) alongside or insteadPollux’s 30% HPO speedup measures a different objective (arXiv:2008.12260); fairness delivery and goodput optimization solve different problems.

One more condition cuts across all rows: the administrative layer has to exist. Ai2’s model works because managers with context set budgets on a regular cadence, with escalation up to the CEO. If your organization cannot staff that budgeting process, a fair-share scheduler becomes a priority queue with extra steps.

Limits of the evidence

Everything measured here is Ai2 measuring Ai2, per its post. The 98% delivery, the 90% worst case, the 2-hours-to-30-seconds debug p90 (with Ai2’s own small-sample variance caveat), and the flat 98% occupancy, all per its post, are one organization’s 30-day window, with no independent replication and no measurement of whether research output improved. The impact layer of Ai2’s own pyramid, whether the most valuable workloads got the GPUs, is asserted by the budgeting process rather than measured against outcomes. Transfer is unproven on clusters smaller than Ai2’s smallest, on inference-heavy fleets, and to demand mixes unlike Ai2’s. And 30 days is one lookback cycle times four; seasonal effects, a major deadline crunch, or a flagship run that monopolizes a cluster could all stress the system in ways this window did not.

My read: for a shared research cluster oversubscribed two to three times, the pilot is justified. The mechanisms are old and implementable, the reported fairness delivery is strong, and the failure modes are documented well enough to defend against. I would not roll the time-slicing contract out to interactive workloads until session checkpoint and restore exists, and I would keep large gang-scheduled training runs behind long protected windows, because the preemption overhead and the placement-starvation risk are the parts of this design the literature and Ai2’s own follow-up investigation both flag. If a lost interactive session means a researcher rebuilding a day’s state by hand, that fix comes before the fairness gains, not after.

Frequently Asked Questions

What specific mechanisms did Ai2 use to replace its priority-based scheduler?

We recently replaced a priority-based scheduler with a system including GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract. As a result, we shifted the debate about how much GPU time each research project deserves from a case-by-case operational task to a transparent administrative budgeting process.

How much GPU time did teams actually receive under the new system?

Over the 30-day test period, teams were delivered 98% of the GPU hours they were owed, and 13 of 15 team allocations received 95% or more with the worst case receiving 90%.

What was the main operational failure mode identified with the new time-slicing contract?

In the old system, a researcher could hold such a session for up to a week. With time-slicing, they were subject to the 8-hour cap on protected runtime, after which a session becomes preemptible if it exceeds its allocation. We did not appreciate the extent to which researchers were dependent on the volatile state of these sessions. Being preempted meant waiting to secure a new session and also rebuilding their state by hand.

References

Follow the links in the article for context. The supporting material is collected here for further reading.

  1. Impactful Schedulinghuggingface.coAccessed
  2. Dynamic Multi-Objective GPU Cluster Schedulingarxiv.orgAccessed
  3. Prediction-Assisted DDL Schedulingarxiv.orgAccessed
  4. Topology-Aware Placement Researcharxiv.orgAccessed
  5. ONES: Elastic Batch-Size Scalingarxiv.orgAccessed
  6. Bridging HPC Clusters with Big-Data Technologiesarxiv.orgAccessed

Join the discussion

Share a useful perspective or ask a question about this article.

Discussion guidelinesComments privacy