Kubernetes has spent years getting very good at answering one fundamental question:
Where should this Pod run?
For a large class of traditional applications, that is exactly the right question.
A stateless Deployment with ten API replicas usually does not care if replica seven starts 20 seconds after replica six. Each Pod makes progress independently. The scheduler evaluates them one at a time, finds a feasible Node, binds the Pod, and moves to the next item in its queue.
That model is simple, scalable, and foundational to how Kubernetes operates.
However, increasingly in modern infrastructure, the entity that needs to make progress is not an individual Pod.
It is a workload composed of multiple Pods that must be placed together under shared constraints around accelerators, network topology, priority, and disruption boundaries.
A distributed AI training job illustrates the fundamental failure mode of Pod-at-a-time scheduling:
flowchart TD
subgraph Job["Distributed Training Job (Requires 32 GPUs / 4 Workers)"]
W1["Worker 1 (8x GPU)<br/>Status: Scheduled"]
W2["Worker 2 (8x GPU)<br/>Status: Scheduled"]
W3["Worker 3 (8x GPU)<br/>Status: Scheduled"]
W4["Worker 4 (8x GPU)<br/>Status: Pending (Insufficient Capacity)"]
end
W1 --- State1["24 GPUs Bound & Holding Memory"]
W2 --- State1
W3 --- State1
State1 --> Result["Result: 0% Application Progress<br/>24 GPUs Locked in Idle State"]
Imagine Kubernetes schedules three workers successfully, but the fourth worker cannot fit on any remaining Node.
From the scheduler’s perspective, it has made three successful Pod placement decisions.
From the application’s perspective, it has made zero progress.
Worse, 24 high-performance GPUs are now locked by a training job that cannot start, while simultaneously blocking other pending jobs from utilizing those accelerators.
The Pods are scheduled. The workload is not.
That distinction is important enough that Kubernetes itself is evolving its core scheduling architecture.
Kubernetes v1.37 promotes the Workload and PodGroup APIs to Beta, introduces workload-aware preemption, and enables the scheduler to treat a PodGroup as an atomic scheduling unit alongside standalone Pods.123
The Pod remains the execution unit. It is simply no longer always the right scheduling unit.
The Original Scheduler Model Is Pod-Centric
The Kubernetes documentation defines scheduling as matching individual Pods to Nodes so that kubelet can execute them.4
At a structural level, the scheduling framework operates as a pipeline centered around a single Pod:
flowchart TD
P["Incoming Pod"] --> Q["Scheduling Queue"]
Q --> PF["PreFilter / Filter Nodes"]
PF --> S["Score Feasible Nodes"]
S --> R["Reserve & Permit"]
R --> B["Bind Pod to Selected Node"]
The scheduling framework phases—PreFilter, Filter, Score, Reserve, Permit, and Bind—are executed sequentially for one Pod at a time.5
That model is optimal when workloads consist of fungible replicas:
flowchart LR
subgraph Web["Stateless Deployment (Fungible Replicas)"]
direction TB
W1["Pod 1"] -->|Bound| NA["Node A"]
W2["Pod 2"] -->|Bound| NC["Node C"]
W3["Pod 3"] -->|Bound| NB["Node B"]
W4["Pod 4"] -->|Pending| Q["Queue"]
end
Web --- Res["Application Status: Operational at 75% Capacity"]
If replica four is delayed or remains pending due to cluster capacity, the application continues to serve traffic. Replica four arriving later does not impede replicas one through three.
Distributed compute workloads exhibit fundamentally different runtime semantics.
A Collection of Pods Is Not a Runnable Workload
8 workers × 8 GPUs = 64 GPUs
A Pod-centric scheduler sees eight discrete scheduling items in its queue. The application sees a single atomic job:
flowchart LR
subgraph Scheduler["Kubernetes Scheduler View"]
direction TB
P1["Pod 1"]
P2["Pod 2"]
P3["Pod 3"]
P4["Pod 4"]
P5["Pod 5"]
P6["Pod 6"]
P7["Pod 7"]
P8["Pod 8"]
end
subgraph Application["Application View"]
direction TB
JOB["Single Distributed Training Job<br/>(64 GPUs Required)"]
end
If the cluster only has capacity for six workers, a Pod-at-a-time scheduler reaches a deadlocked state:
worker-1 Running (Allocated 8 GPUs)
worker-2 Running (Allocated 8 GPUs)
worker-3 Running (Allocated 8 GPUs)
worker-4 Running (Allocated 8 GPUs)
worker-5 Running (Allocated 8 GPUs)
worker-6 Running (Allocated 8 GPUs)
worker-7 Pending (Waiting for Capacity)
worker-8 Pending (Waiting for Capacity)
From an infrastructure accounting perspective, 75% of requested Pods are placed. From the job’s perspective, 0% of compute execution occurs.
This creates workload-level resource fragmentation. The six running workers hold 48 GPUs in an idle state, preventing smaller incoming jobs from running while failing to advance the primary job.
This failure mode is the primary driver behind Kubernetes’ workload-aware scheduling initiative.6
Gang Scheduling Changes the Scheduler Question
Gang scheduling changes the scheduler’s objective function from:
Can I schedule this Pod?
to:
Can I schedule enough of this group for the workload to make progress?
Instead of evaluating and binding workers independently, the scheduler evaluates placement for the gang as a single unit:
flowchart TD
subgraph Individual["Pod-Centric Evaluation (Independent)"]
direction TB
PA["Pod A"] -->|Bind| N1["Node 1"]
PB["Pod B"] -->|Bind| N2["Node 2"]
PC["Pod C"] -->|Bind| N3["Node 3"]
PD["Pod D"] -->|Fail| PEND["Pending"]
N1 & N2 & N3 --- Partial["Partial Placement: Deadlocked"]
end
subgraph Gang["Gang Scheduling Evaluation (Atomic Snapshot)"]
direction TB
PG["PodGroup (minCount: 4)"] --> EVAL{"Evaluate Cluster Snapshot<br/>Can minCount fit?"}
EVAL -->|Yes| BIND_ALL["Bind Pods A, B, C, D Simultaneously"]
EVAL -->|No| REJECT_ALL["Keep Entire Group Pending<br/>(Zero Partial Allocations)"]
end
If the minimum required count (minCount) cannot be satisfied across the cluster, none of the Pods in that group are bound.
Kubernetes natively models this via the PodGroup API (scheduling.k8s.io/v1beta1). A PodGroup represents a set of Pods scheduled together as an atomic unit.2
apiVersion: scheduling.k8s.io/v1beta1
kind: PodGroup
metadata:
name: training-workers
spec:
schedulingPolicy:
gang:
minCount: 4
Kubernetes v1.36 introduced a dedicated PodGroup scheduling cycle. The scheduler captures a single cluster snapshot, evaluates placements for the entire group, and applies the scheduling outcome atomically.6
If the group’s constraints cannot be satisfied, the group is marked unschedulable and returned to the queue without binding partial Pods.6
Kubernetes Has a Workload API for Scheduling Intent
Kubernetes v1.37 promotes the Workload API to Beta. The architecture distinguishes between long-lived policy and runtime state:
flowchart TD
W["Workload API (v1beta1)<br/>Defines application scheduling intent & podGroupTemplates"] -->|Instantiated by Controller| PG["PodGroup API (v1beta1)<br/>Models runtime gang constraints & active scheduling unit"]
PG -->|Manages| P1["Pod 1"]
PG -->|Manages| P2["Pod 2"]
PG -->|Manages| PN["Pod N"]
The Workload resource defines the macro scheduling structure of a multi-Pod application, whereas PodGroup represents an active runtime instance evaluated by the scheduler.72
apiVersion: scheduling.k8s.io/v1beta1
kind: Workload
metadata:
name: training-job
spec:
podGroupTemplates:
- name: workers
schedulingPolicy:
gang:
minCount: 4
Controllers (such as the built-in Job controller) generate runtime PodGroup objects from these templates.7 This provides an explicit API declaration that a collection of Pods constitutes a single scheduling decision rather than isolated queue entries.
The Scheduler Queue Structure
With workload-aware scheduling active, PodGroup objects are queued alongside standalone Pods.3
Under the experimental CompositePodGroup feature gate, the scheduler queue handles multiple distinct primitives:
flowchart LR
subgraph Legacy["Traditional Queue Structure"]
direction TB
T1["Pod A"]
T2["Pod B"]
T3["Pod C"]
T4["Pod D"]
end
subgraph WorkloadAware["Workload-Aware Queue Structure"]
direction TB
W1["Standalone Pod"]
W2["PodGroup (Gang Unit)"]
W3["Standalone Pod"]
W4["CompositePodGroup (Hierarchical Unit)"]
end
The queue item is no longer constrained to a single Pod object. The scheduler accepts larger composite structures when Pod-level queueing fails to represent application semantics.
Preemption at the Workload Level
Traditional Kubernetes preemption functions on a single-Pod basis. When a high-priority Pod cannot be scheduled, the scheduler identifies lower-priority Pods on a single Node and evicts them to reclaim capacity.3
For a gang-scheduled workload, single-Node Pod preemption is often inefficient or ineffective:
flowchart TD
subgraph SinglePod["Traditional Pod Preemption"]
direction TB
PP_POD["Pending Pod"] --> PP_SEARCH["Search Single Node"]
PP_SEARCH --> PP_EVICT["Evict Victims on Node"]
end
subgraph WorkloadPreemption["Workload-Aware Preemption"]
direction TB
WP_GROUP["Pending PodGroup"] --> WP_SEARCH["Search Cluster-Wide Capacity"]
WP_SEARCH --> WP_EVICT["Evict Victims Across Nodes to Satisfy Full minCount"]
end
If a training job requires eight workers across four Nodes, preempting Pods on a single Node does not make the overall workload runnable.
Kubernetes v1.37 promotes workload-aware preemption to Beta.89 Workload-aware preemption evaluates capacity cluster-wide, selecting eviction candidates that satisfy the total resource requirements of the pending PodGroup rather than isolated Nodes.9
Topology Constraints Across Groups
Resource quantities (CPU, Memory, GPU counts) do not fully capture placement requirements. For tightly coupled distributed workloads, physical network topology dictates execution performance.
Consider two placement strategies for four GPU workers:
flowchart TD
subgraph Suboptimal["Placement A: Fragmented Topology (Cross-Zone)"]
direction TB
subgraph ZA["Zone A"]
R1["Rack 1: Worker 1"]
R7["Rack 7: Worker 2"]
end
subgraph ZB["Zone B"]
R3["Rack 3: Worker 3"]
R8["Rack 8: Worker 4"]
end
R1 -.-|High Latency / Cross-Zone Traffic| R3
end
subgraph Optimal["Placement B: Topologically Colocated (Same Rack)"]
direction TB
subgraph ZA2["Zone A"]
subgraph R12["Rack 1"]
W1["Worker 1"]
W2["Worker 2"]
W3["Worker 3"]
W4["Worker 4"]
end
end
W1 <===>|High Bandwidth / NVLink Interconnect| W2
end
Both placements satisfy raw resource requests (4 Pods running, 32 GPUs allocated). However, Placement A introduces cross-zone latencies that severely bottleneck gradient all-reduce steps in distributed training.
Workload-aware scheduling introduces topology constraints directly at the PodGroup level.1 This allows workloads to enforce placement boundaries (e.g., requiring all workers within a PodGroup to land on Nodes sharing a common rack or NVLink domain) during the atomic scheduling cycle.
Composite Workload Hierarchies
Complex distributed applications are rarely composed of uniform worker pools. A standard architecture includes distinct functional roles:
flowchart TD
W["Workload"] --> CPG["CompositePodGroup (Root Schedule Unit)"]
CPG --> PG1["PodGroup: Coordinator<br/>(minCount: 1)"]
CPG --> PG2["PodGroup: GPU Workers<br/>(minCount: 4, Rack Colocation Required)"]
CPG --> PG3["PodGroup: Data Loaders<br/>(minCount: 2)"]
PG1 --> P_C["Coordinator Pod"]
PG2 --> P_G1["GPU Worker 1"]
PG2 --> P_G2["GPU Worker 2"]
PG2 --> P_G3["GPU Worker 3"]
PG2 --> P_G4["GPU Worker 4"]
PG3 --> P_D1["Data Loader 1"]
PG3 --> P_D2["Data Loader 2"]
Each subgroup requires different scheduling semantics:
- The Coordinator requires 1 replica.
- The GPU Workers require a gang
minCountof 4 with rack-level colocation. - The Data Loaders require 2 replicas.
Kubernetes v1.37 introduces the CompositePodGroup primitive to model hierarchical group relationships.21 This allows multi-level gang constraints and topology boundaries to be evaluated together.
Accelerator Integration via Dynamic Resource Allocation (DRA)
Stateless microservices treat CPU and memory as fungible scalar quantities. High-performance accelerators require fine-grained hardware topology modeling:
- Accelerator model and memory capacity
- NVLink interconnect topology
- NUMA node affinity
- RDMA network interface availability
- Multi-Instance GPU (MIG) slice partitioning
flowchart TD
subgraph Legacy["Legacy Extended Resource Path"]
L_POD["Pod"] -->|requests scalar count| L_RES["example.com/gpu: 8"]
end
subgraph Modern["Workload-Aware DRA Architecture (v1.37)"]
M_WORKLOAD["Workload / PodGroup"] --> M_CLAIM["ResourceClaim (Coordinated Device Scope)"]
M_CLAIM --> M_DEV["Hardware Topology: NVLink, NUMA, RDMA"]
M_WORKLOAD --> M_PODS["Managed Pod Group"]
end
In Kubernetes v1.37, Dynamic Resource Allocation (DRA) extended-resource support reached Stable, enabling DRA drivers to satisfy standard resource requests without separate device plugins.10
Concurrently, ResourceClaim integration with Workload and PodGroup APIs reached Beta.8 Devices are claimed and allocated at the workload level rather than forcing each Pod to negotiate hardware topology independently.
Architectural Comparison Matrix
Scheduling Paradigms
| Architectural Dimension | Pod-Centric Scheduling | Workload-Aware Scheduling |
|---|---|---|
| Primary Unit | Pod | PodGroup / Workload / CompositePodGroup |
| Placement Evaluation | Sequential, Pod-by-Pod | Atomic cycle across group snapshot |
| Partial Binding Behavior | Permitted (e.g., 6 of 8 Pods bound) | Blocked until minCount is satisfied |
| Resource Fragmentation | High (idle GPUs held by partially bound jobs) | Eliminated at group boundary |
| Preemption Scope | Single-Node victim eviction | Cluster-wide victim selection for full group capacity |
| Topology Enforcement | Individual Pod affinity / anti-affinity | Group-level placement boundaries (Rack, Switch, Zone) |
| Target Workloads | Stateless APIs, web servers, queue workers | AI/ML training, HPC, distributed inference, batch jobs |
Kubernetes v1.37 Scheduling Primitives
| Abstraction | API Group & Status | Functional Responsibility |
|---|---|---|
Pod | core/v1 (GA) | Container execution unit on a single Node; namespace boundary. |
PodGroup | scheduling.k8s.io/v1beta1 (Beta) | Runtime gang scheduling unit enforcing minCount and atomic binding. |
Workload | scheduling.k8s.io/v1beta1 (Beta) | High-level declaration of multi-Pod scheduling intent and policies. |
CompositePodGroup | scheduling.k8s.io/v1beta1 (Alpha) | Hierarchical tree of PodGroup objects for multi-role job topologies. |
ResourceClaim | resource.k8s.io/v1beta1 (Beta) | Coordinated hardware allocation (GPUs, RDMA) bound at group scope. |
Execution Unit vs Scheduling Unit
This shift does not eliminate the Pod. Pods remain the execution primitive within Kubernetes:
- Created by workload controllers
- Assigned to individual Nodes
- Executed by
kubeletand container runtimes - The boundary for shared namespaces, storage volumes, and cgroups
The core distinction is between the execution unit and the scheduling unit:
Stateless Microservice:
Execution Unit = Pod
Scheduling Unit = Pod
Distributed Batch Workload:
Execution Unit = Pod
Scheduling Unit = PodGroup / Workload
For a 256-GPU training job, Pods remain the container runtime execution boundary on each Node. However, evaluating placement for those Pods independently creates broken operational states. The Pod becomes an execution detail inside a larger scheduling decision.
Architectural Direction
Kubernetes scheduling is shifting from isolated resource placement to holistic workload placement:
flowchart LR
Container["Container"] --> Pod["Pod<br/>(Execution Unit)"]
Pod --> Controller["Controller<br/>(Deployment/Job)"]
Controller --> PodGroup["PodGroup<br/>(Gang Unit)"]
PodGroup --> Workload["Workload API<br/>(Intent Declaration)"]
Workload --> Composite["Composite Workload<br/>(Hierarchical Topology)"]
Traditional scheduling evaluated whether a specific Node possessed sufficient unallocated capacity for a single Pod.
Workload-aware scheduling evaluates whether the cluster can construct a valid placement topology under which an entire application can make execution progress.
Kubernetes v1.37 provides the workloadbuilder library and building block APIs so controller developers can standardize gang scheduling, topology, and disruption primitives across custom CRDs.11
Stateless web services with interchangeable replicas will continue to rely on standard Pod-level scheduling. However, for AI infrastructure, HPC, and tightly coupled systems:
A Pod being schedulable does not mean the workload is runnable.
When that condition holds, scheduling at the individual Pod level is insufficient. The scheduler must reason about the unit of progress—and that unit is the workload.
References
Kubernetes Blog, “Kubernetes v1.37: Advancing Workload-Aware Scheduling”, September 8, 2026. ↩︎ ↩︎ ↩︎
Kubernetes Documentation, “PodGroup API”. The feature is Beta in Kubernetes v1.37 and enabled via the
GenericWorkloadfeature gate. ↩︎ ↩︎ ↩︎ ↩︎Kubernetes Documentation, “Pod Priority and Preemption”. ↩︎ ↩︎ ↩︎
Kubernetes Documentation, “Scheduling, Preemption and Eviction”. ↩︎
Kubernetes Documentation, “Scheduling Framework”. ↩︎
Kubernetes Blog, “Kubernetes v1.36: Advancing Workload-Aware Scheduling”, May 13, 2026. ↩︎ ↩︎ ↩︎
Kubernetes Documentation, “Workload API”. The feature is Beta in Kubernetes v1.37 and enabled via the
GenericWorkloadfeature gate. ↩︎ ↩︎Kubernetes Blog, “Kubernetes v1.37: Garhwal”, August 26, 2026. ↩︎ ↩︎
Kubernetes Documentation, “Workload-Aware Preemption”. ↩︎ ↩︎
Kubernetes Blog, “Kubernetes v1.37: DRA Updates”, September 3, 2026. ↩︎
Kubernetes Documentation, “Scheduling Building Block APIs and the workloadbuilder Library”. ↩︎