Kubernetes has spent years getting very good at answering one fundamental question:

Where should this Pod run?

For a large class of traditional applications, that is exactly the right question.

A stateless Deployment with ten API replicas usually does not care if replica seven starts 20 seconds after replica six. Each Pod makes progress independently. The scheduler evaluates them one at a time, finds a feasible Node, binds the Pod, and moves to the next item in its queue.

That model is simple, scalable, and foundational to how Kubernetes operates.

However, increasingly in modern infrastructure, the entity that needs to make progress is not an individual Pod.

It is a workload composed of multiple Pods that must be placed together under shared constraints around accelerators, network topology, priority, and disruption boundaries.

A distributed AI training job illustrates the fundamental failure mode of Pod-at-a-time scheduling:

flowchart TD
    subgraph Job["Distributed Training Job (Requires 32 GPUs / 4 Workers)"]
        W1["Worker 1 (8x GPU)<br/>Status: Scheduled"]
        W2["Worker 2 (8x GPU)<br/>Status: Scheduled"]
        W3["Worker 3 (8x GPU)<br/>Status: Scheduled"]
        W4["Worker 4 (8x GPU)<br/>Status: Pending (Insufficient Capacity)"]
    end
    
    W1 --- State1["24 GPUs Bound & Holding Memory"]
    W2 --- State1
    W3 --- State1
    
    State1 --> Result["Result: 0% Application Progress<br/>24 GPUs Locked in Idle State"]

Imagine Kubernetes schedules three workers successfully, but the fourth worker cannot fit on any remaining Node.

From the scheduler’s perspective, it has made three successful Pod placement decisions.

From the application’s perspective, it has made zero progress.

Worse, 24 high-performance GPUs are now locked by a training job that cannot start, while simultaneously blocking other pending jobs from utilizing those accelerators.

The Pods are scheduled. The workload is not.

That distinction is important enough that Kubernetes itself is evolving its core scheduling architecture.

Kubernetes v1.37 promotes the Workload and PodGroup APIs to Beta, introduces workload-aware preemption, and enables the scheduler to treat a PodGroup as an atomic scheduling unit alongside standalone Pods.123

The Pod remains the execution unit. It is simply no longer always the right scheduling unit.


The Original Scheduler Model Is Pod-Centric

The Kubernetes documentation defines scheduling as matching individual Pods to Nodes so that kubelet can execute them.4

At a structural level, the scheduling framework operates as a pipeline centered around a single Pod:

flowchart TD
    P["Incoming Pod"] --> Q["Scheduling Queue"]
    Q --> PF["PreFilter / Filter Nodes"]
    PF --> S["Score Feasible Nodes"]
    S --> R["Reserve & Permit"]
    R --> B["Bind Pod to Selected Node"]

The scheduling framework phases—PreFilter, Filter, Score, Reserve, Permit, and Bind—are executed sequentially for one Pod at a time.5

That model is optimal when workloads consist of fungible replicas:

flowchart LR
    subgraph Web["Stateless Deployment (Fungible Replicas)"]
        direction TB
        W1["Pod 1"] -->|Bound| NA["Node A"]
        W2["Pod 2"] -->|Bound| NC["Node C"]
        W3["Pod 3"] -->|Bound| NB["Node B"]
        W4["Pod 4"] -->|Pending| Q["Queue"]
    end
    
    Web --- Res["Application Status: Operational at 75% Capacity"]

If replica four is delayed or remains pending due to cluster capacity, the application continues to serve traffic. Replica four arriving later does not impede replicas one through three.

Distributed compute workloads exhibit fundamentally different runtime semantics.


A Collection of Pods Is Not a Runnable Workload

8 workers × 8 GPUs = 64 GPUs

A Pod-centric scheduler sees eight discrete scheduling items in its queue. The application sees a single atomic job:

flowchart LR
    subgraph Scheduler["Kubernetes Scheduler View"]
        direction TB
        P1["Pod 1"]
        P2["Pod 2"]
        P3["Pod 3"]
        P4["Pod 4"]
        P5["Pod 5"]
        P6["Pod 6"]
        P7["Pod 7"]
        P8["Pod 8"]
    end

    subgraph Application["Application View"]
        direction TB
        JOB["Single Distributed Training Job<br/>(64 GPUs Required)"]
    end

If the cluster only has capacity for six workers, a Pod-at-a-time scheduler reaches a deadlocked state:

worker-1    Running (Allocated 8 GPUs)
worker-2    Running (Allocated 8 GPUs)
worker-3    Running (Allocated 8 GPUs)
worker-4    Running (Allocated 8 GPUs)
worker-5    Running (Allocated 8 GPUs)
worker-6    Running (Allocated 8 GPUs)
worker-7    Pending (Waiting for Capacity)
worker-8    Pending (Waiting for Capacity)

From an infrastructure accounting perspective, 75% of requested Pods are placed. From the job’s perspective, 0% of compute execution occurs.

This creates workload-level resource fragmentation. The six running workers hold 48 GPUs in an idle state, preventing smaller incoming jobs from running while failing to advance the primary job.

This failure mode is the primary driver behind Kubernetes’ workload-aware scheduling initiative.6


Gang Scheduling Changes the Scheduler Question

Gang scheduling changes the scheduler’s objective function from:

Can I schedule this Pod?

to:

Can I schedule enough of this group for the workload to make progress?

Instead of evaluating and binding workers independently, the scheduler evaluates placement for the gang as a single unit:

flowchart TD
    subgraph Individual["Pod-Centric Evaluation (Independent)"]
        direction TB
        PA["Pod A"] -->|Bind| N1["Node 1"]
        PB["Pod B"] -->|Bind| N2["Node 2"]
        PC["Pod C"] -->|Bind| N3["Node 3"]
        PD["Pod D"] -->|Fail| PEND["Pending"]
        N1 & N2 & N3 --- Partial["Partial Placement: Deadlocked"]
    end

    subgraph Gang["Gang Scheduling Evaluation (Atomic Snapshot)"]
        direction TB
        PG["PodGroup (minCount: 4)"] --> EVAL{"Evaluate Cluster Snapshot<br/>Can minCount fit?"}
        EVAL -->|Yes| BIND_ALL["Bind Pods A, B, C, D Simultaneously"]
        EVAL -->|No| REJECT_ALL["Keep Entire Group Pending<br/>(Zero Partial Allocations)"]
    end

If the minimum required count (minCount) cannot be satisfied across the cluster, none of the Pods in that group are bound.

Kubernetes natively models this via the PodGroup API (scheduling.k8s.io/v1beta1). A PodGroup represents a set of Pods scheduled together as an atomic unit.2

apiVersion: scheduling.k8s.io/v1beta1
kind: PodGroup
metadata:
  name: training-workers
spec:
  schedulingPolicy:
    gang:
      minCount: 4

Kubernetes v1.36 introduced a dedicated PodGroup scheduling cycle. The scheduler captures a single cluster snapshot, evaluates placements for the entire group, and applies the scheduling outcome atomically.6

If the group’s constraints cannot be satisfied, the group is marked unschedulable and returned to the queue without binding partial Pods.6


Kubernetes Has a Workload API for Scheduling Intent

Kubernetes v1.37 promotes the Workload API to Beta. The architecture distinguishes between long-lived policy and runtime state:

flowchart TD
    W["Workload API (v1beta1)<br/>Defines application scheduling intent & podGroupTemplates"] -->|Instantiated by Controller| PG["PodGroup API (v1beta1)<br/>Models runtime gang constraints & active scheduling unit"]
    PG -->|Manages| P1["Pod 1"]
    PG -->|Manages| P2["Pod 2"]
    PG -->|Manages| PN["Pod N"]

The Workload resource defines the macro scheduling structure of a multi-Pod application, whereas PodGroup represents an active runtime instance evaluated by the scheduler.72

apiVersion: scheduling.k8s.io/v1beta1
kind: Workload
metadata:
  name: training-job
spec:
  podGroupTemplates:
    - name: workers
      schedulingPolicy:
        gang:
          minCount: 4

Controllers (such as the built-in Job controller) generate runtime PodGroup objects from these templates.7 This provides an explicit API declaration that a collection of Pods constitutes a single scheduling decision rather than isolated queue entries.


The Scheduler Queue Structure

With workload-aware scheduling active, PodGroup objects are queued alongside standalone Pods.3

Under the experimental CompositePodGroup feature gate, the scheduler queue handles multiple distinct primitives:

flowchart LR
    subgraph Legacy["Traditional Queue Structure"]
        direction TB
        T1["Pod A"]
        T2["Pod B"]
        T3["Pod C"]
        T4["Pod D"]
    end

    subgraph WorkloadAware["Workload-Aware Queue Structure"]
        direction TB
        W1["Standalone Pod"]
        W2["PodGroup (Gang Unit)"]
        W3["Standalone Pod"]
        W4["CompositePodGroup (Hierarchical Unit)"]
    end

The queue item is no longer constrained to a single Pod object. The scheduler accepts larger composite structures when Pod-level queueing fails to represent application semantics.


Preemption at the Workload Level

Traditional Kubernetes preemption functions on a single-Pod basis. When a high-priority Pod cannot be scheduled, the scheduler identifies lower-priority Pods on a single Node and evicts them to reclaim capacity.3

For a gang-scheduled workload, single-Node Pod preemption is often inefficient or ineffective:

flowchart TD
    subgraph SinglePod["Traditional Pod Preemption"]
        direction TB
        PP_POD["Pending Pod"] --> PP_SEARCH["Search Single Node"]
        PP_SEARCH --> PP_EVICT["Evict Victims on Node"]
    end

    subgraph WorkloadPreemption["Workload-Aware Preemption"]
        direction TB
        WP_GROUP["Pending PodGroup"] --> WP_SEARCH["Search Cluster-Wide Capacity"]
        WP_SEARCH --> WP_EVICT["Evict Victims Across Nodes to Satisfy Full minCount"]
    end

If a training job requires eight workers across four Nodes, preempting Pods on a single Node does not make the overall workload runnable.

Kubernetes v1.37 promotes workload-aware preemption to Beta.89 Workload-aware preemption evaluates capacity cluster-wide, selecting eviction candidates that satisfy the total resource requirements of the pending PodGroup rather than isolated Nodes.9


Topology Constraints Across Groups

Resource quantities (CPU, Memory, GPU counts) do not fully capture placement requirements. For tightly coupled distributed workloads, physical network topology dictates execution performance.

Consider two placement strategies for four GPU workers:

flowchart TD
    subgraph Suboptimal["Placement A: Fragmented Topology (Cross-Zone)"]
        direction TB
        subgraph ZA["Zone A"]
            R1["Rack 1: Worker 1"]
            R7["Rack 7: Worker 2"]
        end
        subgraph ZB["Zone B"]
            R3["Rack 3: Worker 3"]
            R8["Rack 8: Worker 4"]
        end
        R1 -.-|High Latency / Cross-Zone Traffic| R3
    end

    subgraph Optimal["Placement B: Topologically Colocated (Same Rack)"]
        direction TB
        subgraph ZA2["Zone A"]
            subgraph R12["Rack 1"]
                W1["Worker 1"]
                W2["Worker 2"]
                W3["Worker 3"]
                W4["Worker 4"]
            end
        end
        W1 <===>|High Bandwidth / NVLink Interconnect| W2
    end

Both placements satisfy raw resource requests (4 Pods running, 32 GPUs allocated). However, Placement A introduces cross-zone latencies that severely bottleneck gradient all-reduce steps in distributed training.

Workload-aware scheduling introduces topology constraints directly at the PodGroup level.1 This allows workloads to enforce placement boundaries (e.g., requiring all workers within a PodGroup to land on Nodes sharing a common rack or NVLink domain) during the atomic scheduling cycle.


Composite Workload Hierarchies

Complex distributed applications are rarely composed of uniform worker pools. A standard architecture includes distinct functional roles:

flowchart TD
    W["Workload"] --> CPG["CompositePodGroup (Root Schedule Unit)"]
    CPG --> PG1["PodGroup: Coordinator<br/>(minCount: 1)"]
    CPG --> PG2["PodGroup: GPU Workers<br/>(minCount: 4, Rack Colocation Required)"]
    CPG --> PG3["PodGroup: Data Loaders<br/>(minCount: 2)"]
    
    PG1 --> P_C["Coordinator Pod"]
    PG2 --> P_G1["GPU Worker 1"]
    PG2 --> P_G2["GPU Worker 2"]
    PG2 --> P_G3["GPU Worker 3"]
    PG2 --> P_G4["GPU Worker 4"]
    PG3 --> P_D1["Data Loader 1"]
    PG3 --> P_D2["Data Loader 2"]

Each subgroup requires different scheduling semantics:

  • The Coordinator requires 1 replica.
  • The GPU Workers require a gang minCount of 4 with rack-level colocation.
  • The Data Loaders require 2 replicas.

Kubernetes v1.37 introduces the CompositePodGroup primitive to model hierarchical group relationships.21 This allows multi-level gang constraints and topology boundaries to be evaluated together.


Accelerator Integration via Dynamic Resource Allocation (DRA)

Stateless microservices treat CPU and memory as fungible scalar quantities. High-performance accelerators require fine-grained hardware topology modeling:

  • Accelerator model and memory capacity
  • NVLink interconnect topology
  • NUMA node affinity
  • RDMA network interface availability
  • Multi-Instance GPU (MIG) slice partitioning
flowchart TD
    subgraph Legacy["Legacy Extended Resource Path"]
        L_POD["Pod"] -->|requests scalar count| L_RES["example.com/gpu: 8"]
    end

    subgraph Modern["Workload-Aware DRA Architecture (v1.37)"]
        M_WORKLOAD["Workload / PodGroup"] --> M_CLAIM["ResourceClaim (Coordinated Device Scope)"]
        M_CLAIM --> M_DEV["Hardware Topology: NVLink, NUMA, RDMA"]
        M_WORKLOAD --> M_PODS["Managed Pod Group"]
    end

In Kubernetes v1.37, Dynamic Resource Allocation (DRA) extended-resource support reached Stable, enabling DRA drivers to satisfy standard resource requests without separate device plugins.10

Concurrently, ResourceClaim integration with Workload and PodGroup APIs reached Beta.8 Devices are claimed and allocated at the workload level rather than forcing each Pod to negotiate hardware topology independently.


Architectural Comparison Matrix

Scheduling Paradigms

Architectural DimensionPod-Centric SchedulingWorkload-Aware Scheduling
Primary UnitPodPodGroup / Workload / CompositePodGroup
Placement EvaluationSequential, Pod-by-PodAtomic cycle across group snapshot
Partial Binding BehaviorPermitted (e.g., 6 of 8 Pods bound)Blocked until minCount is satisfied
Resource FragmentationHigh (idle GPUs held by partially bound jobs)Eliminated at group boundary
Preemption ScopeSingle-Node victim evictionCluster-wide victim selection for full group capacity
Topology EnforcementIndividual Pod affinity / anti-affinityGroup-level placement boundaries (Rack, Switch, Zone)
Target WorkloadsStateless APIs, web servers, queue workersAI/ML training, HPC, distributed inference, batch jobs

Kubernetes v1.37 Scheduling Primitives

AbstractionAPI Group & StatusFunctional Responsibility
Podcore/v1 (GA)Container execution unit on a single Node; namespace boundary.
PodGroupscheduling.k8s.io/v1beta1 (Beta)Runtime gang scheduling unit enforcing minCount and atomic binding.
Workloadscheduling.k8s.io/v1beta1 (Beta)High-level declaration of multi-Pod scheduling intent and policies.
CompositePodGroupscheduling.k8s.io/v1beta1 (Alpha)Hierarchical tree of PodGroup objects for multi-role job topologies.
ResourceClaimresource.k8s.io/v1beta1 (Beta)Coordinated hardware allocation (GPUs, RDMA) bound at group scope.

Execution Unit vs Scheduling Unit

This shift does not eliminate the Pod. Pods remain the execution primitive within Kubernetes:

  • Created by workload controllers
  • Assigned to individual Nodes
  • Executed by kubelet and container runtimes
  • The boundary for shared namespaces, storage volumes, and cgroups

The core distinction is between the execution unit and the scheduling unit:

Stateless Microservice:
  Execution Unit  = Pod
  Scheduling Unit = Pod

Distributed Batch Workload:
  Execution Unit  = Pod
  Scheduling Unit = PodGroup / Workload

For a 256-GPU training job, Pods remain the container runtime execution boundary on each Node. However, evaluating placement for those Pods independently creates broken operational states. The Pod becomes an execution detail inside a larger scheduling decision.


Architectural Direction

Kubernetes scheduling is shifting from isolated resource placement to holistic workload placement:

flowchart LR
    Container["Container"] --> Pod["Pod<br/>(Execution Unit)"]
    Pod --> Controller["Controller<br/>(Deployment/Job)"]
    Controller --> PodGroup["PodGroup<br/>(Gang Unit)"]
    PodGroup --> Workload["Workload API<br/>(Intent Declaration)"]
    Workload --> Composite["Composite Workload<br/>(Hierarchical Topology)"]

Traditional scheduling evaluated whether a specific Node possessed sufficient unallocated capacity for a single Pod.

Workload-aware scheduling evaluates whether the cluster can construct a valid placement topology under which an entire application can make execution progress.

Kubernetes v1.37 provides the workloadbuilder library and building block APIs so controller developers can standardize gang scheduling, topology, and disruption primitives across custom CRDs.11

Stateless web services with interchangeable replicas will continue to rely on standard Pod-level scheduling. However, for AI infrastructure, HPC, and tightly coupled systems:

A Pod being schedulable does not mean the workload is runnable.

When that condition holds, scheduling at the individual Pod level is insufficient. The scheduler must reason about the unit of progress—and that unit is the workload.


References


  1. Kubernetes Blog, “Kubernetes v1.37: Advancing Workload-Aware Scheduling”, September 8, 2026. ↩︎ ↩︎ ↩︎

  2. Kubernetes Documentation, “PodGroup API”. The feature is Beta in Kubernetes v1.37 and enabled via the GenericWorkload feature gate. ↩︎ ↩︎ ↩︎ ↩︎

  3. Kubernetes Documentation, “Pod Priority and Preemption”. ↩︎ ↩︎ ↩︎

  4. Kubernetes Documentation, “Scheduling, Preemption and Eviction”. ↩︎

  5. Kubernetes Documentation, “Scheduling Framework”. ↩︎

  6. Kubernetes Blog, “Kubernetes v1.36: Advancing Workload-Aware Scheduling”, May 13, 2026. ↩︎ ↩︎ ↩︎

  7. Kubernetes Documentation, “Workload API”. The feature is Beta in Kubernetes v1.37 and enabled via the GenericWorkload feature gate. ↩︎ ↩︎

  8. Kubernetes Blog, “Kubernetes v1.37: Garhwal”, August 26, 2026. ↩︎ ↩︎

  9. Kubernetes Documentation, “Workload-Aware Preemption”. ↩︎ ↩︎

  10. Kubernetes Blog, “Kubernetes v1.37: DRA Updates”, September 3, 2026. ↩︎

  11. Kubernetes Documentation, “Scheduling Building Block APIs and the workloadbuilder Library”. ↩︎