Learning Journey: Part 02 of 16
Module Goal: Understand how core controllers and provider controllers coordinate asynchronously inside the management cluster via CRD references, status conditions, and provider contracts.
Target Spec: Cluster APIv1.14(v1beta2APIs).
Prerequisites: Part 01: Kubernetes Managing Kubernetes (Declarative Reconciliation & Management vs Workload mental model).
1. The Core Architecture: Asynchronous Controller Coordination
In Part 01, we established the fundamental concept behind Cluster API: a Kubernetes cluster can be expressed as declarative state and managed by controllers.
This raises a practical architectural question: How do multiple independent controllers orchestrate cloud infrastructure, operating system bootstrapping, and control plane quorum without tightly coupling to one another?
A Cluster object does not directly call the AWS EC2 API. A Machine object does not configure kubelet. Instead, Cluster API operates as a decoupled ecosystem of cooperating controllers connected strictly through Kubernetes API objects.
flowchart TB
U["Operator / GitOps / Platform API"]
subgraph MGMT["Management Cluster Control Plane"]
API["Kubernetes API Server\n(etcd State Store)"]
CORE["Cluster API Core Controller"]
BOOT["Bootstrap Provider Controller\n(e.g., CAPBK)"]
CP["Control Plane Provider Controller\n(e.g., KubeadmControlPlane)"]
INFRA["Infrastructure Provider Controller\n(e.g., CAPA / CAPZ / CAPV)"]
API <--> CORE
API <--> BOOT
API <--> CP
API <--> INFRA
end
CLOUD["Cloud Provider / Hypervisor APIs"]
WORKLOAD["Workload Cluster Nodes"]
U -->|kubectl apply| API
INFRA -->|Provision Infrastructure| CLOUD
CP -->|Manage CP Specs| API
CLOUD -->|Hosts| WORKLOAD
The Kubernetes API as a Message Bus
Notice the center of the architecture: The Kubernetes API Server is the sole coordination point.
Controllers do not invoke gRPC or HTTP endpoints on one another. Instead:
- Controller A creates or updates an API custom resource (CR).
- Controller B watches that resource via a Kubernetes Informer.
- Controller B reconciles its local responsibility and updates the resource’s
.statusfield. - Controller A observes the status change and proceeds to the next phase of reconciliation.
This event-driven, eventual-consistency model gives Cluster API resilience against transient network failures, rate limits, and controller restarts.
2. The Four Pillar Controllers & Responsibility Matrix
To maintain separation of concerns, Cluster API divides cluster lifecycle management among four primary controller roles:
flowchart TD
CL["Cluster Resource"] -->|infrastructureRef| IC["InfraCluster Resource\n(e.g., AWSCluster)"]
CL -->|controlPlaneRef| CP["ControlPlane Resource\n(e.g., KubeadmControlPlane)"]
CP -->|owns / creates| M["Machine Resource"]
M -->|infrastructureRef| IM["InfraMachine Resource\n(e.g., AWSMachine)"]
M -->|bootstrap.configRef| BC["BootstrapConfig Resource\n(e.g., KubeadmConfig)"]
Responsibility Breakdown
| Component Role | Target Responsibility | Key CRDs Handled |
|---|---|---|
| Core Provider | Generic lifecycle semantics, machine set scaling, node health checks | Cluster, Machine, MachineSet, MachineDeployment, MachineHealthCheck |
| Infrastructure Provider | Environment-specific compute, networking, subnets, and load balancers | AWSCluster / AWSMachine, AzureCluster / AzureMachine, VSphereCluster |
| Bootstrap Provider | Generating initialization secrets and cloud-init scripts for node join | KubeadmConfig, KubeadmConfigTemplate |
| Control Plane Provider | Managing etcd membership, control plane quorum, and rolling control plane upgrades | KubeadmControlPlane, KubeadmControlPlaneTemplate |
Why Separation Matters
- Core CAPI has zero knowledge of AWS EC2, Azure VMs, or vSphere.
- Infrastructure Providers have zero knowledge of how
kubeadmbootstraps a node. - Bootstrap Providers do not know or care whether the target OS runs on cloud VMs or physical bare metal servers.
3. Provider Contracts & Contractual Decoupling
How can Core Cluster API coordinate with dozens of different cloud providers without importing provider-specific Go packages or schemas?
Through Provider Contracts.
A Provider Contract defines the exact structure, annotation patterns, and status conditions a custom resource must implement to be recognized by Core CAPI controllers.
flowchart TD
Core["Cluster API Core Controllers"] -->|Enforces Provider Contracts| Contract["CRD Status & Schema Contract"]
Contract --> AWS["AWS Provider (CAPA)\nAWSCluster / AWSMachine"]
Contract --> Azure["Azure Provider (CAPZ)\nAzureCluster / AzureMachine"]
Contract --> VSphere["vSphere Provider (CAPV)\nVSphereCluster / VSphereMachine"]
Contract --> Metal["Bare Metal Provider (Metal3)\nMetal3Cluster / Metal3Machine"]
Contract Requirements (e.g., InfrastructureMachine Contract)
Any custom resource pointed to by a Machine’s spec.infrastructureRef must:
- Expose
status.ready(Boolean) to signal when the underlying compute instance is fully provisioned. - Expose
spec.providerIDto map the custom resource to the KubernetesNode.spec.providerID. - Support standard CAPI annotations (such as pause annotations and owner references).
As long as an infrastructure plugin satisfies this contract, Core CAPI can manage its lifecycle seamlessly.
4. End-to-End Creation Walkthrough (The 6 Reconciliation Steps)
Let’s trace what actually happens inside the management cluster when an operator applies a new workload cluster manifest:
sequenceDiagram
autonumber
participant Op as Operator / GitOps
participant API as Kube API Server
participant ClusterCtrl as Cluster Controller
participant InfraCtrl as Infra Provider Controller
participant CPCtrl as Control Plane Controller
participant BootCtrl as Bootstrap Controller
participant MachineCtrl as Machine Controller
Op->>API: 1. Apply Cluster + InfraCluster + ControlPlane manifests
API->>InfraCtrl: 2. Reconcile InfraCluster (Network, VPC, LB)
InfraCtrl-->>API: Update InfraCluster status.ready = true
API->>ClusterCtrl: 3. Cluster Controller marks InfrastructureReady
API->>CPCtrl: 4. Control Plane Controller creates CP Machines
API->>BootCtrl: 5. Reconcile KubeadmConfig -> Generate Bootstrap Data
BootCtrl-->>API: Write Secret containing cloud-init script
API->>MachineCtrl: 6. Reconcile Machine & trigger InfraMachine
MachineCtrl->>InfraCtrl: Provision Cloud Instance with Bootstrap Data
InfraCtrl-->>API: Node Boots, Registers -> Machine Marked Running
Deep Dive into the Steps:
- Manifest Submission: The user submits a
ClusterCR referencing anAWSClusterand aKubeadmControlPlane. - Infrastructure Reconciliation: The AWS Infrastructure Controller sees the
AWSClusterCR and provisions the VPC, Subnets, Security Groups, and Control Plane Load Balancer. Once completed, it setsstatus.ready = true. - Cluster Status Update: The Core Cluster Controller observes
status.ready == trueon theAWSClusterand updates theClusterresource’s status toInfrastructureReady = True. - Control Plane Initialization: The
KubeadmControlPlanecontroller detects that cluster infrastructure is ready and generates the first control planeMachineresource along with a correspondingKubeadmConfig. - Bootstrap Data Generation: The Kubeadm Bootstrap Controller watches the
KubeadmConfigobject, generates thekubeadm initcloud-init script, and stores it as a secure KubernetesSecret. - Machine & VM Provisioning: The Core Machine Controller binds the
Secretto theAWSMachineCR. The AWS Infrastructure Controller reads the secret, calls the AWS EC2 API to launch the VM with thecloud-initpayload, and setsAWSMachine.status.ready = trueonce running.
5. Mental Experiment / Debugging: Tracing Reconciliation Failures
Because CAPI relies on asynchronous controller reconciliation, troubleshooting a stuck cluster requires following the dependency graph rather than sifting through monolithic logs.
The Diagnostic Decision Tree
If a workload cluster fails to provision, use this systematic checklist:
flowchart TD
Start["Cluster Provisioning Stuck?"] --> Q1{"Is Cluster\nstatus.infrastructureReady == true?"}
Q1 -- No --> A1["Check InfraCluster CRD & Infrastructure Provider Logs"]
Q1 -- Yes --> Q2{"Is ControlPlane\nstatus.ready == true?"}
Q2 -- No --> Q3{"Are Control Plane\nMachines created?"}
Q3 -- No --> A3["Check KubeadmControlPlane Controller Logs"]
Q3 -- Yes --> Q4{"Is BootstrapConfig\nstatus.ready == true?"}
Q4 -- No --> A4["Check Bootstrap Provider (CAPBK) Logs & Secret Generation"]
Q4 -- Yes --> Q5{"Is InfraMachine\nstatus.ready == true?"}
Q5 -- No --> A5["Check Cloud Provider API quota, IAM, or VM creation errors"]
Q5 -- Yes --> A6["Check Node registration, kubelet logs, and CNI plugin"]
Q2 -- Yes --> Success["Control Plane & Cluster Fully Reconciled!"]
Pro-Tip: Always inspect
.status.conditionson CAPI custom resources usingkubectl get cluster <name> -o yaml. Cluster API controllers write detailed condition reasons (e.g.,WaitingForInfrastructureReady,WaitingForKubeadmInit) directly into object status.
6. Tooling vs. Runtime: clusterctl vs. Controllers
A common misconception when starting with CAPI is confusing clusterctl with the Cluster API control plane.
flowchart TD
subgraph Tooling["clusterctl (CLI Tool)"]
T1["Initializes Management Cluster"]
T2["Installs Provider CRDs & Managers"]
T3["Generates Workload Cluster Templates"]
end
subgraph Runtime["CAPI Controllers (Background Daemon)"]
R1["Continuous etcd Watch Loops"]
R2["Reconciles Infrastructure & Machines 24/7"]
R3["Self-Heals Node Failures"]
end
Tooling -->|Deploys / Configures| Runtime
clusterctl: A client-side operational CLI tool used to transform a plain Kubernetes cluster into a CAPI Management Cluster (viaclusterctl init), install provider controllers, and generate client templates.- CAPI Controllers: Long-running controller managers running inside the management cluster that continuously enforce desired state 24/7.
7. Key Learning Takeaways
- Decoupled Architecture: Core Cluster API is cleanly decoupled from infrastructure implementation through explicit Provider Contracts.
- Asynchronous Coordination: Controllers do not call each other directly; they communicate by reading and updating Kubernetes custom resource specs and status fields.
- Layered Responsibilities: Core CAPI handles generic machine lifecycle, Infrastructure Providers manage VMs/networking, Bootstrap Providers build initialization scripts, and Control Plane Providers manage etcd & control plane quorum.
- Condition-Based Debugging: Troubleshooting relies on following the object reference graph (
Cluster->InfraCluster->ControlPlane->Machine->InfraMachine) and checking status conditions.
8. Self-Check Questions
Before moving to Part 03, test your comprehension:
- If the AWS Infrastructure Controller crashes, what happens to existing workload clusters managed by CAPI?
- Why does the Kubeadm Bootstrap Controller write
cloud-initscripts into a Kubernetes Secret rather than passing them directly to the Infrastructure Controller? - How does Core Cluster API verify that an
AWSMachinehas finished provisioning without having any AWS SDK code compiled into its binary?
Next Milestone in the Learning Journey
In Part 03: Understanding the Cluster API Object Model, we will dissect the exact YAML specs and field relationships of Cluster, Machine, MachineSet, MachineDeployment, KubeadmControlPlane, and provider templates.