Learning Journey: Part 05 of 16

Module Goal: Build a clear operational mental model of clusterctl: understand its precise architectural boundary against continuous reconciliation, manage provider lifecycles, render workload cluster manifests, inspect object trees, and pivot management cluster state.
Target Spec: Cluster API v1.14 (v1beta2 APIs).
Prerequisites: Part 04: Building Your First Cluster API Environment.


1. The Core Boundary: CLI Tooling vs. Continuous Reconciliation

In Part 04, we used clusterctl to initialize a local management cluster and render a workload cluster manifest.

When getting started with Cluster API (CAPI), it is easy to adopt the wrong mental model:

"clusterctl creates, scales, and manages Kubernetes clusters."

It does not.

clusterctl is primarily a client-side management CLI tool for the Cluster API management plane.

Its responsibilities include:

  • Installing and upgrading Cluster API provider controllers.
  • Rendering parameterized cluster templates into standard Kubernetes YAML.
  • Inspecting the status tree and resource relationships of workload clusters.
  • Retrieving workload cluster credentials (kubeconfig).
  • Migrating (pivoting) CAPI management resources from one management cluster to another.

The continuous, 24/7 lifecycle management of workload clusters—provisioning VMs, bootstrap execution, health monitoring, and scaling—is performed entirely by in-cluster controllers running inside the management cluster.

flowchart TD
    subgraph ClientSide["Client-Side Tooling: Point-in-Time Execution"]
        CTL["clusterctl CLI"]
    end

    subgraph ManagementPlane["Management Cluster: Continuous 24/7 Engine"]
        API["Kubernetes API Server / etcd"]
        CAPI_CTRL["CAPI Core Controllers"]
        INFRA_CTRL["Infrastructure Controllers: CAPD / CAPA"]
        BOOT_CTRL["Bootstrap & Control Plane Controllers"]

        API <--> CAPI_CTRL
        API <--> INFRA_CTRL
        API <--> BOOT_CTRL
    end

    subgraph TargetInfra["Infrastructure Layer"]
        WORKLOAD["Workload Cluster VMs & Nodes"]
    end

    CTL -->|"1. init / upgrade: Installs CRDs & Deployments"| API
    CTL -->|"2. generate cluster: Renders Client YAML"| Templating["Local YAML File"]
    Templating -.->|"kubectl apply"| API
    
    CAPI_CTRL -->|"3. Continuous Reconciliation Loop"| TargetInfra
    INFRA_CTRL -->|"3. Continuous Reconciliation Loop"| TargetInfra
    BOOT_CTRL -->|"3. Continuous Reconciliation Loop"| TargetInfra

Core Insight: clusterctl operates outside the reconciliation loop. Once you apply a Cluster or MachineDeployment spec to etcd, controllers reconcile the target state whether clusterctl is running, closed, or uninstalled.


2. Where clusterctl Fits in the Architecture

To understand where clusterctl fits, divide Cluster API operations into two distinct domains:

clusterctl
    = Setup, manifest generation, object inspection, provider lifecycle management

controllers (capi-system, capd-system, etc.)
    = Continuous state reconciliation, drift correction, self-healing, rolling updates

What Happens After clusterctl Exits?

Consider this common workflow:

# 1. Initialize management cluster
clusterctl init --infrastructure docker

# 2. Render workload manifest
clusterctl generate cluster my-cluster --kubernetes-version v1.37.0 > cluster.yaml

# 3. Apply manifest to management cluster
kubectl apply -f cluster.yaml

If you close your terminal or delete the clusterctl binary at this point:

  1. The management cluster controllers continue running inside Kubernetes.
  2. If a worker node crashes, the MachineSet controller automatically provisions a replacement.
  3. If you scale a MachineDeployment using kubectl scale, the CAPI controllers handle node expansion.
  4. clusterctl does not sit in memory, poll cloud APIs, or listen for events.

3. Management Plane Lifecycle: clusterctl init & Repositories

Before a Kubernetes cluster can act as a CAPI management plane, it must run the required custom resource definitions (CRDs) and controller deployments.

Initializing the Management Plane

The command:

clusterctl init --infrastructure docker

transforms a standard Kubernetes cluster into a Cluster API management cluster.

flowchart LR
    K8S["Standard Kubernetes Cluster"] -->|clusterctl init| Management["CAPI Management Cluster"]
    
    subgraph Management["Installed Provider Components"]
        CAPI["capi-system\n(Core Controller)"]
        KUBE_B["capi-kubeadm-bootstrap-system\n(Bootstrap Controller)"]
        KUBE_C["capi-kubeadm-control-plane-system\n(Control Plane Controller)"]
        CAPD["capd-system\n(Docker Infra Controller)"]
    end

By default, clusterctl init fetches and installs four provider components:

  1. Core Provider: cluster-api (manages Cluster, Machine, MachineSet, MachineDeployment).
  2. Bootstrap Provider: kubeadm (generates cloud-init / kubeadm join data).
  3. Control Plane Provider: kubeadm (manages control plane node topology and upgrades).
  4. Infrastructure Provider: The specified target infrastructure (e.g., docker, aws, azure, vsphere).

You can verify the running provider controller pods with:

kubectl get deployments -A | grep -E 'capi|capd|kubeadm'

Understanding Provider Repositories

How does clusterctl know where to fetch provider CRDs and controller manifests?

clusterctl reads provider release metadata from provider repositories (usually GitHub releases or custom HTTP endpoints).

You can inspect all default and configured repositories using:

clusterctl config repositories

Output highlights:

NAME                    TYPE                    URL
cluster-api             CoreProvider            https://github.com/kubernetes-sigs/cluster-api/releases/latest/core-components.yaml
bootstrap-kubeadm       BootstrapProvider       https://github.com/kubernetes-sigs/cluster-api/releases/latest/bootstrap-components.yaml
control-plane-kubeadm   ControlPlaneProvider    https://github.com/kubernetes-sigs/cluster-api/releases/latest/control-plane-components.yaml
docker                  InfrastructureProvider  https://github.com/kubernetes-sigs/cluster-api/releases/latest/infrastructure-components-docker.yaml
aws                     InfrastructureProvider  https://github.com/kubernetes-sigs/cluster-api-provider-aws/releases/latest/infrastructure-components.yaml

Repository Config vs. Installed Inventory:

  • clusterctl config repositories lists providers clusterctl knows how to fetch.
  • kubectl get providers -A lists providers currently installed in the cluster.

Inspecting Installed Provider Inventory

During initialization, clusterctl creates Provider custom resources in the management cluster to keep track of installed provider versions:

kubectl get providers.clusterctl.cluster.x-k8s.io -A

Or using the short name:

kubectl get providers -A

Example output:

NAMESPACE                           NAME                    TYPE                    VERSION   STATUS
capi-system                         cluster-api             CoreProvider            v1.14.0   Installed
capi-kubeadm-bootstrap-system       bootstrap-kubeadm       BootstrapProvider       v1.14.0   Installed
capi-kubeadm-control-plane-system   control-plane-kubeadm   ControlPlaneProvider    v1.14.0   Installed
capd-system                         infrastructure-docker   InfrastructureProvider  v1.14.0   Installed

This internal inventory tracking is critical when performing provider upgrades via clusterctl upgrade.

Preparing Air-Gapped Environments

For restricted networks or enterprise environments requiring private container registries, clusterctl can pre-calculate required container images before cluster installation:

clusterctl init list-images --infrastructure docker

This outputs every controller image required by the specified providers so they can be mirrored to a private registry.


4. Rendering Workload Manifests: clusterctl generate cluster

Creating workload cluster manifests by hand would require writing hundreds of lines of interconnected YAML (Cluster, DockerCluster, KubeadmControlPlane, DockerMachineTemplate, MachineDeployment, DockerMachineTemplate, KubeadmConfigTemplate).

clusterctl generate cluster automates this by fetching templates and substituting configuration variables.

clusterctl generate cluster capi-quickstart \
  --flavor development \
  --kubernetes-version v1.37.0 \
  --control-plane-machine-count=1 \
  --worker-machine-count=1 \
  > capi-quickstart.yaml

Template Resolution Flow

flowchart TD
    subgraph Inputs["Template & Inputs"]
        P["Provider Release Repository"] -->|Fetch Template| T["cluster-template.yaml"]
        E["Environment Variables\n(e.g., CONTROL_PLANE_MACHINE_COUNT)"]
        F["CLI Flags\n(e.g., --kubernetes-version)"]
        C["clusterctl.yaml Config"]
    end

    subgraph Engine["clusterctl Templating Engine"]
        T & E & F & C --> Sub["Variable Substitution & Manifest Expansion"]
    end

    subgraph Output["Output Manifest"]
        Sub --> Manifest["capi-quickstart.yaml\n(Multi-Document Kubernetes Spec)"]
    end

Provider Flavors

Infrastructure providers publish multiple pre-packaged topologies known as flavors.

For example, the Docker provider might provide:

  • default: Standard multi-node setup.
  • development: Minimal resource footprint for local testing.
  • ha: High-availability control plane (3 nodes) with external load balancing.
  • ipv6: IPv6 networking topology.

To generate a manifest using a specific flavor:

clusterctl generate cluster my-cluster --flavor ha ...

Rendering Custom YAML Files

If you have custom YAML files containing clusterctl environment variables (e.g. ${KUBERNETES_VERSION}), you can substitute variables without using a provider template:

clusterctl generate yaml input.yaml --output output.yaml

5. Inspection and Debugging: clusterctl describe & get kubeconfig

Once a workload cluster spec has been applied to the management cluster, monitoring its provisioning progress using raw kubectl get commands can require querying multiple custom resources.

Higher-Level Resource Trees with clusterctl describe

clusterctl describe cluster aggregates CAPI custom resources into an easy-to-read hierarchy tree showing object health and conditions:

clusterctl describe cluster capi-quickstart

Example visual output:

NAME                                                            STATUS  REASON  AGE  MESSAGE
Cluster/capi-quickstart                                         True            2m   
├─ControlPlane/capi-quickstart-control-plane                    True            2m   
│ └─Machine/capi-quickstart-control-plane-4x8z9                True            2m   
└─Workers                                                                            
  └─MachineDeployment/capi-quickstart-md-0                      True            2m   
    └─2 Machines...                                             True            2m   

Fine-Grained Troubleshooting Flags

By default, clusterctl describe groups identical healthy worker machines to keep output concise. During debugging, group suppression can be disabled:

# Expand all grouped Machine objects individually
clusterctl describe cluster capi-quickstart --disable-grouping

# Show underlying infrastructure (DockerMachine/AWSMachine) and bootstrap resources
clusterctl describe cluster capi-quickstart --disable-no-echo
flowchart TD
    A["1. Run clusterctl describe cluster <name>"] --> B{"Is hierarchy healthy?"}
    B -->|No| C["2. Identify failing resource\n(e.g. Machine/capi-quickstart-md-0-xyz)"]
    C --> D["3. Run kubectl describe <resource>"]
    D --> E["4. Inspect Controller Logs\nkubectl logs -n capd-system deployment/capd-controller-manager"]
    B -->|Yes| F["Cluster Provisioned Successfully"]

Accessing Workload Clusters

Once the control plane is initialized, retrieve the workload cluster’s kubeconfig secret from the management cluster using:

clusterctl get kubeconfig capi-quickstart > capi-quickstart.kubeconfig

You can then interact directly with the workload cluster API server:

kubectl --kubeconfig ./capi-quickstart.kubeconfig get nodes

6. Management Plane Maintenance: clusterctl upgrade

As Cluster API releases new versions (e.g., moving from CAPI v1.13 to v1.14), the provider controllers running in the management cluster need to be upgraded.

Critical Distinction:

  • clusterctl upgrade updates the management plane provider controllers and CRDs.
  • It does not upgrade the Kubernetes version of workload clusters. (Workload Kubernetes upgrades are driven by editing KubeadmControlPlane and MachineDeployment specs).

Step 1: Check Upgrade Availability

clusterctl upgrade plan

clusterctl inspects the installed Provider custom resources in the management cluster, queries configured release repositories, and outputs available upgrade paths:

Checking provider versions...
CoreProvider/cluster-api
  Installed version: v1.13.2
  Latest version:    v1.14.0

InfrastructureProvider/infrastructure-docker
  Installed version: v1.13.2
  Latest version:    v1.14.0

Target versions for upgrade:
  CoreProvider/cluster-api: v1.14.0
  InfrastructureProvider/infrastructure-docker: v1.14.0

Step 2: Apply Provider Upgrades

To perform the provider upgrade across the management cluster:

clusterctl upgrade apply --contract v1beta1 --core capi-system/cluster-api:v1.14.0

7. Management State Migration: clusterctl move (The Pivot Operation)

One of the most powerful features of clusterctl is pivoting: transferring management ownership of workload clusters from a temporary bootstrap cluster to a permanent management cluster.

The Pivot Pattern

flowchart TD
    subgraph Phase1["1. Bootstrap Phase"]
        KIND["Temporary Bootstrap Cluster\n(kind + clusterctl init)"]
        KIND -->|1. Provision| PERM_VM["Permanent Management Cluster\n(Workload Cluster A)"]
    end

    subgraph Phase2["2. Target Preparation"]
        INIT["clusterctl init on Permanent Cluster"]
        PERM_VM --> INIT
    end

    subgraph Phase3["3. Pivot Operation"]
        MOVE["clusterctl move --to-kubeconfig permanent.kubeconfig"]
        KIND -->|Transfer CAPI Objects & State| MOVE
        MOVE -->|Write Objects & Resume Reconciliation| PERM_VM
    end

    subgraph Phase4["4. Teardown"]
        DEL["Delete Temporary kind Cluster"]
    end

    Phase1 --> Phase2 --> Phase3 --> Phase4

Executing clusterctl move

To migrate CAPI objects from the current management cluster to a target management cluster:

clusterctl move --to-kubeconfig target-management.kubeconfig

What Happens During a Move?

  1. Pause Reconciliation: clusterctl adds the cluster.x-k8s.io/paused: "true" annotation to all Cluster resources in the source cluster. This prevents source controllers from deleting cloud infrastructure while state is transferred.
  2. Dependency Graph Discovery: clusterctl scans etcd to build a complete object tree including Cluster, Machine, KubeadmControlPlane, MachineDeployment, infrastructure objects (DockerCluster/AWSCluster), and associated Secrets (kubeconfig credentials, TLS certificates).
  3. Export and Apply: Objects are serialized and created in the target management cluster.
  4. Owner Reference Restoration: Object references and ownership trees are reconciled on the target cluster.
  5. Unpause Reconciliation: The paused annotation is removed on the target management cluster, enabling its controllers to take over continuous reconciliation.
  6. Source Teardown: Objects are cleaned up from the source cluster without deleting the underlying workload infrastructure.

Why clusterctl move is Not a Generic Backup Tool:
clusterctl move is designed specifically for active management handoffs and pivot operations. It requires a quiet, stable workload cluster state. It should not be used as a routine backup/restore solution while cluster modifications or rolling updates are actively in flight.


8. Provider Uninstallation: clusterctl delete

To remove Cluster API provider components from a management cluster:

clusterctl delete --infrastructure docker

clusterctl delete vs. kubectl delete cluster

It is crucial to distinguish between removing provider tooling and deleting a workload cluster:

OperationCommandOperational Effect
Delete Workload Clusterkubectl delete cluster <name>CAPI controllers reconcile teardown: terminate cloud VMs, delete load balancers, clean up storage.
Uninstall Providerclusterctl delete --infrastructure dockerRemoves CAPD controller deployment and CRDs from the management cluster. Workload VMs remain running in cloud, but lose management plane oversight.

[!WARNING] Running clusterctl delete --include-crd removes provider Custom Resource Definitions from etcd. Deleting CRDs causes Kubernetes to delete all instance resources of that CRD type!


9. clusterctl vs. kubectl: Operational Decision Matrix

Use the following reference guide to determine whether to use clusterctl or kubectl for common Cluster API tasks:

Goal / TaskToolExact Command
Check local CLI versionclusterctlclusterctl version
List known provider repositoriesclusterctlclusterctl config repositories
Turn K8s cluster into CAPI Management Planeclusterctlclusterctl init --infrastructure <provider>
Pre-fetch provider container imagesclusterctlclusterctl init list-images ...
List installed management plane providerskubectlkubectl get providers -A
Render workload cluster YAML from templateclusterctlclusterctl generate cluster <name> ...
Apply workload cluster manifestkubectlkubectl apply -f cluster.yaml
Inspect object hierarchy tree & statusclusterctlclusterctl describe cluster <name>
Inspect low-level CRD fields & status conditionskubectlkubectl describe machine <machine-name>
Fetch workload cluster access credentialsclusterctlclusterctl get kubeconfig <name>
Check for available provider upgradesclusterctlclusterctl upgrade plan
Upgrade management plane provider controllersclusterctlclusterctl upgrade apply ...
Upgrade workload cluster Kubernetes versionkubectlkubectl patch kubeadmcontrolplane ...
Pivot management state to another clusterclusterctlclusterctl move --to-kubeconfig ...
Uninstall CAPI provider controllersclusterctlclusterctl delete --infrastructure ...
Teardown workload cluster infrastructurekubectlkubectl delete cluster <name>

10. Key Learning Takeaways

  • Management Tooling, Not Reconciler: clusterctl prepares, inspects, and moves management plane configurations. It does not run inside the continuous reconciliation loop.
  • Provider Inventory Management: clusterctl init creates Provider custom resources that track installed controller versions for subsequent clusterctl upgrade operations.
  • Client-Side Manifest Generation: clusterctl generate cluster is a template rendering engine that outputs standard Kubernetes YAML for consumption by kubectl apply.
  • Tree-Based Inspection: clusterctl describe cluster provides an aggregated view of Cluster API object trees, making it the fastest entry point for debugging provisioning issues.
  • Pivoting State: clusterctl move pauses reconciliation and safely transfers complex object graphs and secrets from a temporary bootstrap cluster to a permanent management plane.

Next: What Actually Happens When You Create a Cluster?

We now have a complete mental model of CAPI architecture, the custom resource object graph, local environment setup, and clusterctl management tooling.

In Part 06, we will trace a single operation end-to-end through the system:

kubectl apply -f cluster.yaml

We will follow the exact controller sequence:

  1. Cluster reconciliation and InfrastructureCluster initialization.
  2. KubeadmControlPlane orchestration and bootstrap secret creation.
  3. MachineSet expansion and InfrastructureMachine VM provisioning.
  4. kubeadm join execution and Node registration back to the workload control plane.

References