Upgrading Clusters on VMware vSphere

This document explains how to upgrade Kubernetes clusters on VMware vSphere after the platform-side distribution upgrade is complete. The documented workflow focuses on updating the control plane and worker nodes through Cluster API resources. The provider behavior and rules on this page are validated against VMware vSphere Provider v1.0.17. The rules the admission webhooks enforce are listed in Provider Requirements.

INFO

Where this page fits in the full ACP upgrade flow

This page covers only the Kubernetes step of the upgrade. The full ACP upgrade flow — including upgrade artifact synchronization, ACP Core upgrade through CVO, Aligned plugin upgrades, and Agnostic plugin upgrades from Marketplace — is documented in the ACP product documentation. Complete those steps before you start the Kubernetes step on this page:

Use this page when the same cluster runs on an immutable operating system, because the Kubernetes step on immutable OS replaces nodes from a new VM template rather than upgrading binaries in place.

Upgrade Sequence

Upgrade VMware vSphere clusters in the following order:

  1. (Prerequisite) Upgrade the ACP platform on the management cluster first. This brings the cluster-api-provider-vsphere controller and the related CAPI components to versions that understand the new schema. The provider updates its CRDs and migrates existing resources to the new field layout itself — no manual YAML change is required. See Upgrading the Provider. Trigger workload-cluster upgrades only after the management-side controllers have rolled out and become Ready.
  2. Complete the distribution-version upgrade described in Upgrading Clusters.
  3. Verify that the existing KubeadmControlPlane and MachineDeployment manifests satisfy the rollout rules. A manifest that does not can block the very first edit of the upgrade. See Check the manifests against the admission rules.
  4. Verify that the control plane is healthy and the current cluster is stable.
  5. Upgrade Kube-OVN to the chart version required by the target ACP release and wait for the AppRelease to reach Success.
  6. Upgrade the control plane Kubernetes version.
  7. Upgrade worker nodes to the target Kubernetes version.

Prerequisites

Before you begin, ensure the following conditions are met:

  • The distribution-version upgrade is complete.
  • The control plane is healthy and reachable.
  • All nodes are in the Ready state.
  • A current etcd backup has been taken and verified by using the supported ACP backup procedure.
  • The VM template built from the Alauda OS image published for the target ACP release is present in the vSphere environment. The upgrade fails if the template is not present when the new VSphereMachineTemplate is applied.
  • For cross-version upgrades that span more than one Kubernetes minor, the intermediate-version Core images and VM templates are pre-staged. See Cross-Version Upgrade Preparation.
  • The target Kubernetes version is compatible with your workloads and add-ons.
  • The machine config pools have enough capacity for rolling updates.
  • The KubeadmControlPlane sets spec.rolloutStrategy.rollingUpdate.maxSurge: 0 and spec.replicas to 3 or more, and every MachineDeployment sets spec.strategy.rollingUpdate.maxSurge: 0 with maxUnavailable of 1 or more. These values let the replacement reuse the same slot and disk identity, and the admission webhooks enforce them. Verify them with Check the manifests against the admission rules before you start.
  • The management cluster is not a disaster-recovery standby. Against a standby, the provider skips non-global objects and requeues every 30 seconds, so the rollout appears not to start rather than to fail. See Standby Management Clusters.
  • You know the KubeadmControlPlane name for this cluster. Read it from Cluster.spec.controlPlaneRef.name; the manifests in Creating Clusters on VMware vSphere name it <cluster_name>-kcp.
  • If the cluster uses a Self-built VIP (controlPlaneLoadBalancer.type: internal), VSphereCluster.status.conditions[SelfBuiltLoadBalancerReady] is True, and the new VM template satisfies the Alive kernel prerequisites (the ip_vs module family is loadable, net.ipv4.vs.conntrack=1 is permitted). Replacement control-plane nodes are built from that template, and the provider neither injects nor validates these settings.
  • Review the Kubernetes upgrade path and version skew policy.
WARNING

The rollout rules can reject the first edit of an existing cluster

The rules apply to UPDATE as well as CREATE, so an existing cluster can hold values that are no longer accepted:

  • a KubeadmControlPlane whose rolloutStrategy.rollingUpdate.maxSurge is unset — the Cluster API default is 1;
  • a KubeadmControlPlane with replicas: 1;
  • a MachineDeployment whose strategy.rollingUpdate.maxUnavailable is unset — the Cluster API default is 0.

The rejection lands on the version bump itself, so the upgrade stops before anything is replaced. Correct these values in a separate, earlier edit. A KubeadmControlPlane at replicas: 1 on a pool-backed template cannot be upgraded at all until it is scaled to 3; a single-replica control plane is not a supported fixed-IP topology. Scaling it is a two-part change: add two more slots to the control-plane VSphereMachineConfigPool — hostnames, static addresses, and the required persistent disks — and then raise replicas and set maxSurge in one patch, because both rules are checked on the same update. See Check the manifests against the admission rules.

WARNING

Disk Preservation Model

Upgrades rely on Cluster API's rolling replacement mechanism. Each cluster has five disk classes; only the pool-managed persistent class and external CSI volumes survive a delete-recreate.

Disk classDeclared inSurvives upgrade?Use for
System disk (root volume)The VM template used for spec.template.spec.template❌ NeverOS + kubelet/kubeadm/containerd. Rebuilt from the new template every replacement.
Template-local disksVSphereMachineTemplate.spec.template.spec.* (additional disks declared in the template)❌ NeverEphemeral cache. Destroyed with the old VM.
Pool-managed ephemeral disksVSphereMachineConfigPool.spec.configs[].ephemeralDisks[]❌ Never — the VMDK is deleted with the VM and a new empty disk is created for the replacementRebuildable node-local data that should not sit on the system disk, including the standard /var/lib/kubelet, /var/lib/containerd, and /var/lib/etcd disks.
Pool-managed persistent disksVSphereMachineConfigPool.spec.configs[].persistentDisks[]✅ Detached from old VM and reattached to the new VM at the same slotPlatform state such as /var/cpaas.
External CSI volumes (vSphere CSI, etc.)Workload PVCs / CSI driver✅ Unrelated to node lifecycleApplication data.

Unlike pool-managed persistent disks, ephemeral disks take no part in reclaim and never hold a slot release, so they never delay a rollout. See Choose persistent or ephemeral.

Which class your disks are in is fixed at cluster creation. ephemeralDisks[] is new in provider v1.0.17. A cluster created on an earlier provider declares /var/cpaas, /var/lib/containerd, and /var/lib/etcd as persistent disks, with wipeFilesystem: true on the etcd disk, and an upgrade does not change that: the provider migrates no disk, and an allocated slot rejects the removal or the reshaping of a disk in either list, so a disk cannot be moved from one list to the other. Do not restructure the pool's disk lists as part of an upgrade — the standard persistent-plus-ephemeral split applies to newly created clusters only.

"Preserved" means the same disk identity is reattached — it does not mean the disk's contents are time-traveled. Anything written to a pool-managed disk during the upgrade window stays after the upgrade and stays after a rollback.

WARNING

Templates Cannot Be Modified In Place

VSphereMachineTemplate.spec.template.spec is immutable. The vSphere admission webhook rejects any update with the message "VSphereMachineTemplate spec.template.spec field is immutable. Please create a new resource instead." Every upgrade step on this page therefore creates a new VSphereMachineTemplate with a new metadata.name, applies it, and then patches the controlling resource's infrastructureRef.name to the new template. Keep the previous template until the new rollout is healthy in case rollback is required.

INFO

Fleet Essentials boundary

Fleet Essentials 1.0.4 and later can request the ACP 4.3-and-later Distribution Version upgrade through CVO. It does not perform the vSphere Kubernetes and Alauda OS replacement described on this page. Complete Phase 1 with the ACP workflow, then use the YAML procedure below for Phase 2.

WARNING

An upgraded cluster cannot use the Self-built VIP

The Self-built VIP is available only to clusters created with provider v1.0.17 or later on ACP v4.4 or later. A cluster that was created on an earlier provider keeps its external LoadBalancer after the provider is upgraded, and cannot be moved onto a Self-built VIP: the provider injects the VIP while the first control-plane node bootstraps, and there is no in-place migration for a running control plane. See Control Plane Endpoint Modes.

The endpoint itself does not move, and an upgrade is not an opportunity to change it. The control plane endpoint — the mode and the VIP or LoadBalancer address — is chosen when the cluster is deployed and is fixed from then on; it cannot be adjusted while the cluster runs or during an upgrade. An existing cluster needs no action and gets none: spec.controlPlaneLoadBalancer stays unset, which is exactly what an external LoadBalancer looks like, and writing it now is rejected. The field is optional for this reason — a cluster created before the Self-built VIP existed keeps its endpoint configuration untouched across the upgrade. See Upgrading the Provider.

Required Values From the OS Support Matrix

The authoritative mapping between an ACP release, the Kubernetes version, and the matching CoreDNS, etcd, and Kube-OVN versions lives in OS Support Matrix. Locate the row that corresponds to the target ACP version before you start; the row supplies the component values the steps below need.

The VM template name is not in the matrix. It is the name the template was created under in vSphere from the Alauda OS image published for the target ACP release — see Downloading Alauda OS — and it is what VSphereMachineTemplate.spec.template.spec.template must be set to on both the control plane and worker templates.

The cells you read from that row map to the upgrade manifests as follows:

OS Support Matrix columnUsed to setWhere it lands
Kubernetes VersionKubeadmControlPlane.spec.version and MachineDeployment.spec.template.spec.versionBoth control plane and worker
corednsKubeadmControlPlane.spec.kubeadmConfigSpec.clusterConfiguration.dns.imageTagControl plane only. Set only the tag — the provider sets dns.imageRepository itself, derived from the -v<acp> suffix in that tag, and also adds the CoreDNS and kube-proxy skip annotations to your KubeadmControlPlane. See Provider-managed fields on your manifests
etcdKubeadmControlPlane.spec.kubeadmConfigSpec.clusterConfiguration.etcd.local.imageTagControl plane only
kube-ovn (chart)Cluster.metadata.annotations["cpaas.io/kube-ovn-version"] and the provider-managed cni-kube-ovn AppReleaseComplete the shared provider-version-specific Kube-OVN procedure before changing KubeadmControlPlane. The provider version and target chart version determine the chart name and supported reconciliation flow.

The CoreDNS and etcd image tags are control-plane-only because clusterConfiguration is a KubeadmControlPlane field. Worker nodes inherit container image versions from the new VM template; the MachineDeployment does not carry its own dns/etcd tags. The Kube-OVN annotation lives on the Cluster resource, not on KubeadmControlPlane, because the vSphere provider watches it independently of the Kubernetes control plane rollout.

Steps

Check the manifests against the admission rules

Read the current values before you change anything:

kubectl -n <namespace> get kubeadmcontrolplane <cluster_name>-kcp \
  -o jsonpath='replicas={.spec.replicas} maxSurge={.spec.rolloutStrategy.rollingUpdate.maxSurge}{"\n"}'

kubectl -n <namespace> get machinedeployment -o custom-columns='NAME:.metadata.name,TYPE:.spec.strategy.type,SURGE:.spec.strategy.rollingUpdate.maxSurge,UNAVAIL:.spec.strategy.rollingUpdate.maxUnavailable'

An empty maxSurge or UNAVAIL value means the field is unset, and unset means the Cluster API default — 1 and 0 respectively. Neither default is valid for a pool-backed controller. Correct them before the version bump:

# replicas is already 3 or more: correct maxSurge on its own.
kubectl -n <namespace> patch kubeadmcontrolplane <cluster_name>-kcp --type='merge' \
  -p='{"spec":{"rolloutStrategy":{"type":"RollingUpdate","rollingUpdate":{"maxSurge":0}}}}'

# replicas is 1: set replicas and maxSurge in the same patch. Declare the two extra
# control-plane slots first.
kubectl -n <namespace> patch kubeadmcontrolplane <cluster_name>-kcp --type='merge' \
  -p='{"spec":{"replicas":3,"rolloutStrategy":{"type":"RollingUpdate","rollingUpdate":{"maxSurge":0}}}}'

kubectl -n <namespace> patch machinedeployment <cluster_name>-md-0 --type='merge' \
  -p='{"spec":{"strategy":{"type":"RollingUpdate","rollingUpdate":{"maxSurge":0,"maxUnavailable":1}}}}'

The webhook evaluates the maxSurge and replicas rules on every KubeadmControlPlane update and reports them together, so on a replicas: 1 cluster a patch that corrects only one of them is rejected by the other. Never patch replicas on a control plane that already has three or more; the value above is a floor, not a target.

A MachineDeployment that uses strategy.type: OnDelete is exempt from both rules, because OnDelete already replaces machines delete-first.

Verify that the pools are healthy before you start:

kubectl -n <namespace> get vspheremachineconfigpool
kubectl -n <namespace> get vspheremachineconfigpool <cluster_name>-cp-pool \
  -o jsonpath='{range .status.conditions[*]}{.type}={.status} ({.reason}){"\n"}{end}'

Expect Ready=True, MembersValid=True, MembersUnique=True, and PersistentDisksReady=True. SlotAvailable can legitimately be False on a fully allocated pool; it is not part of the Ready summary.

Upgrade Kube-OVN Before the Control Plane

Follow Upgrade Kube-OVN Before the Control Plane and select the VMware vSphere tab that matches the installed provider version: a cluster still on v1.0.15 or earlier follows the legacy tab, and a Kube-OVN v4.4 or later target requires upgrading the provider to v1.0.16 or later first. The shared procedure includes the vSphere-specific network annotation check, the chart-name migration at Kube-OVN v4.4, and the required AppRelease health checks. Do not patch the chart source by hand; the provider reconciles it from the annotation.

Create the target machine templates

Before you start the rolling upgrade, create new VSphereMachineTemplate resources for the control plane and workers.

  1. Export the existing control plane template

    kubectl get vspheremachinetemplate <cluster_name>-control-plane -n <namespace> -o yaml > new-cp-template.yaml
  2. Modify the control plane template

    Edit new-cp-template.yaml:

    • Set metadata.name to a new unique name (for example, <cluster_name>-control-plane-v2)
    • Update spec.template.spec.template to the target VM template name
    • Update CPU, memory, or disk settings if needed
    • Remove server-generated fields: metadata.resourceVersion, metadata.uid, metadata.generation, metadata.creationTimestamp, metadata.managedFields, metadata.annotations["kubectl.kubernetes.io/last-applied-configuration"], and status
    • Leave spec.template.spec.providerID unset. The vSphere provider sets providerID to the VM's BIOS UUID once the VM is created; pre-filling it in the template breaks the controller's identity binding.
  3. Export and modify the worker template

    kubectl get vspheremachinetemplate <cluster_name>-worker -n <namespace> -o yaml > new-worker-template.yaml

    Edit new-worker-template.yaml:

    • Set metadata.name to a new unique name (for example, <cluster_name>-worker-v2)
    • Update spec.template.spec.template to the target VM template name
    • Update CPU, memory, or disk settings if needed
    • Remove the same server-generated fields listed above
  4. Apply both new templates

    kubectl apply -f new-cp-template.yaml
    kubectl apply -f new-worker-template.yaml

Upgrade the control plane

Before you start, complete the shared Upgrade Kube-OVN Before the Control Plane procedure and verify that the cni-kube-ovn AppRelease is at the target revision with phase=Success. Then collect every required control-plane value from the target ACP row in the OS Support Matrix as described in Required Values From the OS Support Matrix.

  1. Patch the KubeadmControlPlane with the target Kubernetes values

    Update the KubeadmControlPlane resource in a single edit to keep spec.version, the CoreDNS image tag, the etcd image tag, and the infrastructure template reference consistent with the same VM template:

    • spec.version ← Kubernetes Version from the OS Support Matrix row

    • spec.kubeadmConfigSpec.clusterConfiguration.dns.imageTag ← coredns column from the same row

    • spec.kubeadmConfigSpec.clusterConfiguration.etcd.local.imageTag ← etcd column from the same row

    • spec.machineTemplate.infrastructureRef.name ← the new VSphereMachineTemplate name created above

    • When the target is Kubernetes 1.35 or later, update /etc/kubernetes/patches/kubeletconfiguration0+strategic.json in spec.kubeadmConfigSpec.files in this same edit, as described in Required kubelet patch for Kubernetes 1.35

      kubectl edit kubeadmcontrolplane <cluster_name>-kcp -n <namespace>

    Updating only spec.version is not sufficient. The CoreDNS and etcd image tags must move together with the Kubernetes version because they are built from the same release; leaving them at the previous values can result in CoreDNS and etcd pods that do not match the new Kubernetes minor version.

    INFO

    After the provider reconciles, expect the CoreDNS and kube-proxy skip annotations and a controller-set dns.imageRepository on the object you just edited. Do not remove them, and do not re-apply a stored manifest that lacks them. The workload cluster's kube-proxy DaemonSet and CoreDNS Deployment repositories are reconciled separately and only after this rollout completes, so a lag there is expected. See Provider-managed fields on your manifests.

    On a Self-built VIP cluster, Alive keeps the VIP throughout the rollout. The bootstrap VIP is injected only on the original kubeadm init node; replacement control-plane nodes join through the VIP and never receive it. As control-plane membership changes, the provider updates the Alive backend list from the current control-plane node addresses. Keep maxSurge: 0 — the VIP and etcd quorum both depend on one-at-a-time replacement.

  2. Monitor the control plane rollout

    kubectl -n <namespace> get kubeadmcontrolplane <cluster_name>-kcp -w
    kubectl -n <namespace> get machine -l cluster.x-k8s.io/control-plane

Upgrade the worker nodes

After the control plane upgrade completes, update the MachineDeployment to reference the new worker template and the target Kubernetes version.

Typical changes include:

  • spec.template.spec.version — the target Kubernetes version
  • spec.template.spec.infrastructureRef.name — the new VSphereMachineTemplate name
  • spec.template.spec.bootstrap.configRef.name — the new KubeadmConfigTemplate name. This is required for Kubernetes 1.35 or later so the worker receives the required kubelet patch; for earlier versions, change it only when other bootstrap settings must change. See Updating Bootstrap Templates.

The patch below changes only spec.template.spec. If the pre-flight check showed that the MachineDeployment strategy needed correcting, that correction must already be applied — an update carrying a non-compliant strategy is rejected in full, including the version field.

Apply the changes:

The following command includes the bootstrap reference required for Kubernetes 1.35 or later. For an earlier target with no bootstrap change, omit the bootstrap object.

kubectl patch machinedeployment <cluster_name>-md-0 -n <namespace> \
  --type='merge' -p='{
    "spec": {
      "template": {
        "spec": {
          "version": "<target_kubernetes_version>",
          "infrastructureRef": {
            "name": "<new-worker-template-name>"
          },
          "bootstrap": {
            "configRef": {
              "name": "<new-bootstrap-template-name>"
            }
          }
        }
      }
    }
  }'

Monitor the worker rollout:

kubectl -n <namespace> get machinedeployment <cluster_name>-md-0 -w
kubectl -n <namespace> get machine
kubectl --kubeconfig=/tmp/<cluster_name>.kubeconfig get nodes -o wide

OS Image Update Without a Kubernetes Version Change

Use this procedure to move nodes onto a new Alauda OS image while the Kubernetes version stays the same — for example, to pick up an OS security fix, or to retrofit the pod log directory relocation onto existing nodes.

  1. Create new VSphereMachineTemplate resources with the export, scrub, and rename procedure in Create the target machine templates. Change only spec.template.spec.template to the new VM template name.
  2. Apply both new templates.
  3. Patch KubeadmControlPlane.spec.machineTemplate.infrastructureRef.name, wait for the control-plane rollout to finish, then patch MachineDeployment.spec.template.spec.infrastructureRef.name.

Do not touch spec.version, clusterConfiguration.dns.imageTag, clusterConfiguration.etcd.local.imageTag, or the cpaas.io/kube-ovn-version annotation. An OS-image rollout is not a Kubernetes hop, so it does not need the Kube-OVN-first ordering.

The same rollout rules apply: maxSurge: 0 on both controllers and maxUnavailable of 1 or more on each MachineDeployment. Pool-managed persistent disks reattach exactly as they do in a version upgrade, and pool-managed ephemeral disks are destroyed and recreated empty.

To recover, point the controlling resource back at the previous template. This is stage 2 of Recovering From a Failed Phase 2 Upgrade.

Rollout Holds That Are Not Failures

Some rollout pauses are the provider protecting data. Do not clear them by hand.

Attachment-safety hold. When a VSphereVM is deleted, the provider scans the datacenter for VMs that still have the slot's persistent-disk backings attached. While any of them is attached, it keeps the finalizer and requeues. The same check guards pool reclaim. A rollout that appears stalled at the deletion of the first old machine is frequently this hold rather than a defect. Do not remove the finalizer.

Read the current state:

kubectl -n <namespace> get vspheremachineconfigpool <pool_name> \
  -o jsonpath='{range .status.persistentDiskStatuses[*]}{.hostname}{" "}{.name}{" "}{.phase}{" owner="}{.ownerMachineName}{" err="}{.lastError}{"\n"}{end}'
kubectl -n <namespace> get vspherevm -o custom-columns='NAME:.metadata.name,DELETED:.metadata.deletionTimestamp,FINALIZERS:.metadata.finalizers'

The normal handoff is Attached, then Available once the old VM is gone, then Attached again on the replacement. ownerMachineName and ownerMachineUID are preserved across the delete and recreate so the replacement reattaches the same disks.

Reclaimed is a tombstone. Do not delete a phase: Reclaimed entry to unstick a rollout. See Troubleshooting on the cluster-creation page for why removing it produces a re-seed loop.

SlotAvailable=False is expected. With maxSurge: 0 and a fully allocated pool, no slot is free until an old machine is deleted. SlotAvailable is excluded from the pool's Ready summary precisely so this does not read as a failure.

A node powered off in vCenter stays off. Once VSphereVM latches InitialPowerOnCompleted=True, the controller no longer forces a powered-off VM back on. A replacement VM created during a rollout has an unset latch, so a node that was deliberately powered off before the upgrade comes back powered on after replacement.

Recovering From a Failed Phase 2 Upgrade

Do not treat a Kubernetes minor downgrade as an ordinary rollback. Choose the recovery path from the rollout stage:

  1. No target-version control-plane Machine has been created: restore the previous Kube-OVN annotation and the previous KubeadmControlPlane and MachineDeployment manifest values. This cancels the target rollout before a new control-plane data format is introduced.
  2. Only the machine template or OS image changed, and the Kubernetes minor did not change: point the controlling resource back to the previous template. Cluster API performs another replacement rollout. Keep the Kubernetes minor unchanged.
  3. A control-plane Machine on the target Kubernetes minor has joined the cluster: do not patch Kubernetes, CoreDNS, or etcd back to the previous minor. Stop further rollout, repair forward on the target minor, or restore the cluster from the verified pre-upgrade backup by using the supported ACP recovery procedure.

If a target-minor control-plane Machine was created but never joined, first restore healthy etcd quorum and determine whether the failed replacement can be removed safely. Do not assume that changing the version fields alone is sufficient.

Keep these infrastructure facts in mind during any recovery:

  • The old VMs are gone. They were destroyed during the upgrade. Template recovery builds a fresh set of replacement machines; it does not restore the original VMs.
  • The old VSphereMachineTemplate resource must still exist. Do not delete the previous template until the new rollout is healthy. If you already deleted it, recreate it from version control or backup before attempting same-minor template recovery.
  • Pool-managed disk identity is preserved, but data state is not. Disks declared in VSphereMachineConfigPool.spec.configs[].persistentDisks[] reattach to the replacement machines at the same slot, but data written during the upgrade window remains on those disks.
  • Ephemeral disks are gone. Data on VSphereMachineConfigPool.spec.configs[].ephemeralDisks[] is destroyed with every replaced VM and cannot be recovered from the pool. Only pool-managed persistent disks reattach.
  • Do not reverse a rollout-rule correction. If stage 1 requires restoring earlier manifest values, leave maxSurge at 0 and maxUnavailable at 1 or more. Restoring the older, non-compliant values is not accepted.

For stage 1, use the vSphere rule in Restore Kube-OVN During Stage-1 Recovery. The provider restores the chart from the annotation. Wait until the restored AppRelease passes the shared verification before changing control-plane manifests.

The KubeadmControlPlane controller can block replacement while etcd is unhealthy. Recover quorum before retrying any safe replacement action.

Verification

Read the pool state back after the upgrade:

kubectl -n <namespace> get vspheremachineconfigpool
kubectl -n <namespace> get vspheremachineconfigpool <pool_name> \
  -o jsonpath='{range .status.persistentDiskStatuses[*]}{.hostname}{" "}{.name}{" "}{.phase}{"\n"}{end}'

Confirm the following results:

  • KubeadmControlPlane reaches the target version and desired replica count.
  • MachineDeployment reaches the target version and desired replica count.
  • Control plane and worker nodes return to the Ready state.
  • The vSphere CPI daemonset remains available in the workload cluster.
  • Each VSphereMachineConfigPool reports Ready=True and PersistentDisksReady=True, its ALLOCATED count equals the number of running Machines drawn from it, and its TOTAL count is unchanged.
  • Every persistent disk on an allocated slot reports phase: Attached.

Troubleshooting

IssueWhat to check
The kubectl edit or kubectl patch on KubeadmControlPlane or MachineDeployment is rejectedRun the pre-flight check. For the full message-to-cause map, see Admission rejections.
The rollout stalls after the first machine is deletedUsually the attachment-safety hold rather than a defect. See Rollout Holds That Are Not Failures.
The pool reports Ready=True but SlotAvailable=FalseNormal on a fully allocated pool during a maxSurge: 0 rollout. SlotAvailable is excluded from the Ready summary.
A replaced node's disks are emptyConfirm which list the disk is declared in. On a cluster created on v1.0.17 or later, /var/lib/kubelet, /var/lib/containerd, and /var/lib/etcd are ephemeral by design and start empty on every replacement. On an older cluster they are persistent disks, so also check wipeFilesystem, which is true on the etcd disk by design.
Nothing reconciles, no conditions change, and no error appearsThe management cluster may be a disaster-recovery standby. See Standby Management Clusters.
A workload node is powered off and does not come backExpected once InitialPowerOnCompleted has latched. Power it on in vCenter. See Rollout Holds That Are Not Failures.
A new control-plane machine never bootstrapsRead the BootstrapReady condition on VSphereVM and VSphereMachine; the reasons are BootstrapSecretGetFailed and BootstrapSecretContentInvalid.
The Kube-OVN annotation changed but the chart name or revision did not convergeFollow the shared Kube-OVN procedure. Do not patch the chart source manually.
The workload kube-proxy or coredns is still on the old image repository after the control plane reaches the target versionExpected lag — those are reconciled only after the control-plane rollout completes. See Provider-managed fields on your manifests.
A VSphereMachineTemplate update is rejected as immutableCreate a new template under a new name. See the warning near the top of this page.

Next Steps

After the Kubernetes upgrade is complete, continue with routine node operations in Managing Nodes on VMware vSphere.