VMware vSphere Provider Installation

This document explains how to install the VMware vSphere Infrastructure Provider and verify that the required Cluster API components are available on the global cluster. The provider version, CRD set, and admission webhooks on this page are validated against VMware vSphere Provider v1.0.17.

INFO

v1.0.17 is the provider packaging version. It bundles the CAPV controller, the vSphere CPI image, the provider chart, and the admission webhooks that enforce the fixed-IP rollout rules. For those rules, see Provider Requirements. For the change history, see Release Notes.

Prerequisites

Before installing the provider, ensure the following conditions are met:

  • You can access the global cluster.
  • You can access Customer Portal to download provider packages.
  • You know which platform release and provider package version you plan to install. This page targets v1.0.17.
  • The vCenter account that clusters will reference through VSphereCluster.spec.identityRef can list, create, and set VM custom attributes. Grant the Global → Manage custom attributes and Global → Set custom attribute privileges. The provider requires them whenever a VSphereCluster or its CAPI Cluster carries any vsphere.cluster.x-k8s.io/* annotation, and a missing privilege fails VM reconciliation rather than skipping the metadata write.
  • The global cluster is the active management cluster, not a disaster-recovery standby. See Standby Management Clusters.

Steps

Download the required packages

To use VMware vSphere with Immutable Infrastructure, download the following cluster plugins from Customer Portal:

  1. Alauda Container Platform Kubeadm Provider
  2. Alauda Container Platform VMware vSphere Infrastructure Provider

Upload the packages

Upload the downloaded packages to the platform repository.

For the detailed upload procedure, refer to Upload Packages.

Install the provider on the global cluster

Install both cluster plugins on the global cluster.

For the detailed cluster-plugin procedure, refer to Cluster Plugin.

Verify the installation

Check the module state:

kubectl get minfo -l cpaas.io/module-name=cluster-api-provider-vsphere
kubectl get minfo -l cpaas.io/module-name=cluster-api-provider-kubeadm

Both plugins must be present and in the Running state.

Check the registered CRDs:

kubectl get crd | grep -E '^vsphere[a-z]+\.infrastructure\.cluster\.x-k8s\.io\b'

Expected output includes all nine CRDs:

  1. vsphereclusters.infrastructure.cluster.x-k8s.io
  2. vsphereclusteridentities.infrastructure.cluster.x-k8s.io
  3. vsphereclustertemplates.infrastructure.cluster.x-k8s.io
  4. vspheredeploymentzones.infrastructure.cluster.x-k8s.io
  5. vspherefailuredomains.infrastructure.cluster.x-k8s.io
  6. vspheremachines.infrastructure.cluster.x-k8s.io
  7. vspheremachineconfigpools.infrastructure.cluster.x-k8s.io
  8. vspheremachinetemplates.infrastructure.cluster.x-k8s.io
  9. vspherevms.infrastructure.cluster.x-k8s.io

Check the registered admission webhooks. The ValidatingWebhookConfiguration object name varies by packaging, so match on the webhook names instead:

kubectl get validatingwebhookconfiguration \
  -o jsonpath='{range .items[*]}{range .webhooks[*]}{.name}{"\n"}{end}{end}' \
  | grep -E 'vsphere|capv'

Expected output includes:

  • validation.vspherecluster.infrastructure.cluster.x-k8s.io
  • validation.vspheremachineconfigpool.infrastructure.cluster.x-k8s.io
  • validation.vspheremachinetemplate.infrastructure.cluster.x-k8s.io
  • validation.kubeadmcontrolplane.capv.cluster.x-k8s.io
  • validation.machinedeployment.capv.cluster.x-k8s.io

The last two enforce the fixed-IP rollout rules described in Provider Requirements. Where they are not registered, the same values are still required by the slot model — a non-compliant rollout stalls for want of a free slot instead of being rejected up front. The vspheremachineconfigpool webhook also intercepts DELETE, which is what turns the deletion of a pool that still owns InUse slots into an immediate rejection rather than a resource held in Terminating.

Check the management-cluster permissions if your platform restricts ClusterRoles. Beyond the upstream set, the provider chart grants capv-manager-role the following:

  • Read access to clustermodules, moduleconfigs, and moduleplugins in cluster.alauda.io.
  • Write access to moduleinfoes and read access to moduleinfoes/status.
  • Read access to clusters in clusterregistry.k8s.io and platform.tkestack.io.
  • patch and update on controlplane.cluster.x-k8s.io/kubeadmcontrolplanes.
  • Access to vspheremachineconfigpools.

The kubeadmcontrolplanes patch permission is what lets the provider set the CoreDNS and kube-proxy skip annotations described in Provider-managed fields on your manifests.

Upgrading the Provider

Upgrade the provider the same way you installed it: upload the newer package and update the cluster plugin on the global cluster.

The upgrade also updates the provider CRDs. Some fields move between releases, but existing resources are migrated for you — do not hand-edit VSphereMachineConfigPool or any other provider object to match a new schema, and do not delete and recreate one. On its first reconcile after the upgrade the controller brings each object up to the current layout:

  • The control plane endpoint is not touched. An existing cluster leaves spec.controlPlaneLoadBalancer unset, which is what an external LoadBalancer looks like, and it stays that way: the API server address does not change and nothing is deployed into the workload cluster.
  • Per-disk state that used to be recorded on spec.configs[].persistentDisks[].volumePath and .diskUUID, and per-slot reclaim state under status.configStatuses[].reclaimStatus, are folded into status.persistentDiskStatuses[]. A reclaim that was in flight keeps its phase and task reference and continues, and reclaimStatus is then cleared so status has a single source of truth.
  • The deprecated spec fields are left in the object. Treat them as history and read status.persistentDiskStatuses[] instead.
  • The migration never overwrites a value that is already set, so it is safe to repeat. The disk part is status-only bookkeeping and runs before vCenter is contacted, so it completes even while vCenter is unreachable.
  • It covers the provider's own resources only. The rollout values the admission webhooks require on the Cluster API objects — maxSurge, replicas, maxUnavailable — are validated, never rewritten, so an existing KubeadmControlPlane or MachineDeployment that does not carry them still has to be corrected by hand. See Check the manifests against the admission rules.

Nothing in the workload cluster is touched, and no VM is replaced. Three consequences are worth knowing before you upgrade:

  • A cluster created on an earlier provider keeps its external LoadBalancer and cannot adopt the Self-built VIP. The endpoint is a deploy-time choice: it cannot be adjusted while the cluster runs, and an upgrade is not an opportunity to change it — filling in spec.controlPlaneLoadBalancer on a running cluster is rejected. See Control Plane Endpoint Modes.
  • A pool keeps the disk lists it was created with. spec.configs[].ephemeralDisks[] is new in this release, but the upgrade moves no existing disk into it: a cluster created on an earlier provider keeps /var/cpaas, /var/lib/containerd, and /var/lib/etcd as persistent disks, and a slot that is InUse or Released rejects the removal or the reshaping of a disk in either list. The persistent-plus-ephemeral split documented for new clusters is a creation-time layout. See Standard data disks.
  • Guest-side changes delivered through guestinfo user data, such as the pod-log relocation, reach a node only when it is replaced. See Relocating the pod log directory.

Confirm the result with the checks in Verify the installation, then read one pool back:

kubectl -n <namespace> get vspheremachineconfigpool <pool_name> \
  -o jsonpath='{range .status.persistentDiskStatuses[*]}{.hostname}{" "}{.name}{" "}{.phase}{"\n"}{end}'

Each disk that was already provisioned before the upgrade appears with a phase, and status.configStatuses[].reclaimStatus is gone. A disk on a slot that no Machine has ever allocated has no VMDK yet, so it has no record; that is expected rather than an incomplete migration.

Standby Management Clusters

While the etcd-sync ConfigMap exists in the kube-public namespace, the provider treats the management cluster as a disaster-recovery standby:

  • Objects owned by the reserved global cluster are still reconciled.
  • Every other object is skipped and requeued every 30 seconds.
  • An object whose cluster name cannot be determined is skipped as well.

Do not create or upgrade a workload cluster against a standby management cluster. The reconcilers no-op silently — no condition changes and no error appears on the resource, so the workflow looks stalled rather than blocked.

Check before you start:

kubectl -n kube-public get configmap etcd-sync -o name 2>/dev/null || echo "not standby"

A common mistake is to look in kube-system. The ConfigMap is in kube-public.

For the disaster-recovery workflow itself, see Global Cluster Disaster Recovery.

Next Steps

After the provider is installed, continue with: