VMware vSphere Provider Installation
This document explains how to install the VMware vSphere Infrastructure Provider and verify that the required Cluster API components are available on the global cluster. The provider version, CRD set, and admission webhooks on this page are validated against VMware vSphere Provider v1.0.17.
v1.0.17 is the provider packaging version. It bundles the CAPV controller, the vSphere CPI image, the provider chart, and the admission webhooks that enforce the fixed-IP rollout rules. For those rules, see Provider Requirements. For the change history, see Release Notes.
Prerequisites
Before installing the provider, ensure the following conditions are met:
- You can access the
globalcluster. - You can access Customer Portal to download provider packages.
- You know which platform release and provider package version you plan to install. This page targets
v1.0.17. - The vCenter account that clusters will reference through
VSphereCluster.spec.identityRefcan list, create, and set VM custom attributes. Grant the Global → Manage custom attributes and Global → Set custom attribute privileges. The provider requires them whenever aVSphereClusteror its CAPIClustercarries anyvsphere.cluster.x-k8s.io/*annotation, and a missing privilege fails VM reconciliation rather than skipping the metadata write. - The
globalcluster is the active management cluster, not a disaster-recovery standby. See Standby Management Clusters.
Steps
Download the required packages
To use VMware vSphere with Immutable Infrastructure, download the following cluster plugins from Customer Portal:
- Alauda Container Platform Kubeadm Provider
- Alauda Container Platform VMware vSphere Infrastructure Provider
Upload the packages
Upload the downloaded packages to the platform repository.
For the detailed upload procedure, refer to Upload Packages.
Install the provider on the global cluster
Install both cluster plugins on the global cluster.
For the detailed cluster-plugin procedure, refer to Cluster Plugin.
Verify the installation
Check the module state:
Both plugins must be present and in the Running state.
Check the registered CRDs:
Expected output includes all nine CRDs:
vsphereclusters.infrastructure.cluster.x-k8s.iovsphereclusteridentities.infrastructure.cluster.x-k8s.iovsphereclustertemplates.infrastructure.cluster.x-k8s.iovspheredeploymentzones.infrastructure.cluster.x-k8s.iovspherefailuredomains.infrastructure.cluster.x-k8s.iovspheremachines.infrastructure.cluster.x-k8s.iovspheremachineconfigpools.infrastructure.cluster.x-k8s.iovspheremachinetemplates.infrastructure.cluster.x-k8s.iovspherevms.infrastructure.cluster.x-k8s.io
Check the registered admission webhooks. The ValidatingWebhookConfiguration object name varies by packaging, so match on the webhook names instead:
Expected output includes:
validation.vspherecluster.infrastructure.cluster.x-k8s.iovalidation.vspheremachineconfigpool.infrastructure.cluster.x-k8s.iovalidation.vspheremachinetemplate.infrastructure.cluster.x-k8s.iovalidation.kubeadmcontrolplane.capv.cluster.x-k8s.iovalidation.machinedeployment.capv.cluster.x-k8s.io
The last two enforce the fixed-IP rollout rules described in Provider Requirements. Where they are not registered, the same values are still required by the slot model — a non-compliant rollout stalls for want of a free slot instead of being rejected up front. The vspheremachineconfigpool webhook also intercepts DELETE, which is what turns the deletion of a pool that still owns InUse slots into an immediate rejection rather than a resource held in Terminating.
Check the management-cluster permissions if your platform restricts ClusterRoles. Beyond the upstream set, the provider chart grants capv-manager-role the following:
- Read access to
clustermodules,moduleconfigs, andmodulepluginsincluster.alauda.io. - Write access to
moduleinfoesand read access tomoduleinfoes/status. - Read access to
clustersinclusterregistry.k8s.ioandplatform.tkestack.io. patchandupdateoncontrolplane.cluster.x-k8s.io/kubeadmcontrolplanes.- Access to
vspheremachineconfigpools.
The kubeadmcontrolplanes patch permission is what lets the provider set the CoreDNS and kube-proxy skip annotations described in Provider-managed fields on your manifests.
Upgrading the Provider
Upgrade the provider the same way you installed it: upload the newer package and update the cluster plugin on the global cluster.
The upgrade also updates the provider CRDs. Some fields move between releases, but existing resources are migrated for you — do not hand-edit VSphereMachineConfigPool or any other provider object to match a new schema, and do not delete and recreate one. On its first reconcile after the upgrade the controller brings each object up to the current layout:
- The control plane endpoint is not touched. An existing cluster leaves
spec.controlPlaneLoadBalancerunset, which is what an external LoadBalancer looks like, and it stays that way: the API server address does not change and nothing is deployed into the workload cluster. - Per-disk state that used to be recorded on
spec.configs[].persistentDisks[].volumePathand.diskUUID, and per-slot reclaim state understatus.configStatuses[].reclaimStatus, are folded intostatus.persistentDiskStatuses[]. A reclaim that was in flight keeps its phase and task reference and continues, andreclaimStatusis then cleared so status has a single source of truth. - The deprecated
specfields are left in the object. Treat them as history and readstatus.persistentDiskStatuses[]instead. - The migration never overwrites a value that is already set, so it is safe to repeat. The disk part is status-only bookkeeping and runs before vCenter is contacted, so it completes even while vCenter is unreachable.
- It covers the provider's own resources only. The rollout values the admission webhooks require on the Cluster API objects —
maxSurge,replicas,maxUnavailable— are validated, never rewritten, so an existingKubeadmControlPlaneorMachineDeploymentthat does not carry them still has to be corrected by hand. See Check the manifests against the admission rules.
Nothing in the workload cluster is touched, and no VM is replaced. Three consequences are worth knowing before you upgrade:
- A cluster created on an earlier provider keeps its external LoadBalancer and cannot adopt the Self-built VIP. The endpoint is a deploy-time choice: it cannot be adjusted while the cluster runs, and an upgrade is not an opportunity to change it — filling in
spec.controlPlaneLoadBalanceron a running cluster is rejected. See Control Plane Endpoint Modes. - A pool keeps the disk lists it was created with.
spec.configs[].ephemeralDisks[]is new in this release, but the upgrade moves no existing disk into it: a cluster created on an earlier provider keeps/var/cpaas,/var/lib/containerd, and/var/lib/etcdas persistent disks, and a slot that isInUseorReleasedrejects the removal or the reshaping of a disk in either list. The persistent-plus-ephemeral split documented for new clusters is a creation-time layout. See Standard data disks. - Guest-side changes delivered through
guestinfouser data, such as the pod-log relocation, reach a node only when it is replaced. See Relocating the pod log directory.
Confirm the result with the checks in Verify the installation, then read one pool back:
Each disk that was already provisioned before the upgrade appears with a phase, and status.configStatuses[].reclaimStatus is gone. A disk on a slot that no Machine has ever allocated has no VMDK yet, so it has no record; that is expected rather than an incomplete migration.
Standby Management Clusters
While the etcd-sync ConfigMap exists in the kube-public namespace, the provider treats the management cluster as a disaster-recovery standby:
- Objects owned by the reserved
globalcluster are still reconciled. - Every other object is skipped and requeued every 30 seconds.
- An object whose cluster name cannot be determined is skipped as well.
Do not create or upgrade a workload cluster against a standby management cluster. The reconcilers no-op silently — no condition changes and no error appears on the resource, so the workflow looks stalled rather than blocked.
Check before you start:
A common mistake is to look in kube-system. The ConfigMap is in kube-public.
For the disaster-recovery workflow itself, see Global Cluster Disaster Recovery.
Next Steps
After the provider is installed, continue with: