Creating Clusters on Bare Metal
This document explains how to create Kubernetes clusters on physical servers using the bare-metal provider. The workflow is YAML-only — there is no Fleet Essentials UI for bare-metal clusters at this time.
TOC
Prerequisites1. Required Plugin Installation2. OS Images Imported and Image Catalog Populated3. Network Connectivity4. Host Time Synchronization5. TPM Decision6. Workload Cluster Image Registry7. Host Disk and Boot PreparationCluster Creation WorkflowResolving Placeholder ValuesStep 1: Build the SeedImage and Register HostsSize COS_STATE before the first installOptional: Prepare Managed Data DisksStep 2: CreateMachineInventoryPool ResourcesStep 3: Create the Control-Plane Cluster ResourcesStep 4: Deploy Worker NodesCluster VerificationUsing kubectlVerify the Control Plane EndpointExpected ResultsCommon Failure ModesNext StepsAppendixComplete KubeadmControlPlane ConfigurationPrerequisites
Before creating clusters, ensure all of the following prerequisites are met.
1. Required Plugin Installation
Install the following plugins on the global cluster:
- Alauda Container Platform Kubeadm Provider
- Alauda Container Platform Bare Metal Infrastructure Provider (umbrella chart that installs both the bare-metal manager and
elemental-operator)
See the Installation Guide for details.
2. OS Images Imported and Image Catalog Populated
Before creating any cluster resources, import both bare-metal OS images into the target platform registry, then add the target Kubernetes version to elemental-image-catalog. Neither step happens on its own: the provider package contains no OS image, and the provider chart creates the image catalog with no entries.
Importing only one image is not sufficient. Adding an image-catalog entry or a SeedImage reference does not import the corresponding image.
Import the images. The Bare Metal OS download for your ACP release is an archive of one OCI image layout that holds both images under the same tag; see Image Format by Provider. Run the import as root on a host that has containerd and can reach the platform registry, such as a global control-plane node; Alauda OS ships ctr but not skopeo. /var/lib/containerd needs free space of at least the archive size. Import the archive into a dedicated containerd namespace and list the image names it creates:
The output lists baremetal-base-image:<os-image-tag> and baremetal-base-image-iso:<os-image-tag>. Push both images and confirm that the registry lists the tag for each of them. --skip-verify requires --local:
The last command removes the local copies after both pushes succeed. For the registry of the global cluster, the admin password is stored in the cpaas-system/registry-admin Secret. In a disaster recovery pair, import the images into the registry of both global clusters; see Prepare the global Cluster as a Management Cluster.
After the import, <base-image> is <registry-address>/tkestack/baremetal-base-image:<os-image-tag> and <base-image-iso> is <registry-address>/tkestack/baremetal-base-image-iso:<os-image-tag>. Take both images from the same archive. Do not point cluster resources at a build registry, pair images from different archives, or substitute an independently built image.
Populate the image catalog. The elemental-image-catalog ConfigMap in cpaas-system maps Machine.spec.version to the elemental upgrade image used to (re)provision a node. The provider chart creates it with no entries, so no node can be provisioned until you add the target version. Patch the ConfigMap rather than replacing it, so that existing entries are kept, then check the result:
Every value used as Machine.spec.version (for both the control plane and worker MachineDeployment resources) must appear as a key in this ConfigMap, with the leading v preserved, and resolve to the imported base-image. The provider resolves the image at reprovision time by substituting the platform registry address for the registry portion of the entry. If a target version is missing, no reprovision plan is written and BaremetalMachine ends up in Failed / Reason=ImageCatalogMiss. Adding the entry afterwards does not clear that state; delete the affected Machine so that its owner recreates it.
3. Network Connectivity
Bare-metal hosts are registered, installed, and joined entirely over the network, so every path in the table below has to be open before the first host boots the SeedImage ISO.
The host firewall on an Alauda OS node is preconfigured and maintained by the platform — you do not manage it, and this table is not a node firewall rule set. See Managing Firewall Ports. What you do have to plan is the network between the hosts and the endpoints they reach: site firewalls, security groups, and router ACLs between the host subnet, the platform, and the registry.
443 alone is not enough for a workload host
elemental-register and elemental-system-agent both reach the platform through the platform access address, so registration and plan polling need only the platform HTTPS port. That is not the whole requirement. The same host still resolves names through DNS, pulls its operating system and platform images from the registry on 11443, joins the API server on the control-plane endpoint port, and exchanges cluster traffic with the other nodes. A firewall approval that covers only 443 produces hosts that register successfully and then fail during install or join, where the cause is much harder to see.
Scope of the table. These are the paths for workload-cluster hosts, which is what this page creates. Two neighbouring cases are documented elsewhere and are not covered above:
- Bootstrap. While
globalitself is still being installed, its hosts register against the temporary bootstrap endpoint and a different set of ports applies — see Bootstrap host network. globalhosts after handoff. Aglobalhost's system agent uses the Global control-plane endpoint rather than the platform ingress path, and in a DR pair it uses the local control-plane VIP on6443directly. See Endpoint and Identity.
A SeedImage built during the bootstrap phase points at the bootstrap endpoint and must not be reused after handoff; create new MachineRegistration and SeedImage objects on the active global cluster. A workload SeedImage created on the global cluster stays valid across a handoff, because its registration URL is the platform domain rather than a bootstrap address.
Control-plane endpoint.
- For
InternalSelf-built VIP, the VIP must live in the same Layer-2 broadcast domain as the control-plane node IPs. Thevridmust be unique in that domain, the network must allow VRRP and gratuitous ARP updates, and the node image must expose IPVS and allow Alive to setnet.ipv4.conf.all.arp_accept=1andnet.ipv4.vs.conntrack=1. - For
ExternalLoadBalancer, provision the listener and control-plane backends before cluster creation. Follow Plan the Control Plane Endpoint.
First-boot addressing.
If the target VM or physical host does not receive an address from DHCP while booted into the live ISO, configure the network manually from the host console before waiting for registration. Check the NetworkManager connection name first, then apply the site-specific address, gateway, and DNS values:
4. Host Time Synchronization
Every host that will become a cluster node — control plane and worker alike — must already carry the correct time when it boots the SeedImage ISO. The platform's cross-node requirement is a unified time zone on every cluster node and a time synchronization error between nodes of no more than 10 seconds; see Node Checks. That requirement describes the cluster's nodes, so it applies here, but nothing on this path verifies it for you: the quick-configuration script and node checks on that page target traditional-operating-system machines that the platform joins over SSH, and they never run against an Alauda OS host installed by elemental-operator.
Treat time synchronization as a gate before Step 1:
- Make the site NTP servers reachable from the host subnet, and open UDP
123as listed in Network Connectivity. A production cluster should not depend on public NTP. - Set the clock from firmware, not from an operating system. At this point the install disk has been wiped as described in Host Disk and Boot Preparation, so there is no OS on the host to run a command in. Correct the time in the BIOS/UEFI setup or through the BMC, and keep the hardware clock on UTC so that the time zone is applied consistently once Alauda OS is installed.
- Compare the hosts against each other, not only against a reference. Across all hosts intended for the same cluster, the spread must stay inside the 10-second requirement before any of them boots the ISO.
- After the nodes reach
Ready, confirm the result on each node withtimedatectl— it reports both the time zone and whether the clock is synchronized — andchronyc trackingfor the source and current offset.
The whole cluster PKI inherits the clock of the host that runs kubeadm init. kubeadm sets the certificate authority's NotBefore from that host's current time, every leaf certificate inherits that NotBefore, and each certificate's NotAfter is derived from it. A host running ahead therefore issues a PKI that the other nodes and any client treat as not yet valid until their own clocks catch up, and a host running behind issues one that expires earlier than the configured validity period implies. A server whose hardware clock has never been synchronized is typically wrong by minutes or hours rather than seconds, so the consequence is a control plane that will not come up at all — and it surfaces at control-plane bring-up, long after the registration step where the wrong clock was introduced.
A Day-2 chrony MachineConfig does not replace this
Configuring the Chrony Time Service writes /etc/chrony.conf to nodes that are already running and joined. It is the correct way to point a cluster at the site NTP servers for ongoing operation, and it is a different requirement from this one: it cannot correct the clock of a host that is still installing Alauda OS or running kubeadm. Deferring to it leaves the install and join window unprotected.
5. TPM Decision
Set MachineRegistration.spec.config.elemental.registration.emulate-tpm from whether the host exposes a real hardware TPM (/dev/tpm0), not from whether it is physical or virtual:
emulate-tpm: false(or omit the field) — only when the host has a working hardware TPM, soelemental-registercan use it forauth: tpm.emulate-tpm: truewithemulated-tpm-seed: -1— for any host without a hardware TPM. This includes both virtual machines and physical servers that ship without a TPM module (for example a Dell R620). On such a host,auth: tpmcombined withemulate-tpm: falsemakeselemental-registerfail TPM attestation and never send the registration request, soelemental-operatorrecords zero registration POSTs.
6. Workload Cluster Image Registry
Bringing the cluster up needs no registry credential. Bare-metal hosts install the operating system and join the cluster using the platform registry address alone, so you can reach a cluster whose Nodes are all Ready without configuring one.
A credential matters at the next step — installing platform components on the new cluster. Those components pull their images from a registry, and a registry that requires authentication needs credentials the cluster can use.
By default the cluster uses the registry of the global cluster, backed by the public-registry-credential Secret when that registry is authenticated. To have it pull platform component images from a dedicated registry instead, follow Choose the Image Registry for a Workload Cluster.
The credential can wait, but the choice of registry cannot. A cluster is bound to its registry when it is created, and an existing cluster cannot be moved to another one. Decide before you apply the manifests below, even if you do not install platform components until later.
Neither choice affects the operating system images described above. Hosts pull base-image and base-image-iso from the platform registry during elemental install and every reprovision, before any cluster-level binding exists.
7. Host Disk and Boot Preparation
Classify every disk and virtual disk (VD) on each host before you boot the SeedImage ISO. Record which device is the OS install target and which devices, if any, must retain application data:
- Wipe the selected OS install disk and any other obsolete boot disks so that no bootable previous operating system remains. The host must boot the ISO, not an old on-disk OS. Set
MachineRegistration.spec.config.elemental.install.deviceto the exact OS disk. - Do not wipe a data disk that you intend to manage with the
Adoptpolicy. Back it up independently, record its stable ID and filesystem UUID, and make sure it cannot be selected as the install device. - A disk intended for
InitializeIfBlankmust be disposable and objectively blank. Formatting still requires a separate initialization approval after the host registers; the registration manifest does not authorize it. - Remove residual Elemental partition labels —
COS_STATE,COS_PERSISTENT,COS_OEM,COS_RECOVERY— from all disks, not only the intended install disk. Elemental resolves partitions by label (blkid -L COS_STATE); a stale label left on a second disk (for example from a previous install or a leftover multipath member) makes it resolve the wrong device and the reprovision snapshotter fails. - On hosts where the same disk can appear through more than one path (multipath), make sure none of those paths exposes a disk that still carries a
COS_*label; a residual label on any path can be resolved ahead of the intended install disk. Clean the disk rather than disabling multipath — the image keeps it enabled for network-attached storage boot. - On a host with multiple disks, do not leave
install.deviceempty and do not use/dev/sdaas a persistent identity. Follow Select a Fixed System Disk on Multi-Disk Bare-Metal Hosts to reuse one ISO while selecting each host's system disk by WWN. - Set the boot order so the host boots the virtual CD / ISO first.
Use only stable storage identities reported by the registered inventory. Linux paths such as /dev/sdX, /dev/vdX, /dev/nvmeXnY, and /dev/mapper/mpathX are runtime paths and must not be placed in MachineInventory.spec.storage.
Cluster Creation Workflow
When using YAML, import the paired OS images first, then proceed through five required resource steps, with an optional storage-preparation step before pool membership. Every Kubernetes resource must be applied in the cpaas-system namespace.
Important Namespace Requirement
All bare-metal resources must be applied in the cpaas-system namespace. The provider and elemental-operator only reconcile objects in that namespace.
Workload Cluster Naming
The workload cluster-name must not be global. That name is reserved for the global cluster, and reusing it causes the workload cluster's resources to collide with global cluster resources in cpaas-system. As a convention, keep the CAPI Cluster and BaremetalCluster named exactly <cluster-name>, and prefix dependent resources (KubeadmControlPlane, KubeadmConfigTemplate, MachineDeployment, machine templates, pools, registrations) with <cluster-name>-.
Resolving Placeholder Values
The example manifests below use <placeholder> syntax for environment-specific values:
Step 1: Build the SeedImage and Register Hosts
Create a MachineRegistration that describes the registration URL and first-install cloud-config, and a SeedImage that points elemental-operator at the matching ISO base image.
Set SeedImage.spec.baseImage to the final target-platform registry reference of the imported base-image-iso. It must be paired with the imported base-image used by the image catalog for later reprovisioning: take both images from the same OS archive, never from different archives.
Size COS_STATE before the first install
The state partition holds the running system plus every retained snapshot. The Elemental default of 8192 MiB is not enough: snapshots of one image share extents, but an upgrade to a different image has to hold two unrelated systems of roughly 3.5 GiB at once and fails part-way through with no space left on device, leaving a partial snapshot behind on every retry. Every cluster upgrade goes through elemental upgrade, and partition sizes are fixed at install time, so this only surfaces later — on a cluster already in production that can no longer be repartitioned without a reinstall.
Size it to at least 20480 MiB (20 GiB) at install time.
The fragment must sit in SeedImage.spec.cloud-config under stages.boot. Putting it in MachineRegistration.spec.config.cloud-config, or under any other stage, leaves an 8 GiB partition and reports no error.
Verify the result on the first installed host before rolling out the rest of the fleet:
Do not add /etc/resolv.conf to SeedImage.spec.cloud-config. Keep the ISO generic, and put site-specific resolver configuration in MachineRegistration.spec.config.cloud-config only when the first-boot registration path needs it. During the tested Global deployment flow, the ISO did not carry resolver files; node DNS was configured later by the kubeadm bootstrap data.
install.device, install.eject-cd, and install.reboot are intentional. If the target disk is omitted, elemental install can select an unintended device. On multi-disk physical hosts, use the fixed system disk workflow instead of a changing /dev/sdX name. If eject-cd or reboot is false, a host can remain in the live environment after the first install and never become usable inventory for Cluster API.
Apply the manifest and wait for the SeedImage build to finish:
Do not use status.state as the success gate; the current controller does not populate it. Continue only when SeedImageReady=True has reason SeedImageBuildSuccess and both status.downloadURL and status.checksumURL are non-empty. A SeedImage whose download lifetime has expired can still report SeedImageReady=True, but its reason changes and its URLs are cleared.
Boot every target host from this ISO. elemental-register runs first (creates the MachineInventory and uploads observedNetwork), then elemental install writes the on-disk OS. After install completes, the host stays available for plan execution. If the live ISO environment has no DHCP address, configure NetworkManager manually on the host console as described in Network Connectivity before waiting for the MachineInventory.
Confirm registration:
Every inventory you intend to use must:
- Show
Ready=True. - Have a non-empty
status.plan.secretRef.name. - Have a
spec.observedNetworkthat matches the host's expected NIC (only required when you want the install-time IP to survive across reprovisions). - When managed data disks are required, run an observer-capable OS image, report a fresh
status.observedStorage, and reachstatus.storage.phase=Preparedbefore pool allocation.
Record the exact MachineInventory names — they are referenced by name in the next step.
Optional: Prepare Managed Data Disks
Managed storage belongs to the long-lived MachineInventory, not to a CAPI Machine or BaremetalMachineTemplate. Configure it while the inventory is unallocated and before adding that inventory to a production pool.
Follow Manage Data Disks on Bare-Metal Hosts to:
- Verify
elemental-storage-observer.serviceand inspectstatus.observedStorage. - Select each device by a stable ID and choose
AdoptorInitializeIfBlank. - Use the matching
storagectlbinary to validate and atomically patch the declaration and, when required, its initialization approval. - Wait for
StoragePrepared=True/AllRequiredVolumesPreparedandstatus.storage.phase=Prepared.
If no disks should be managed, leave spec.storage absent or set volumes: []. The storage controller converges that inventory to Unmanaged with StoragePrepared=True/NoManagedVolumes and performs no disk operation.
Step 2: Create MachineInventoryPool Resources
Create one pool per role. The pool reconciler validates that every member exists, computes capacity counters, and writes the baremetal.alauda.io/pool=<pool-name> annotation onto the inventory.
Key parameters:
Apply and verify:
A healthy pool reports Ready=True, total = len(spec.machineInventories), and available = total - allocated - preparing - reprovisioning - unavailable. Inventories listed in spec.machineInventories that fail validation (missing, plan secret missing, Ready=False, or storage not prepared) raise the pool's unavailable counter and surface in the MembersValid condition. A non-empty storage declaration is an allocation gate: the inventory is not Available to CAPI until Prepare succeeds against a fresh observation.
Size the control-plane pool to at least KubeadmControlPlane.spec.replicas. Size the worker pool to at least MachineDeployment.spec.replicas. For rolling upgrades the pool must hold the entire replica count — the provider uses delete-then-add semantics from the same pool, never both at once.
Step 3: Create the Control-Plane Cluster Resources
Create the BaremetalCluster (declares the selected control-plane endpoint mode), the control-plane BaremetalMachineTemplate (points at the control-plane pool), the KubeadmControlPlane (replicas + kubeadm config), and the CAPI Cluster.
The Bare Metal API supports both endpoint modes in the same workflow:
Internal: the provider deploys Alive and reconciles the VIP and control-plane backend membership.External: the provider skips Alive; the load balancer owner maintains the listener, health check, and backend membership.
There is no legacy provider-version tab because Bare Metal has no earlier provider release to preserve. Select the mode with <control-plane-load-balancer-type> in the manifest below.
Use the Kubernetes images built into the OS image
The supported bare-metal OS image preloads the kubeadm control-plane, CoreDNS, and etcd images under cloud.alauda.io/alauda. Keep that repository in the KCP configuration and use the component tags from the OS Support Matrix. Do not replace it with <registry-address>/tkestack: that produces image references that do not match the images preloaded in the OS. The OS configures its built-in pause image through containerd; do not add pod-infra-container-image to the KCP kubelet arguments.
Full Configuration Reference
The example below uses a minimal KubeadmControlPlane. For the full hardening profile recommended in production — admission, audit, kubelet patches, encryption provider — see Complete KubeadmControlPlane Configuration in the Appendix.
Cluster annotations. The bare-metal provider relies on a small set of Cluster annotations during reconcile. Authoritative ones the operator must set:
BaremetalCluster parameters:
Apply and watch:
Each new control-plane BaremetalMachine advances Pending → Allocated → Reprovisioning → Running. Watch:
BaremetalMachine.status.machineInventoryRef.name— which inventory was picked.BaremetalMachine.status.planSecretRef.name— plan secret being driven. The secret carries the annotationbaremetal.alauda.io/plan.type=reprovision.MachineInventory.status.plan.state—Appliedonce the host completescloud-init clean,elemental upgrade, reboot, andkubeadm init/join.BaremetalCluster.status.conditions[EndpointReady]— true once the configured control-plane endpoint is reachable.
The bare-metal provider does not support single-node control planes. Provision at least three control-plane replicas (KubeadmControlPlane.spec.replicas: 3) so that etcd retains quorum. In Internal mode, the same nodes also participate in Alive VIP arbitration.
Step 4: Deploy Worker Nodes
After the control plane is Ready, create the worker BaremetalMachineTemplate, the worker KubeadmConfigTemplate, and the MachineDeployment. The full worker YAML and parameter table are in Managing Nodes on Bare Metal → Worker Node Deployment.
Cluster Verification
Using kubectl
Verify the Control Plane Endpoint
Confirm the configured mode and endpoint:
For Internal, verify Alive and the node prerequisites:
Run the following command on each control-plane node, not on the management cluster:
For External, verify the load balancer frontend and each backend according to the External LoadBalancer contract. Confirm that the load balancer owner has registered every current control-plane node and that HTTPS /healthz returns HTTP 200 for each backend.
Expected Results
A successfully created cluster shows:
Cluster.status.conditions[Ready]=True.KubeadmControlPlanereplicas allReady.- Every
BaremetalMachine.status.phase=Runningandstatus.ready=true. - Every used
MachineInventory.status.plan.state=Appliedwith the annotationbaremetal.alauda.io/plan.type=reprovisionon its plan secret (an annotation, not a label —-o jsonpath='{.metadata.annotations}'). - Every used inventory with non-empty
spec.storage.volumes[]reportsstatus.storage.phase=Active,StoragePrepared=True, andStorageActive=True; the correspondingBaremetalMachinereportsStorageReady=True/AllVolumesReady. MachineInventoryPool.statussatisfiesavailable + allocated + preparing + reprovisioning + unavailable = total.- Kubernetes Nodes Ready.
- For
Internal, Alive Pods are Ready and the API is reachable through the Self-built VIP. - For
External, the load balancer frontend is reachable and its backend list matches the current control-plane nodes.
Common Failure Modes
For the full operator-side state machine reference (every condition reason and recovery action), see Provider Overview → clean / reprovision plans.
Next Steps
After creating a cluster:
Appendix
Complete KubeadmControlPlane Configuration
The hardened configuration recommended for production bare-metal clusters — admission control, audit policy, kubelet patches, encryption provider, and IPv6 bind addresses. Substitute the placeholders from the table in Resolving Placeholder Values.
imagePullCredentialsVerificationPolicy: NeverVerify is required only starting with Kubernetes 1.35. Omit this parameter when creating a cluster with Kubernetes 1.34 or earlier.
Worker bootstrap is symmetric — see Managing Nodes on Bare Metal → Bootstrap Template for the worker KubeadmConfigTemplate.