Set Up Disaster Recovery Replication
Disaster Recovery (spec.activeRedis.mode: peerof) is the hot-standby topology: an upstream instance fans out directed replication links to one or more downstream instances, each ready to be promoted if the upstream datacenter is lost. Setting it up involves two parts — enabling cross-datacenter replication on each Redis instance, and establishing a connection from the downstream instance to the upstream instance.
For the alternative topology, in which every datacenter accepts writes, see Set Up Active-Active Replication.
In a disaster recovery group, the role of each Redis instance is not fixed; it can be either an upstream or a downstream. The downstream does not restrict data writing, but data written to the downstream will be overwritten by data synchronized from the upstream or cause synchronization failure due to data type conflicts.
TOC
Choose the procedure for your Redis versionRedis 7.2 — new moduleStep 1: Create the peer-auth credential in each datacenterA dedicated account (recommended)The instance's default account (fallback)Step 2: Enable replication on the upstream instanceStep 3: Expose the upstream proxyStep 4: Enable replication on the downstream instanceStep 5: Create the connection on the downstream sideStep 6: VerifyRedis 6.0 — legacy moduleUpstream sideUse LoadBalancer as the access address for the upstream ProxyDownstream sidePre-flight inspectionRemoving a connectionCredential rotationChoose the procedure for your Redis version
The setup procedure differs between the two module generations. Follow the section that matches your Redis version.
Replication between Redis 6.0 and Redis 7.2 is not supported. See Module generations and version applicability.
Redis 7.2 — new module
The new module requires a peer-auth credential: every instance in the replication group binds a RedisUser through spec.activeRedis.redisUserRef. The operator pushes that credential to every node, and both inbound and outbound peer links authenticate with it. Enabling replication without a binding is rejected at admission, and the module rejects all inbound peer replication.
The module accepts an inbound peer only when the credential it presents matches the local peer-auth credential. Every member of the group therefore has to carry the same username and the same password value. Wiring a link between two datacenters whose credentials differ is refused at admission:
The bound account may be a custom one you create for the purpose, or the instance's own default account. A system account is always refused.
Step 1: Create the peer-auth credential in each datacenter
A dedicated account (recommended)
A custom RedisUser keeps the replication credential separate from the instance password: binding the default account widens what that password grants — anything holding it can act as a replication peer — and ties credential rotation to instance-password rotation.
Create the same account in every datacenter. Password Secrets are one-to-one with RedisUser objects, so each datacenter needs its own Secret object; what has to match across datacenters is the password value, not the Secret name.
The password is policy-checked at admission: 8 to 32 characters, and it must mix letters, digits and special characters. A value that does not qualify is rejected when the RedisUser is created, with password should consists of letters, number and special characters.
Wait for the account before binding it — the binding is rejected unless its phase is Success:
The instance's default account (fallback)
The instance's own default account may be bound instead. It is the account Redis 6.0 links have always used, so an existing group already carries the matching password on every member.
Every instance has a default-account RedisUser, but it holds a password Secret only when the instance was created with spec.passwordSecret. On an instance created without one the account exists and reports Success while carrying no credential, and binding it leaves the instance's ActiveRedis in Failed:
Set an instance password first, or bind a dedicated account as above.
Binding it requires every member of the group to carry the same default password. That is not automatic — each instance's password is whatever its own spec.passwordSecret holds — and each member needs its own Secret object: a Secret shared by several instances is contended by their RedisUser controllers and never settles. Set the same password value on every member before you wire the links.
The default-account RedisUser is named after the instance:
Check that it is provisioned in both datacenters:
The binding itself is set in Step 2, along with the rest of the replication configuration.
Step 2: Enable replication on the upstream instance
The range of serviceID is [0-15] and it must be unique within the replication group. It cannot be changed later, and replication cannot be turned off once enabled.
The instance performs a rolling restart of its data pods when replication is first enabled — the proxy is created and the module is delivered to each pod. The instance leaves Ready, and its ActiveRedis sits in Pending while the pods roll:
Wait for both to settle before continuing. The same applies to the downstream in Step 4, and Step 5 dials the upstream, so it needs the upstream settled too.
The instance must also satisfy the new module's constraints, which are enforced by admission:
customConfig.appendonlymust not beyes— the new module is RDB-only.customConfig.databasesmust be<= 16— the module refuses to load beyond 16 databases, and the pods would crash-loop.
Step 3: Expose the upstream proxy
After enabling replication, the proxy provides a NodePort access address by default. A single NodePort-based access address is not highly available. In a production environment, use a LoadBalancer address:
Each instance creates a proxy Service named activeredis-proxy-<instance-name>.
For Sentinel instances, the instance name needs to be prefixed with
rfr-.
The proxy Service of a peerof instance exposes a single port, 6379. The proxy carries the RESP control plane and the replication stream on it, telling the two apart per session, so the downstream needs to reach only that one address — the one you put in spec.addresses. No additional port is created in Disaster Recovery mode, and none has to be opened between datacenters.
Connection admission dials the endpoint it resolves for the replication stream and rejects the connection if it is unreachable. With spec.peerPort left unset, that is the address in spec.addresses[0] itself.
Step 4: Enable replication on the downstream instance
Create the downstream's peer-auth account as in Step 1 — same username, same password value as the upstream's. Then enable replication with a different serviceID:
This restarts the downstream's data pods too. Wait for the instance to be Ready before Step 5.
Step 5: Create the connection on the downstream side
The ActiveRedisConnection is created in the downstream cluster. spec.instance names the local (downstream) instance, and spec.addresses points at the upstream proxy's RESP endpoint.
The connection name becomes the module's peer name, which is stricter than Kubernetes naming:
- it must start with a letter —
1connis rejected; reset,clear,all, andlistare reserved words and are rejected.
An instance may hold at most one ActiveRedisConnection — its own upstream link. A one-upstream-to-many-downstreams fan-out is built by creating a connection in each downstream cluster, all pointing at the same upstream.
Step 6: Verify
PEERS counts upstream links and DPEERS downstream ones, so each side prints one and leaves the other blank — a zero count is not rendered. The upstream instance's fan-out is status.downstreamPeerCount.
Inspect the connection for per-shard synchronization detail:
status.shards[].status indicates the connection status of the shard, and status.shards[].syncStatus indicates the data synchronization status: PartialSync (incremental synchronization of Oplog) or FullSync (RDB is being synchronized). offset and opId advance as the stream is applied — a stalled pair on a Connected shard is the signal to look further.
status.shards[].versionState is the module's own verdict on the upstream's version — unknown, verified, or violating. A shard that is being held for version skew also carries a non-zero versionHoldSince.
status.credentialMode: peer-auth-global records that the module-side peer records hold no inline credential — outbound dials read the node-local peer-auth credential, so a rotation is picked up on the next re-dial.
The native Redis Sentinel mode does not have the concept of shards. Here, a primary-replica pair of Sentinel is abstracted as
shard 0to be compatible with the same data structure as the Redis cluster mode.
Redis 6.0 — legacy module
Redis 6.0 carries the frozen legacy module, kept for compatibility with existing instances. It supports Disaster Recovery only — no Active-Active mode, and no module-level peer authentication. Enabling replication on Redis 6.0 returns an admission warning recommending Redis 7.2. For new deployments, use Redis 7.2.
Upstream side
You need to create a Redis instance first.
Enable Disaster Recovery
The command returns an admission warning recommending Redis 7.2; the patch is applied. The instance then performs a rolling restart of its data pods to bring up the proxy, so wait for it to be Ready again before going on.
Use LoadBalancer as the access address for the upstream Proxy
After enabling disaster recovery support for the instance, the disaster recovery Proxy provides a NodePort access address by default. A single NodePort-based access address is not highly available. In a production environment, you can use a LoadBalancer address to provide access to the Proxy.
Each disaster recovery instance will create a Proxy Service with a name that follows this format: activeredis-proxy-<instance-name>.
For Sentinel instances, the instance name needs to be prefixed with
rfr-
Downstream side
Enable Disaster Recovery
The range of serviceID is [0-15]. In the same disaster recovery cluster, the serviceID cannot be repeated.
Wait for the instance and the peer-auth binding
This side rolls its data pods too. The connection created next is rejected while the instance is still rolling — the RedisUser it authenticates as has to report Success, and it does not while its nodes are coming back:
The operator adds the binding on a following reconcile, and the connection is rejected while it is still empty. Wait until the instance is Ready and this prints a name:
Configure Disaster Recovery Connection
Note to replace the upstream address.
A Redis 6.0 link authenticates as the instance's default account. The operator writes that down for you: the first time it reconciles an instance that has replication enabled and carries no binding, it sets spec.activeRedis.redisUserRef to the instance's own default-account RedisUser — drc-acl-<instance-name>-default for a cluster instance, rfr-acl-<instance-name>-default for a Sentinel one. Nothing is provisioned: that account already exists, and the password in it is the one the link uses. Both datacenters must therefore carry the same default password — the same value, each in its own Secret object.
There is no credential field on the connection. spec.secretName has been removed, and a manifest that still sets it is rejected with unknown field "spec.secretName".
Until v5.1, Disaster Recovery on Redis 6.0 was the only cross-datacenter replication this product offered, and the link credential lived on the connection: ActiveRedisConnection.spec.secretName named a Secret holding a password and nothing else.
There was never a username in it. The legacy module sends a one-argument AUTH <password> when no username is set, and the proxy resolves that to the instance's default account — so a Redis 6.0 link has always authenticated as default, whatever the Secret was called. Both datacenters have therefore always had to carry the same default password.
v5.1 removes spec.secretName from ActiveRedisConnection and ActiveRedisInspection, and takes the credential from spec.activeRedis.redisUserRef alone. Existing groups need no intervention: on its first reconcile after the upgrade the operator writes the binding on each replicating instance, pointing it at that instance's own default-account RedisUser. Nothing changes on the wire — same account, same password, read from somewhere else — and running links stay up across the upgrade.
What does change is your manifests. A stored ActiveRedisConnection keeps working, but a manifest that still sets secretName is rejected on the next apply, and the field is gone from kubectl explain. Drop it, and manage the credential through the instance's binding instead.
A custom RedisUser can be bound on Redis 6.0 as well — create it as shown in Step 1 of the Redis 7.2 procedure and set spec.activeRedis.redisUserRef to it in the same patch that enables replication. A binding that is already set is never replaced by the automatic one, and the link then authenticates as that account rather than as default.
What you do not get is rotation safety. A Redis 6.0 link carries its credential inside the module's peer record, frozen there when the link is wired, and the legacy path has no self-healing re-wire. Changing the password leaves the running link working until its session next drops — and then it stays down:
Recovering means deleting the ActiveRedisConnection and creating it again, which re-wires the peer with the current credential. Plan a rotation on Redis 6.0 as a short outage of the link, or upgrade to Redis 7.2, where the credential is read per dial instead.
Check Disaster Recovery Connection Status
Here status.shards[0].status indicates the connection status of the shard, and status.shards[0].syncStatus indicates the data synchronization status. The synchronization status can be PartialSync (indicating incremental synchronization of Oplog) or FullSync (indicating that RDB is being synchronized).
The native Redis Sentinel mode does not have the concept of shards. Here, a primary-replica pair of Sentinel is abstracted as
shard 0to be compatible with the same data structure as the Redis cluster mode.
Pre-flight inspection
Before a connection is accepted, the platform runs a set of pre-flight checks against the upstream. The Web Console exposes them through the Inspect button; they also run automatically during ActiveRedisConnection admission, and a failing check rejects the connection with the corresponding message. The checks are backed by the ActiveRedisInspection resource.
On Redis 7.2 the inspection dials with the peer-auth credential the data path will actually use, so a successful inspection also confirms that both datacenters carry the same credential. It additionally performs a plain TCP dial against the endpoint it resolves for the replication stream — the address in spec.addresses[0], unless spec.peerPort splits the two.
Removing a connection
Deleting an ActiveRedisConnection tears the link down according to its spec.teardownPolicy:
Decommission is the correct teardown for a datacenter that is gone for good — a detached-but-never-returning peer keeps holding garbage-collection floors on the survivors. It cannot be changed back to Detach:
The decision is recorded against the peer's serviceID and survives a restart. What it changes is what this instance keeps for that peer: its Oplog retention, tombstone garbage-collection and lag accounting for that serviceID are released and never re-armed. Nothing stops a connection to that peer from being created again, and one created straight away still resumes from whatever the two ends happen to hold — so re-creating the link is not a way to undo the decommission, and a peer that is meant to come back should be detached, not decommissioned.
Credential rotation
On Redis 7.2, peer links carry no inline credential: outbound dials read the node-local peer-auth credential that the operator re-pushes on every reconcile. To rotate, update the password Secret of each datacenter's peer-auth RedisUser in place, to the same new value, one datacenter at a time. The RedisUser controller replays the ACL, the operator re-pushes the credential, and the proxy accepts the new credential on the next rebuild. Established links stay healthy and keep replicating throughout — including the window in which the datacenters hold different values — and a link that drops after the rotation re-dials with the current credential.
That window is not a free-for-all, though: while the values differ, creating a link is refused, because admission dials the upstream with the credential the data path will use. Finish the rotation on every member before wiring anything new.
When a dedicated custom account is bound, rotating an instance's own password Secret is independent and does not affect peer links.
When the instance's default account is bound, its password Secret is the peer-auth Secret: rotating the instance password rotates the replication credential with it, and the rest of the group has to follow. That is the coupling a dedicated account avoids.
None of the above applies to Redis 6.0. Its links hold the credential inline in the module's peer record, fixed when the link was wired, and there is no self-healing re-wire on that path — so a rotated link keeps running only until its session next drops, then reports shard 0 status is Disconnected and stays there. Delete and re-create the ActiveRedisConnection to re-wire it with the current credential.