Redis Persistence and Replication: RDB, AOF, and Failover

Updated on
9 min read

Redis is often introduced as a fast cache, but production systems may also depend on it for sessions, queues, counters, and other live state. Redis persistence and replication address different risks: persistence helps reconstruct data after a restart, while replication keeps copies on other servers for read scaling and availability. This guide explains what each mechanism protects, what it cannot guarantee, and how to test the recovery path.

What Is Redis Persistence and Replication?

Redis is an in-memory data structure server. Commands usually operate on the server’s in-memory dataset, so a process restart or host failure can otherwise remove the data. Persistence writes enough state to storage to reload the dataset later. Replication sends changes from a primary server to one or more replicas.

Persistence and replication are complementary, not interchangeable. A primary and its replicas can all lose data after a faulty command, and a persisted file does not provide a live server to take over during an outage. Redis describes its persistence options and replication behavior separately because they solve different operational problems.

The Problem These Mechanisms Solve

RAM provides fast access but is not by itself a recovery plan. If a Redis process restarts without saved state, its dataset may be empty. A machine or availability-zone failure can also make a standalone server unreachable. These failures have different remedies:

  • Process restart or host loss: Load a valid snapshot or append-only log from persistent storage.
  • Server outage: Route clients to an available replica, using an appropriate failover mechanism.
  • Accidental deletion or corrupted data: Restore a known-good backup; a replica may copy the same mistake.
  • Read pressure: Send selected reads to replicas, while accounting for replication lag.

The design starts with recovery objectives. A recovery point objective (RPO) defines how much recent data the service may lose; a recovery time objective (RTO) defines how long it may take to restore service. Neither persistence nor replication should be enabled by habit without deciding which failures and loss windows matter.

How Redis Persistence and Replication Work

Redis offers two principal persistence formats: RDB snapshots and the append-only file (AOF). Replication is a separate stream of updates from a primary to its replicas.

Mechanism What it stores or sends Main benefit Main limitation
RDB snapshot A point-in-time representation of the dataset Compact files and efficient restarts or backups Writes since the last completed snapshot can be lost
AOF A log of write operations that can rebuild the dataset A configurable, smaller recovery window More disk I/O and log-rewrite management
Replication Updates sent from a primary to replicas Additional copies for read scaling and failover Normally asynchronous; replicas can lag or miss recent writes
Sentinel or Cluster Control and routing around replicated servers Automated monitoring and failover or partitioning Does not replace persistence or independent backups

RDB snapshots

An RDB file captures the dataset at a point in time. Redis can create snapshots on configured save conditions or when an operator requests one. Background snapshotting lets the primary continue serving commands while a child process writes the file. The operating system’s copy-on-write behavior helps the child work from a consistent view, but concurrent writes can increase memory use while pages are copied.

RDB files are useful for backups and fast bulk recovery. The trade-off is the interval between snapshots: if the server fails after a snapshot, updates since that snapshot may not be present in the restored dataset. Snapshot success also depends on storage health and sufficient disk space.

AOF logging

With AOF enabled, Redis records write operations so it can replay them when loading state. The appendfsync setting controls when buffered log data is synchronized to storage. always requests synchronization for each write and can add latency; everysec is a common balance that can leave a short recent-write window at risk; no leaves synchronization timing to the operating system.

As the log grows, Redis can rewrite it into a more compact sequence that reconstructs the same dataset. Rewriting reduces replay time and storage use, but it is not a substitute for checking free disk space, monitoring rewrite failures, or preserving independent backups. When RDB and AOF are both enabled, Redis uses the AOF during restart because it generally contains a more complete recent history.

Replication and failover

A replica initially synchronizes with its primary, commonly by loading an RDB snapshot, then applies the continuing stream of updates. If a connection drops, partial resynchronization can send updates from the primary’s replication backlog when the replica’s offset is still available. Otherwise, the replica needs another full synchronization.

Replication is asynchronous by default. A primary may acknowledge a write before a replica has received or applied it, so a replica read can be stale and a primary failure can lose writes that had not propagated. Redis replication transfers data over TCP. TCP provides an ordered byte stream while a connection is working, as described in RFC 9293, but transport reliability does not make a Redis write durable or guarantee that a replica has applied it.

Redis Sentinel monitors non-clustered primary/replica deployments and can coordinate promotion of a replica when the primary is unavailable. Redis Cluster distributes keys across shards and can use replicas for shard failover. Neither mechanism changes the underlying consistency and loss trade-offs of asynchronous replication.

Components and Key Concepts

  • Primary: Accepts writes and normally serves authoritative reads for the replicated dataset.
  • Replica: Applies updates received from a primary. It can serve reads, but a read may not include the primary’s latest write.
  • Replication offset and backlog: Track the update stream and retain a bounded history for partial resynchronization after a brief disconnect.
  • RDB file and AOF: Recovery artifacts whose validity, retention, disk location, and restore procedures must be managed.
  • Sentinel or Cluster: Availability and topology mechanisms. Sentinel monitors a primary/replica group; Cluster partitions the keyspace and manages shard membership.
  • WAIT: A command that can ask Redis to wait until writes from the current client have reached a requested number of replicas. It reduces some replication uncertainty, but does not force replicas to fsync those writes or provide consensus-level durability.

Persistence files should be on storage that survives the process or container lifecycle. A container’s writable layer alone is not an adequate durable location. Also distinguish a live replica from a backup: replication mirrors current state, including unwanted changes, while a backup should provide a separate, retained restore point.

Real-World Use Cases

  • Disposable cache: Replication may help keep reads available, while persistence may be unnecessary if the cache can be rebuilt safely. Define the expected behavior during a cold restart.
  • Sessions and rate limits: Losing recent state can affect user sessions or enforcement. Select a persistence policy and recovery window that match the impact of missing entries.
  • Queues and streams: Confirm which Redis data structures and acknowledgment semantics the application relies on. AOF, replicas, and a tested restore procedure can reduce risk, but do not assume they provide exactly-once processing.
  • Read-heavy service: Replicas can absorb reads, but code must tolerate lag and stale results. Direct read-after-write paths to the primary when the application requires fresh state.
  • Availability-sensitive service: Sentinel or Cluster can help restore service after node failure, while persistence and independent backups address restart and recovery needs.

For general database recovery objectives and restore testing, see Database Backup and Recovery Strategies. For broader replication models across database systems, see Database Replication Patterns.

Getting Started: Configure and Verify Redis

For a local Docker experiment, create a named volume so the data directory survives container replacement. This enables AOF with a one-second synchronization policy and configures periodic snapshots:

docker volume create redis-lab-data
docker run -d --name redis-lab \
  -p 127.0.0.1:6379:6379 \
  -v redis-lab-data:/data \
  redis:7 \
  redis-server --appendonly yes --appendfsync everysec --save 60 1000

Write a key and inspect persistence status:

docker exec redis-lab redis-cli SET recovery:check "saved"
docker exec redis-lab redis-cli INFO persistence
docker exec redis-lab redis-cli CONFIG GET appendonly

Restart the container, then confirm the key is still present:

docker restart redis-lab
docker exec redis-lab redis-cli GET recovery:check

This checks a basic restart, not host-loss recovery or backup restoration. Do not publish an unauthenticated Redis instance to an untrusted network; this example binds the published port to localhost and is for development only.

To see asynchronous replication in a separate local lab, create a private Docker network and start a primary and replica:

docker network create redis-replication-lab
docker volume create redis-primary-data
docker run -d --name redis-primary --network redis-replication-lab \
  -p 127.0.0.1:6381:6379 \
  -v redis-primary-data:/data \
  redis:7 redis-server --appendonly yes --appendfsync everysec
docker run -d --name redis-replica --network redis-replication-lab \
  -p 127.0.0.1:6380:6379 \
  redis:7 redis-server --replicaof redis-primary 6379

Write through the primary, read from the replica, and inspect its link state:

docker exec redis-primary redis-cli SET replication:check "copied"
docker exec redis-replica redis-cli GET replication:check
docker exec redis-replica redis-cli INFO replication

The read may need a moment after the write because replication is asynchronous. A local two-container test demonstrates the data path; it does not test failover, network partitions, storage loss, or production security. Remove the lab containers and network when finished; keep or remove the named volume deliberately based on whether the data is needed.

Common Misconceptions

“A replica is a backup”

A replica follows the live primary. A mistaken delete or corrupted write can reach replicas as well. Keep separate backups with retention, protect them from the primary’s failure domain, and periodically restore them.

“Enabling AOF means no writes can be lost”

AOF’s loss window depends on the synchronization policy, operating system, storage device, and failure mode. Even always cannot protect against every hardware or operational failure. Define the required RPO and test the actual storage path.

“Failover makes Redis strongly consistent”

Failover changes which server is primary; it does not retroactively recover writes that never reached the promoted replica. Applications must account for stale replica reads and possible loss of recent asynchronous writes.

TBO Editorial

About the Author

TBO Editorial writes about the latest updates about products and services related to Technology, Business, Finance & Lifestyle. Do get in touch if you want to share any useful article with our community.