Conflict-Free Replicated Data Types (CRDTs) Explained
Conflict-free replicated data types (CRDTs) let multiple copies of data accept changes independently and later converge when those changes are exchanged. They are useful to developers building collaborative and offline-capable software, but convergence is not magic conflict resolution: the data model determines what happens when people make concurrent edits. This guide explains the guarantees, common structures, and operational costs so you can decide where a CRDT fits.
Why CRDTs Are Being Talked About
Applications increasingly let people edit the same data from several devices, including when a device is disconnected. A conventional client-server design can serialize every write through one authority, but that adds a network dependency to editing and can make offline writes difficult. Allowing devices to accept local changes improves responsiveness, yet those copies can diverge.
CRDTs provide a principled way to merge certain kinds of concurrent changes. They underpin many local-first and collaborative systems, from shared text and lists to counters and maps. The foundational CRDT research formalizes the conditions that allow replicas to converge despite different message timing. CRDTs are one choice within a broader local-first software architecture, not a requirement for every offline application.
What Is a CRDT?
A CRDT is a data structure designed so independently updated replicas can be combined consistently. If replicas start from compatible state, apply valid local changes, and eventually receive the relevant updates, they converge without requiring a central coordinator to order every change.
The structure encodes the merge rules. A counter can add contributions from different replicas; a set can preserve which additions a removal has observed; a shared text type can track inserted characters and their relationships. The application usually manipulates familiar values, while the CRDT implementation stores identifiers and causal metadata needed to merge them.
The CRDT project overview introduces the family of designs, while the Yjs documentation shows how a practical library exposes shared types. CRDTs are not a network protocol or a synchronization service: an application still needs to authenticate peers, persist data, and deliver updates.
The Problem CRDTs Solve
Suppose two people open the same list while online, then disconnect. One adds “coffee” and the other adds “tea.” A plain last-write-wins replacement of the whole list can keep one copy and discard the other. A manual conflict screen preserves both edits but asks users to resolve a routine merge.
With a suitable replicated set, each device can record its addition locally. Later, the replicas exchange updates and combine them. The result includes both elements regardless of which update arrived first. This is particularly useful when the system must accept local edits during a network partition, one of the availability and consistency choices explained in the CAP theorem guide.
That benefit has boundaries. A CRDT cannot decide whether a doctor should accept two mutually exclusive prescriptions, whether two reservations exceed available inventory, or which of two incompatible legal edits reflects intent. Such decisions need domain rules, coordination, or human review. The type can preserve and converge on operations; it cannot invent the correct product semantics.
How CRDTs Work
CRDT designs are commonly divided into state-based and operation-based approaches:
| Approach | What replicas exchange | Convergence requirement | Main trade-off |
|---|---|---|---|
| State-based (CvRDT) | Full state or a compact state delta | Merge is associative, commutative, and idempotent; state advances monotonically | A state can be larger to transmit, but duplicate or reordered delivery is manageable |
| Operation-based (CmRDT) | Operations such as “insert after this element” | Concurrent operations commute; delivery generally preserves required causal order | Updates can be small, but transport and causal delivery assumptions matter |
| Centralized serialization | Requests sent to one authority that chooses an order | The authority durably orders accepted writes | Simpler semantics, but writes depend on reachability and capacity of that authority |
In a state-based design, replicas merge state using a join operation. Associativity means grouping merges does not change the result; commutativity means the order does not matter; idempotence means applying the same state twice is harmless. Those properties make the merge resilient to duplicated and reordered state transfer. Many state-based designs use a semilattice: the state only moves forward in a partial order.
An operation-based design sends changes rather than repeatedly sending a whole data structure. Concurrent operations must be compatible, and the messaging layer must satisfy the delivery assumptions the type requires, often including causal ordering. A transport that drops an operation permanently or an implementation that applies an update outside its assumptions can still leave replicas inconsistent.
Neither design requires devices to communicate at edit time. Both do require eventual delivery or another recovery mechanism if every replica is expected to converge. The details vary by library: for example, Yjs documents shared types and update exchange, and Automerge’s documentation describes a CRDT document model and its synchronization options.
Key Concepts and Data Types
- Replica identity and causal history: A replica needs enough metadata to distinguish its changes from changes it has already seen. Counters, logical clocks, version vectors, or unique operation identifiers can track this history.
- Grow-only counter: Each replica records its own increment count; the value is the sum of those counts. Since only positive contributions are added, merging takes the maximum seen for each replica before summing.
- Positive-negative counter: A decrement is represented separately from increments, often as two grow-only counters. The visible value is increments minus decrements. The structure converges, but it does not enforce a business rule such as “balance must stay above zero.”
- Replicated set: An add-wins observed-remove set can remove only additions it has observed. If another replica concurrently adds the same value, that unseen addition can remain after merge. Other variants choose different add/remove precedence, so the exact semantics matter.
- Sequence or text: Concurrent inserts need stable identities and a deterministic order. Deletions must be represented so replicas agree which element was removed; tombstones or equivalent metadata may remain until safe compaction is possible.
- Map or document: A map composes other replicated types. Per-field rules determine whether concurrent assignments preserve, select, or expose competing values. “JSON-shaped” alone does not mean that a document is safe to merge.
A standard patch format is not automatically a CRDT. The JSON Merge Patch RFC defines how a patch modifies a JSON document, but it does not define how independently produced concurrent patches converge. A synchronization design must choose data types with merge semantics that match the application.
Real-World Use Cases
CRDTs work well when local responsiveness and automatic convergence are more important than a single globally ordered edit history:
- Collaborative editors: A sequence CRDT can merge text insertions and deletions from participants editing concurrently.
- Offline note-taking and task lists: A device can accept changes without a connection, then sync later. The note-taking synchronization architecture guide discusses where merge policies fit into an application’s sync pipeline.
- Shared collections and preferences: Maps and sets can merge independent field changes or membership updates from multiple devices.
- Presence and counters: Some CRDTs suit transient collaboration state or aggregations where duplicate delivery and temporary divergence are expected.
They are less suitable when correctness depends on a globally serialized invariant, such as enforcing a strict inventory limit or allocating a unique identifier from a bounded pool. Those workflows generally need a trusted authority, reservations, or another coordination protocol. A system can use CRDTs for documents while using transactions or a leader-based service for critical account or inventory changes.
Getting Started with a CRDT
Before choosing a library, define what concurrent edits should mean for each field. Decide whether removals win over concurrent additions, whether simultaneous text insertions both remain, and how long old replicas may stay offline. Then test duplicate delivery, reordered updates, reconnects, and storage recovery.
The following small Node.js example uses Yjs to merge independent updates to a shared map. It demonstrates data convergence in memory; it does not provide persistence, authentication, or a network transport.
npm install yjs
Save this as crdt-demo.mjs and run it with node crdt-demo.mjs:
import * as Y from 'yjs'
const left = new Y.Doc()
const right = new Y.Doc()
left.getMap('settings').set('theme', 'dark')
right.getMap('settings').set('fontSize', 16)
const leftUpdate = Y.encodeStateAsUpdate(left)
const rightUpdate = Y.encodeStateAsUpdate(right)
Y.applyUpdate(left, rightUpdate)
Y.applyUpdate(right, leftUpdate)
const snapshot = (doc) => Object.fromEntries(
[...doc.getMap('settings').entries()].sort(([a], [b]) => a.localeCompare(b)),
)
const leftState = snapshot(left)
const rightState = snapshot(right)
if (JSON.stringify(leftState) !== JSON.stringify(rightState)) {
throw new Error('Replicas did not converge')
}
console.log(leftState)
Both replicas should contain fontSize: 16 and theme: 'dark'. Production software must also decide how updates are stored and exchanged. Yjs provides separate providers and persistence integrations; a local-first design still needs application-level identity, authorization, backup, schema evolution, and recovery policies. Do not expose a demo relay or unauthenticated sync endpoint as a production service.
Measure update size, memory use, merge time, and metadata growth with realistic documents and prolonged offline periods. Test two clients making the same kind of edit concurrently, not just independent field updates. For critical values, verify the invariant directly instead of assuming that successful convergence implies correctness.
Common Misconceptions
“Conflict-free means users cannot make conflicting edits.”
People can still make changes that conflict in meaning. A CRDT guarantees defined merge behavior and convergence for a modeled type; it does not guarantee that the merged result matches user intent.
“CRDTs are always eventually consistent and therefore always available.”
An application can choose to accept local writes during disconnection, but a particular library or product may impose additional checks or require a server. The distributed failure guide explains why availability during a partition depends on system policy, not just on the data structure.
“A CRDT removes the need for coordination or operational care.”
Some types avoid coordination for specific updates, but authentication, permissions, schema changes, compaction, backup, and business invariants remain. Metadata may grow with history or replica count, and deleting that history safely can require knowing which replicas have observed it.
Related Articles
- Local-First Software Architecture explains the broader offline-first application model.
- CAP Theorem Explained describes availability and consistency trade-offs during partitions.
- Note-Taking App Synchronization Architecture covers sync pipelines and conflict-resolution choices.
- Distributed System Failures: Partitions Explained covers partial failures and recovery assumptions.
Changelog
- Initial publication: explains CRDT convergence, common replicated types, and their practical trade-offs.

