Transactional Outbox Pattern: Reliable Event Delivery

Updated on
10 min read

The transactional outbox pattern helps application developers and system architects publish events without losing the connection between a database change and its notification. It addresses a common failure in event-driven systems: a service commits business data to a database, then fails before sending the corresponding message. This explainer covers the pattern’s transaction boundary, relay choices, duplicate handling, ordering limits, and a practical PostgreSQL design.

Why the Transactional Outbox Pattern Is Being Discussed

Microservices often update a database and notify other services through a broker such as Kafka, RabbitMQ, or a cloud event bus. These are separate systems, and a routine application operation cannot usually commit to both in one atomic transaction. As teams move more workflows to asynchronous events, the gap between those two writes becomes a reliability problem rather than an edge case.

The outbox is one of the core patterns for making that boundary explicit. It does not make the broker and database one distributed transaction; it records the intent to publish alongside the business change and lets a separate relay deliver that intent.

What Is the Transactional Outbox Pattern?

An outbox is a database table that stores events waiting to be published. When a service changes business data, it inserts the corresponding event record in the same local database transaction. A separate process, called a relay or publisher, reads committed outbox records and sends them to a message broker.

The important property is atomicity inside the database: either both the business update and its outbox record commit, or neither does. The broker may be temporarily unavailable, but the event remains in durable storage until the relay can retry it.

The transactional outbox pattern description on microservices.io outlines this database-plus-relay design and its delivery trade-offs. An outbox is not itself an event bus, event store, or guarantee of exactly-once effects. It is a way to preserve the publisher’s intent across a local transaction boundary.

The Problem: Database and Broker Dual Writes

Consider an order service that writes an order row and then publishes OrderPlaced. If it saves the row first and crashes before publishing, the order exists but inventory and billing never hear about it. Reversing the order creates the opposite failure: a message can be delivered even though the database transaction rolls back.

Retrying both operations does not remove the race. A retry may publish duplicates, and there is no safe ordering that covers a crash between independent commits. A relational database can guarantee atomicity for its own transaction, as described in the PostgreSQL transaction tutorial, but that guarantee does not extend to an unrelated broker.

The AWS transactional outbox guidance also calls out these dual-write failure modes. The practical objective is not to promise that a message is sent exactly once; it is to ensure that a committed business change has a durable event waiting to be delivered, then make retries safe.

How the Outbox Works

The pattern has three stages:

  1. Write business data and event intent together. The application updates its normal tables and inserts an outbox row in one database transaction.
  2. Relay committed rows. A polling worker or change-data-capture (CDC) connector reads the outbox only after the transaction commits.
  3. Publish and record progress. The relay sends an event, then marks the row delivered or advances its CDC position. Consumers handle redelivery safely.

For example, an order transaction might persist an order with status placed and an OrderPlaced event containing the order ID and relevant payload. If the transaction rolls back, neither is visible. If it commits but the broker is down, the relay can retry later.

There is still a small failure window after broker acceptance but before the relay records success. The relay may send the event again after restarting. That is why outbox delivery is normally at least once, and why a stable event ID and idempotent consumers matter.

The outbox table should contain only the information required by the relay and consumers: event ID, aggregate type and ID, event type, payload, creation time, and delivery or lease metadata. Avoid treating it as an unbounded archive. Define retention and cleanup around delivery, replay, and audit requirements.

Relay Choices and Trade-offs

Consideration Polling publisher CDC or transaction-log relay Direct dual write
Atomic link to business update Yes, when both rows share one database transaction Yes, when the connector reads the committed outbox change No; database and broker writes can diverge
Delivery latency Depends on polling interval and batch size Often low, based on transaction-log consumption Potentially low, but unreliable on partial failure
Operational needs Worker, query/index tuning, retries, cleanup Connector, log retention, offsets, schema and connector operations Fewer components initially, but recovery logic is application-specific
Duplicate handling Required after publish/mark crash window Required when the connector or consumer replays Required for retries, without solving lost updates
Ordering Must be deliberately preserved by the relay Depends on database log order, connector behavior, and broker partitioning No dependable cross-system order

A polling relay is straightforward to understand and can work well at moderate volume. CDC avoids repeatedly querying the table and can reduce delivery latency, but it introduces connector operations and database log-retention considerations. Debezium’s outbox event router documentation describes one way to map outbox-table changes into broker records. The right choice depends on throughput, latency targets, database support, and the team’s operational experience.

Key Components and Design Decisions

  • Outbox schema: Store a stable event ID, event type, aggregate key, payload, and a creation or sequence value. Keep sensitive fields out unless downstream consumers need them.
  • Relay: Poll rows with short, bounded batches or consume committed changes from the database log. Apply backoff for transient broker failures and expose permanent failures for investigation.
  • Delivery state: Polling relays can track a publish timestamp or lease. CDC relays generally track connector offsets and still need policies for replay and retained rows.
  • Consumer idempotency: A consumer may receive the same event more than once. Use the event ID as a deduplication key and commit the deduplication record with the consumer’s own database changes when possible.
  • Ordering: A broker key such as aggregate_id can keep events for one aggregate in the same partition, but it does not fix a relay that publishes those events out of order. If per-aggregate order matters, include a sequence or aggregate version and serialize publication for that key.
  • Retention and monitoring: Alert on oldest unpublished event, pending-row count, retry rate, and relay lag. Retain delivered rows only as long as replay, audit, or debugging needs justify.

Real-World Use Cases

An order service can commit an order and OrderPlaced event together, allowing inventory and payment consumers to proceed even if the broker is temporarily unreachable. A billing service can record a payment state and publish PaymentAuthorized for fulfillment without making its database transaction depend on every downstream service.

The same pattern works for account changes that update a search index, subscription changes that trigger provisioning, or transactional records that feed analytics. It is most useful when a database is the authoritative write store and other systems need to react asynchronously. It does not replace a saga for coordinating a multi-step business process, nor does it remove the need to define what happens when a downstream action fails permanently.

Practical Guide: A PostgreSQL Outbox

The following example uses PostgreSQL and a polling relay. The orders table represents business state; outbox_events stores event intent and simple lease metadata. A stable event_id lets publishers and consumers identify retries.

CREATE TABLE orders (
  id text PRIMARY KEY,
  status text NOT NULL
);

CREATE TABLE outbox_events (
  event_id uuid PRIMARY KEY DEFAULT gen_random_uuid(),
  aggregate_type text NOT NULL,
  aggregate_id text NOT NULL,
  aggregate_version bigint NOT NULL,
  event_type text NOT NULL,
  payload jsonb NOT NULL,
  created_at timestamptz NOT NULL DEFAULT clock_timestamp(),
  published_at timestamptz,
  lease_token text,
  lease_until timestamptz,
  attempts integer NOT NULL DEFAULT 0
);

CREATE INDEX outbox_pending_created
  ON outbox_events (created_at, event_id)
  WHERE published_at IS NULL;

Write the business row and its event in one transaction:

BEGIN;

INSERT INTO orders (id, status)
VALUES ('ord-1001', 'placed');

INSERT INTO outbox_events (
  aggregate_type, aggregate_id, aggregate_version, event_type, payload
)
VALUES (
  'order',
  'ord-1001',
  1,
  'OrderPlaced',
  '{"orderId":"ord-1001","status":"placed"}'::jsonb
);

COMMIT;

A polling worker can claim a batch with a short database transaction, commit the leases, publish the returned rows, and mark each successful event as published. SKIP LOCKED allows multiple workers to claim different rows without waiting on each other:

WITH claimable AS (
  SELECT event_id
  FROM outbox_events
  WHERE published_at IS NULL
    AND (lease_until IS NULL OR lease_until < clock_timestamp())
  ORDER BY created_at, event_id
  LIMIT 100
  FOR UPDATE SKIP LOCKED
)
UPDATE outbox_events AS event
SET lease_token = 'worker-1:batch-7',
    lease_until = clock_timestamp() + interval '30 seconds',
    attempts = attempts + 1
FROM claimable
WHERE event.event_id = claimable.event_id
RETURNING event.*;

Use a unique lease token for each claimed batch in a real relay. Commit the claim before making a network call so locks are not held while the broker responds. After a successful publish, mark delivery only if the row still has that lease token. If the process crashes after publish but before that update, the lease expires and the event can be sent again; consumers must deduplicate by event_id. If publish fails, release or allow the lease to expire and retry with bounded backoff.

Check the oldest pending event and backlog to spot a stalled relay:

SELECT count(*) AS pending_events,
       min(created_at) AS oldest_pending_at
FROM outbox_events
WHERE published_at IS NULL;

For production, add a cleanup policy for delivered rows, alerts on growing backlog and repeated failures, and a dead-letter or operator workflow for poison events. For CDC, configure the connector to capture the outbox table and map its columns to the broker’s key, headers, and payload; test schema changes and connector restarts before relying on it.

Common Misconceptions

“The outbox guarantees exactly-once delivery.”

It makes the database update and event intent atomic, but a relay can publish and crash before saving its progress. Duplicate delivery is expected; consumer-side idempotency prevents the duplicate from repeating the business effect.

“Using an outbox means events are delivered immediately.”

Polling interval, connector lag, broker availability, and retry backoff all affect latency. The event is durable after the database transaction commits, but delivery is asynchronous. Monitor end-to-end lag against the workflow’s requirements.

“The outbox removes the need for ordering rules.”

Rows can be claimed by multiple workers and published concurrently. Broker partitioning preserves order only within a partition and in the order records arrive there. If order per entity is required, design the relay around an entity key and enforce its sequence rather than relying on timestamps alone.

Changelog

  • 2026-10-09: First publication.
TBO Editorial

About the Author

TBO Editorial writes about the latest updates about products and services related to Technology, Business, Finance & Lifestyle. Do get in touch if you want to share any useful article with our community.