The Transactional Outbox: Making Events Reliable

The problem
A service saves an order, then publishes an order.created event. If the process crashes between the two steps, the order exists but nobody hears about it. Reverse the order and you can announce an event for data that was never saved. This is the dual-write problem, and retries do not fix it.
The pattern
Instead of publishing directly, write the event to an outbox table in the same database transaction as the business data. Either both are committed or neither is. A separate relay process reads unpublished rows, delivers them, and marks them as sent.
- Atomic write: the event exists if, and only if, the data does.
- At-least-once delivery: the relay may deliver an event twice after a crash, so consumers must be idempotent.
- Ordering: preserve order per aggregate (for example per order id) rather than globally.
What to watch in production
- Relay lag: alert on the age of the oldest unpublished row, not only on errors.
- Poison events: cap retries, then park the event for inspection instead of blocking the queue.
- Cleanup: delete or archive delivered rows so the table stays small.
- Tracing: carry a trace id from the request into the event and on to the webhook, so one action can be followed end to end.
How we use it
Rumuze Core is a NestJS API built around this pattern: domain events go through an EventBus into a transactional outbox, and from there to a webhook engine for outbound delivery and to Socket.IO for realtime fan-out through Redis. Inbound webhooks are verified before they are accepted. The same idea is worth using in any system where a missed event means a support ticket.
Enjoyed this article? Share it with your network.