Sixteen workers, one send: deduplicating with a ledger
Why Rowfire sends only when an INSERT reports it inserted a row, and how that one rule keeps 16 racing workers from sending the same message 16 times.
- PostgreSQL
- engineering
- postgresql
- reliability
A rule that says once per account per week is a promise to a customer: they hear about a declined card once, not three times. Keeping that promise is easy with one worker and one poll. It gets hard when polls overlap and several workers run at once. Rowfire keeps it with one rule:
A send happens if, and only if,
claim()inserted a row.
Not “if the row looked new”. If the INSERT reported that it inserted.
The obvious version, and its window
The natural way to write deduplication is to look first, then write:
SELECT 1 FROM fire WHERE rule_name = $1 AND dedup_key = $2 AND dedup_bucket = $3;
-- nothing found, so:
INSERT INTO fire (rule_name, dedup_key, dedup_bucket) VALUES ($1, $2, $3);
-- and send
Between the SELECT and the INSERT, a second worker can run the same
SELECT, find nothing, and send too. With one worker you never see it. Under
load you do. We measured it by racing 16 workers at one row:
| strategy | workers that would send |
|---|---|
SELECT-then-INSERT |
16 |
INSERT … ON CONFLICT DO NOTHING RETURNING |
1 |
Let the unique index decide
The fire ledger has a unique constraint on (workspace_id, rule_name, dedup_key, dedup_bucket), and a claim is a single statement (simplified):
INSERT INTO fire (workspace_id, rule_name, dedup_key, dedup_bucket)
VALUES ($1, $2, $3, $4)
ON CONFLICT DO NOTHING
RETURNING id;
If it returns a row, this worker won and sends. If it returns nothing, another
worker already claimed this fire, and this one moves on. There is no window,
because PostgreSQL resolves the race under the unique index. workspace_id is
part of that key from the start, so going multi-tenant later doesn’t mean
rebuilding a unique index on the busiest table.
The ledger is keyed on the rule, not the trigger. If it were keyed on the trigger, the first rule to fire for an order would mark it fired, and a second rule on the same trigger would see “already” and never send: one automation silently suppressing another.
One constraint for every cadence
The dedup_bucket column is what lets one constraint serve every kind of rule:
| cadence | dedup_bucket |
|---|---|
| once ever | '' (empty) |
| once per day, week or month | the start of the period, e.g. 2026-08-03 |
| every nth time | the ordinal |
Once ever and once per period are pure uniqueness, so replaying a run that crashed halfway is harmless: the second attempt hits the same conflict. Every nth time is different because it counts. A run that died after counting but before sending would shift every later nth on retry, so the count only advances for an occurrence that is genuinely new.
Why overlapping polls are fine
Rows don’t always arrive in clock order: a row can land with an event_time
slightly in the past. So each poll deliberately re-scans a lookback window
behind its watermark, and the watermark only advances when a poll succeeds. A
failed poll covers its window again instead of skipping it.
Re-scanning would be dangerous with a weaker ledger. With this one, it is cheap. Re-polling the same 120-day window of the fixture database a second time gives:
matched=697 fired=0 already=697
Every row matched, none fired, and all 697 were recognized as already claimed.
Scaling out without coordination
Workers lease due triggers with SELECT … FOR UPDATE SKIP LOCKED, so two
workers never poll the same trigger at once. Every send still goes through
the claim, so even if they did, the ledger would let only one through. Adding
workers is docker compose up --scale worker=3, with no leader election or
locks to configure.
Shadow mode takes the same path: the message is rendered and a delivery is recorded either way. The only difference is whether the HTTP request is made. So what you watch in shadow is what live will do.
You can see the effect in the declined-card use case, where 211 declined charges become 105 tickets, or watch rules fire in the live demo.
Questions
Why not check whether a row was already sent, then send?
Between the check and the insert there is a window where another worker does the same check. In a test with 16 workers racing at one row, check-then-insert let all 16 send.
Do workers need a coordinator?
No. Workers lease due triggers with SKIP LOCKED, and the unique index on the ledger decides every race, so you can run more workers without any coordination between them.