The setup
Our event store had grown organically since 2019. Sixty million rows, one table, every stream in the product writing to it. Index bloat was real, and a few hot streams pushed P99 query latency past 1.5s.
The shape of the problem: 94% of reads touched events from the last ninety days, while 100% of the index had to be kept warm for the other 6%.
We decided to split the table into a write-optimised "live" table and a read-optimised "archive" table, with a cutover nobody outside the team would notice.
Act one: dual writes
For two weeks every append went to both tables inside the same transaction. This is the boring part and the part you cannot skip.
Three things made it survivable:
A feature flag per stream family, so we could turn dual-writing off for one noisy stream without reverting the deploy.
A counter on both sides, exported to the dashboard, with an alert on divergence above zero.
A decision, made early, that the old table stayed the source of truth for the whole window — the new one was allowed to be wrong.
Act two: verifying parity
Backfilling the historical rows took eleven hours in batches of 5,000, throttled to keep replication lag under two seconds.
Verification was a nightly job comparing checksums per stream, not per row: hashing the ordered event ids of a stream is cheap and catches both missing and reordered events, which is the failure mode that actually scared us.
We found two mismatches. Both were streams written by a legacy importer that bypassed the repository layer — exactly the sort of thing a row count would never have shown.
Act three: the cutover
Reads moved over behind a flag, one stream family at a time, starting with the least critical. Each step stayed in place for at least 24 hours.
Total downtime: zero. Total drama: also surprisingly low — turns out RailsEventStore's linker is forgiving, and the dual-write window meant the rollback plan was "flip the flag back", not "restore a backup".
We kept the old table read-only for a further month before dropping it. Nobody needed it. We slept better anyway.
What we would tell you to skip
The rehearsal on staging. We did it, it passed, and it told us nothing — staging had 400,000 rows and none of the hot streams. The only useful rehearsal was the first stream family in production, behind the flag.
The elaborate rollback runbook. Eleven pages, three of which were about restoring a backup we were never going to restore, because the dual-write window meant the real rollback was a flag flip. We should have written one page and spent the time on the divergence alert.
What we would not skip: the throttle on the backfill. Replication lag is the thing that turns a long migration into an incident, and the batch job that respects it is thirty lines.
The tooling, briefly
The dual-write layer is a decorator around the repository: it takes the primary result, enqueues the secondary write in the same transaction, and increments two counters. About 200 lines including the flag lookup.
The verification job walks streams in id order, hashes the ordered event ids in pages of 10,000, and writes one row per stream per night. A mismatch pages nobody — it opens a task, because in nine cases out of ten the fix can wait for morning.
Neither is open source yet. Both will be, once we have removed the two places where they assume our table names.
The aftermath
P99 latency is now 130ms. The live table is 4× smaller and its indexes fit comfortably in memory. The archive table is queried about forty times a day and nobody cares how fast it is.
We will write up the dual-write tooling separately — it is about 200 lines and most of it is the divergence counter.
If you take one thing from this: verify with checksums over ordered ids, not row counts. The bug you are afraid of is ordering, and a count will never see it.
