Two doctors both went off call
Each transaction read a valid state and made a legal decision. The combination broke a rule neither of them could see.
A hospital rota has one rule: at least one doctor stays on call. Alice and Bob are both on call. Both decide, at roughly the same moment, to go off.
Each request opens a transaction, counts how many doctors are currently on call, sees 2, concludes that removing one is safe, and removes one. Both commit. Nobody is on call.
Two transactions, one invariant
The rule is that at least one doctor stays on call. Step through and watch who checks it.
Neither transaction did anything wrong on its own. Each read a consistent state, applied the rule correctly to what it read, and wrote one row. The constraint they broke exists only in the combination, and neither of them wrote to a row the other one read.
This is write skew, and it's permitted by the isolation level your ORM almost certainly uses.
Where the levels differ
Step through the demo again on read committed and watch transaction 2's second read. It returns 1, because transaction 1 committed in between. The same query, in the same transaction, returned 2 a moment earlier.
Switch to repeatable read and that second read returns 2. Repeatable read in Postgres is snapshot isolation: the transaction sees the database as it was when it started, for its whole lifetime, no matter what commits around it.
That fixes the inconsistent read. It does not fix the skew. Both transactions still commit, and there is still nobody on call.
What each level permits in Postgres
| anomaly | read committed | repeatable read | serializable |
|---|---|---|---|
| Dirty read | prevented | prevented | prevented |
| Non-repeatable read | allowed | prevented | prevented |
| Phantom read | allowed | prevented | prevented |
| Lost update | allowed | prevented | prevented |
| Write skew | allowed | allowed | prevented |
Two transactions read overlapping rows, each writes a different row, and the combination breaks a constraint neither one could see alone.
Write skew is the row that matters here, because it's the only one that survives snapshot isolation. Every other anomaly on that table is gone by the time you reach repeatable read, which is why repeatable read feels safe and mostly is.
Why snapshot isolation can't catch it
Snapshot isolation resolves conflicts by looking at writes. Two transactions that write the same row conflict, and one of them aborts. That's first-updater-wins, and it's why lost update disappears at repeatable read.
Write skew involves no write conflict. Transaction 1 writes Alice. Transaction 2 writes Bob. Different rows, no overlap, nothing to detect. The dependency runs through the reads: each transaction read a set of rows that the other one then modified.
Serializable in Postgres is serializable snapshot isolation, and it tracks exactly that. It records which rows each transaction read, watches for a cycle in the read/write dependency graph, and aborts one of the participants when it finds one. Run the demo on serializable and transaction 2 fails with a serialization error rather than committing.
The cost of serializable
Serializable transactions can fail for reasons that have nothing to do with your code being wrong. Two correct transactions can form a dependency cycle, and Postgres resolves it by killing one.
That means every serializable transaction needs a retry loop:
for (let attempt = 0; attempt < 3; attempt++) {
try {
return await db.transaction({ isolationLevel: 'serializable' }, run)
} catch (error) {
if (error.code !== '40001') throw error // serialization_failure
}
}Retrying is only safe if the transaction body has no side effects outside the database. A transaction that charges a card, sends an email, or writes to another service cannot simply be run again. That constraint, more than the performance overhead, is what usually decides against serializable.
The alternatives
You don't have to raise the level for the whole application. Write skew needs a shared thing to conflict on, and you can create one.
SELECT ... FOR UPDATE on the rows you read turns the read into a write conflict, which snapshot isolation already handles. In the rota case, locking both doctor rows before counting makes the second transaction wait, and it then reads 1 and refuses.
A constraint the database can check is better still, when the invariant fits one. A count maintained in a summary row with a CHECK (on_call_count >= 1) gives you a single row that both transactions must write, which turns the skew into an ordinary conflict.
What doesn't work is checking the invariant in application code before the write. That's exactly what both transactions in the demo did.
Finding it in your own code
The shape to look for is a read that decides whether a write is allowed, where the thing being read is not the thing being written.
Booking the last seat. Withdrawing when the balance permits. Enforcing a unique username without a unique index. Approving when there are enough approvers. Any "check then act" over a set of rows rather than one row.
Each of those works perfectly in every test you'll write, because a test runs one transaction at a time.