RFA-328 · Case file with fixtures · Case 300 of 694 · Runtime evidence
An RwLock Writer Panic Poisons the Lock
A panic while an RwLock is held exclusively can interrupt a logical mutation, so later read and write acquisitions report poison. Inspect and validate the guarded value before deliberately clearing that advisory flag.
- Reviewed
- Rust
- Rust 1.98.1, edition 2024
- Targets
- targets with std::sync::RwLock
- Profiles
- dev, release with panic=unwind, test
Direct answer
What this Rust failure means
- Why it happens
- Panic during exclusive write access can interrupt a logical mutation, so the standard lock records poison for subsequent readers and writers.
- First discriminating check
- Bind the write guard across the panic, recover the error's guard, validate the domain invariant, and clear poison only after repair.
The reader-side RwLock case in this Atlas shows that a panic under a read guard does not poison the lock. The complementary rule is just as important: a panic while a write guard is alive can poison it.
The failing program changes a protected integer and panics before the write guard drops normally. A later read() returns PoisonError.
Poison follows interrupted exclusive access
The RwLock poisoning documentation states that the lock may become poisoned only when a panic occurs while it is locked exclusively in write mode. Safe code can mutate the protected T through that guard, so unwind may have interrupted a multi-step invariant.
The lock is released during unwinding. Poison does not leave it permanently locked. Instead, later acquisition returns a result that asks the caller to make a recovery decision.
This distinction prevents a silent return to business as usual after partial mutation.
The protected data is still there
RwLock::read returns a LockResult. Its error contains the guard that would otherwise have been returned.
The repaired program uses PoisonError::into_inner to inspect the value. It sees the write performed before panic. This is why poison cannot automatically roll back state: the integer changed, and a more complex structure might contain only half of an intended transition.
Recovering the guard is access, not proof of validity.
Repair the invariant before clearing poison
After validation, the fixture calls clear_poison. A later read then succeeds normally.
In production, the order matters:
- acquire the poisoned guard;
- inspect or rebuild the protected state;
- make the domain invariant true;
- clear poison;
- allow ordinary operations to resume.
Clearing first removes the warning before repair is complete. Blindly calling unwrap_or_else(PoisonError::into_inner) everywhere has the same problem: it converts an important signal into routine access.
Poison is advisory rather than transactional
An RwLock does not know the semantic invariant of T. A writer may panic before changing anything, after completing all changes, or halfway through them. Poison records the guard and panic relationship, not the exact health of data.
External effects are even less visible. A writer may have sent a message or updated a file before panic. Repairing only the in-memory value may not restore system consistency.
For important transitions I prefer designs that construct a new valid state separately and swap it under a short write lock. This reduces the amount of fallible work performed while exclusive access is held.
Reader and writer failures belong in paired tests
Rust's asymmetry is deliberate. A read guard ordinarily cannot mutate T, so reader panic alone does not indicate interrupted exclusive mutation. A write guard can.
I keep paired tests because “panic under RwLock poisons” is too imprecise. The mode and lifetime of the guard decide whether the event is recorded.
Interior mutability can make the application story more complex: a reader may cause effects through atomics or inner locks, but the outer RwLock still applies its documented rule. Those inner invariants need their own signals.
Guard lifetime must be visible in the reproduction
A temporary guard may drop at the end of a statement before a later panic. Then the panic did not occur while the lock was held, and poison is not expected.
The fixture binds the write guard, mutates through it, and panics in the same scope. The evidence therefore tests the guard relationship rather than merely placing a panic somewhere near a lock call.
For asynchronous code I also remember that std::sync::RwLock guards and blocking acquisition have constraints different from async-aware lock types. I do not carry this poisoning rule onto every third-party lock without reading its contract.
Panic strategy and process boundaries matter
This recovery path requires unwinding. With panic=abort, the whole process ends and no later reader observes in-memory poison. Recovery then comes from restart, persistence, or an upstream supervisor.
Even under unwind, a panic crossing an FFI or task boundary may follow different rules. I test the real deployment profile and keep unsafe code from treating poison as a memory-safety guarantee.
My recovery checklist
I ask which fields could have changed, what invariant joins them, whether side effects escaped, whether a known-good snapshot exists, and whether requests can wait during repair. If validation cannot prove safety, I fail the component rather than clear the flag optimistically.
The core principle is that a lock protects access, while poison reports one kind of interrupted responsibility. An RwLock writer panic makes later callers stop and inspect. The data remains recoverable, but deciding it is trustworthy belongs to the application that understands the invariant.