đź”§ Thread-Safe Per Operation Isn’t Thread-Safe Per Sequence
Switching from `Dictionary` to `ConcurrentDictionary` correctly makes each individual read or write safe from corruption – but a common pattern of “check if a key exists, then decide what to do” is TWO separate operations, and another thread can run its own read and write in between them, producing a race condition that looks identical to the exact bug the switch to `ConcurrentDictionary` was meant to fix.
🔎 The Problem
private readonly ConcurrentDictionary<string, int> _counters = new();
public void Increment(string key)
{
if (_counters.TryGetValue(key, out int current))
{
_counters[key] = current + 1; // Read, then a SEPARATE write.
}
else
{
_counters[key] = 1;
}
}
// Each individual TryGetValue and indexer write is internally thread-safe -
// but between the TryGetValue and the write, another thread can run its
// OWN complete Increment() call. Two threads incrementing "requests" from
// 5 at the same instant can both read 5, both write 6 - and one increment
// is silently lost, with no exception and no obvious symptom.
âś… Fix: Use the Atomic Compound Operations Built for This
- `_counters.AddOrUpdate(key, 1, (k, oldValue) => oldValue + 1)` performs the entire check-then-update as a single atomic operation – this is exactly the method `ConcurrentDictionary` provides specifically because the naive read-then-write pattern above isn’t safe, even on a concurrent collection.
- For a simple counter specifically, `Interlocked.Increment` against a dedicated field (or a `ConcurrentDictionary<string, StrongBox<int>>` combined with `Interlocked`) can be faster still, since it avoids even the dictionary’s own internal locking for the hot increment path.
- As a general rule when working with any concurrent collection: any operation described as “check, then act based on what I found” needs one of the collection’s dedicated atomic methods (`GetOrAdd`, `AddOrUpdate`, `TryUpdate`) – a series of individually-safe calls is not the same guarantee as one atomic compound call.
⚠️ Why This Passes Testing and Fails at Scale
- Under light concurrent load, the window between the read and the write is narrow enough that two threads rarely land inside it at the same moment – the bug becomes measurable only once real concurrent traffic increases the odds of the race actually being hit.
- The resulting symptom – a counter that’s occasionally a little lower than it should be – looks like a plausible business-logic discrepancy rather than a threading bug, which is exactly what makes this class of race condition so easy to chase down the wrong path.
A concurrent collection protects each individual operation – it was never going to protect a decision you make BETWEEN two of them.
