A product page usually reads from Redis. The database sees only occasional cache fills, so its connection pool looks comfortably sized. Then one popular entry expires during a traffic burst. Fifty requests observe the same miss and independently run the same database query.
Redis has not become slow. The cache has temporarily stopped suppressing duplicate work. If the origin slows under the extra load, the miss window grows, admitting still more duplicate requests.
This is a cache stampede. Fixing it requires three separate decisions: how many callers may refresh one key, whether expirations across different keys should align, and what other callers receive while a refresh runs. A lock, a random TTL, and a stale response address different parts of that problem.
We will compare those mechanisms, then run a Python implementation against a real local Redis server. The recorded versions are Python 3.13.4, redis-py 8.1.0, and Redis 8.2.9. The origin is an explicitly simulated asynchronous read. The experiment counts duplicate origin calls; it is not a database throughput benchmark.
Start with the miss window, not the average hit rate
Consider a hypothetical key receiving 400 requests per second whose refresh takes 250 milliseconds. Roughly 100 requests can arrive during that refresh interval. If each independently loads the missing value, a single expired entry can create a burst much larger than the usual origin traffic suggests.
That calculation assumes approximately steady arrivals over the interval. It is a way to reason about amplification, not a measured production result. Real traffic bursts, connection limits, queueing, and origin latency distributions change the outcome.
An excellent daily hit ratio can hide this failure because misses are concentrated in time and around a few keys. Measure refresh duration, simultaneous refreshes per key, waiters, origin calls, and refresh failures. Those describe the load the database must absorb when the cache stops helping.
Three mechanisms with different jobs
| Mechanism | What it changes | What remains |
|---|---|---|
| Single flight / request coalescing | Concurrent requests for one key share a refresh. | Waiters still need deadlines; coordination has a scope. |
| TTL jitter | Different keys tend to expire at different times. | One hot key can still receive many simultaneous misses. |
| Stale-while-revalidate | Callers receive an eligible older value while refresh runs. | Refreshes still need coordination and a maximum stale age. |
A process-local single-flight map is often the cheapest first layer. Go’s singleflight package is a concrete example of suppressing duplicate calls by key. A local map cannot coordinate independent application processes, however. Twenty replicas can still produce twenty refreshes.
The implementation below uses a short Redis lease to coordinate clients sharing the same Redis instance. While an owner remains within its lease, other callers wait for its cache fill. If the lease expires, another caller may refresh. This is a cache-efficiency mechanism with explicit failure limits, not a guarantee that the origin is called exactly once.
A refresh lease must protect publication too
A lease is acquired with SET lease-key token NX PX duration. NX requires the key to be absent, and PX gives it a millisecond expiry. Acquisition and expiry are configured by one command; see the Redis SET reference.
The token identifies this particular acquisition. A previous owner must not delete a replacement owner’s lease. Redis’s locking guidance describes checking the token before release. The same ownership question matters when publishing the loaded value.
Suppose owner A pauses long enough for its lease to expire. Owner B obtains a new lease and publishes a newer result. When A resumes, an unconditional cache write would replace B’s value. Comparing the token only during cleanup is too late.
Our publication script checks ownership, stores the value with its TTL, and removes the lease without allowing another client’s command to interleave. Redis documents this execution property in its Lua scripting introduction. The script does only bounded key operations; the origin read stays outside Redis.
The complete cache-fill implementation
Save this as cache.py. The loader must be an asynchronous, repeatable read that returns a string. Use a Redis client configured with decode_responses=True.
import asyncio
import hashlib
import random
import secrets
from contextlib import suppress
from redis.exceptions import RedisError
PUBLISH = """
if redis.call('GET', KEYS[1]) ~= ARGV[1] then return 0 end
redis.call('SET', KEYS[2], ARGV[2], 'PX', ARGV[3])
redis.call('DEL', KEYS[1])
return 1
"""
RELEASE = """
if redis.call('GET', KEYS[1]) == ARGV[1] then
return redis.call('DEL', KEYS[1])
end
return 0
"""
def keys_for(identity: str) -> tuple[str, str]:
digest = hashlib.sha256(identity.encode()).hexdigest()
return f"stampede:{{{digest}}}:data", f"stampede:{{{digest}}}:lease"
async def get_or_load(redis, identity, loader, *, ttl_ms=10_000,
jitter_ms=1_000, lease_ms=1_500, wait_s=2.0):
"""Cache a string; loader must be an async, repeatable read operation."""
for value in (ttl_ms, jitter_ms, lease_ms):
if type(value) is not int:
raise ValueError("millisecond settings must be integers")
if ttl_ms <= 0 or lease_ms <= 0 or not 0 <= jitter_ms < ttl_ms:
raise ValueError("invalid TTL, jitter, or lease duration")
if wait_s <= 0:
raise ValueError("wait_s must be positive")
data_key, lease_key = keys_for(identity)
async with asyncio.timeout(wait_s):
while True:
cached = await redis.get(data_key)
if cached is not None:
return cached
token = secrets.token_hex(16)
acquired = await redis.set(lease_key, token, nx=True, px=lease_ms)
if acquired:
try:
# Another owner may have filled between GET and SET NX.
cached = await redis.get(data_key)
if cached is not None:
return cached
value = await loader()
if not isinstance(value, str):
raise TypeError("loader must return a string")
ttl = ttl_ms - secrets.randbelow(jitter_ms + 1)
accepted = await redis.eval(
PUBLISH, 2, lease_key, data_key, token, value, ttl)
if accepted:
return value
# The lease expired: discard this result and read again.
finally:
# Best effort only; expiry remains the crash-recovery path.
with suppress(RedisError, TimeoutError):
async with asyncio.timeout(0.2):
await redis.eval(RELEASE, 1, lease_key, token)
await asyncio.sleep(random.uniform(0.010, 0.025))
The second cache read is intentional. A previous owner may finish after the initial miss but before this caller acquires the lease. Rechecking prevents a redundant origin call in that interval.
Empty strings are valid hits, so the condition is cached is not None. A database “not found” result needs its own explicit representation if you want negative caching; it should not accidentally share a representation with a cache miss or infrastructure failure.
Both keys include the same hash tag. Redis Cluster uses matching hash tags to place keys in the same slot, which matters for a multi-key script. See the cluster specification. This key layout alone does not make the example cluster-tested: the experiment uses a standalone server and a standalone client.
Include tenant, authorization scope, locale, query parameters, and representation version in the cache identity when they affect the result. Hashing an incomplete identity does not fix it. Every writer for these keys must follow the same ownership and data-format contract.
Reproduce a synchronized cold-cache burst
Download the complete laboratory: source code, pinned dependency, runner, and all 15 tests (ZIP). The runner starts its own disposable Redis process on a dynamically selected loopback port and stops it after verification. The inline example below can also be run against a separate laboratory instance.
Install the tested client version in an isolated environment:
python3 -m venv .venv
source .venv/bin/activate
python -m pip install redis==8.1.0
With a Redis 8.2.9 binary available, start a disposable loopback instance in a separate terminal. Persistence is disabled for this laboratory:
redis-server --bind 127.0.0.1 --port 6387 --save "" --appendonly no
Save the following as demo.py beside cache.py, then run python demo.py from the activated environment. The default URL is redis://127.0.0.1:6387/0; REDIS_URL can select a different laboratory instance.
import asyncio
import os
import uuid
from redis.asyncio import Redis
from cache import get_or_load, keys_for
async def main():
url = os.environ.get("REDIS_URL", "redis://127.0.0.1:6387/0")
clients = [Redis.from_url(url, decode_responses=True,
socket_connect_timeout=0.5, socket_timeout=0.5)
for _ in range(2)]
identity = "demo:" + uuid.uuid4().hex
calls = 0
barrier = asyncio.Barrier(50)
async def origin():
nonlocal calls
calls += 1
await asyncio.sleep(0.05) # Simulated origin, not a database benchmark.
return '{"name":"example-product"}'
async def naive(i):
client = clients[i % 2]
data_key, _ = keys_for(identity)
value = await client.get(data_key)
await barrier.wait() # Force all 50 requests to observe the cold cache.
if value is None:
value = await origin()
await client.set(data_key, value, px=10_000)
return value
try:
await asyncio.gather(*(naive(i) for i in range(50)))
assert calls == 50
print(f"naive cold burst: origin_calls={calls}")
await clients[0].delete(*keys_for(identity))
calls = 0
values = await asyncio.gather(*(
get_or_load(clients[i % 2], identity, origin) for i in range(50)))
assert len(set(values)) == 1
assert calls == 1
print(f"coordinated cold burst: origin_calls={calls}, responses={len(values)}")
finally:
await clients[0].delete(*keys_for(identity))
for client in clients:
await client.aclose()
if __name__ == "__main__":
asyncio.run(main())
The actual output was:
naive cold burst: origin_calls=50
coordinated cold burst: origin_calls=1, responses=50
The barrier forces every request in the naive variant to observe the same cold cache before any origin result is available. The coordinated variant sends the same number of requests through two independent Redis client pools. The simulated origin finishes comfortably inside the lease.
The result demonstrates duplicate suppression under those conditions. It does not establish production latency, maximum throughput, or behavior through Redis failover. Adding a random TTL alone would not change that first synchronized miss: there is no cached value whose expiry can be spread yet.
The clients close their pools explicitly. That lifecycle is documented in the redis-py asyncio examples. A long-running service should normally reuse its clients rather than construct a new pool for every request.
Separate the caller deadline from the lease duration
The request budget limits how long this caller waits. The lease duration limits how long other clients treat its refresh claim as current. The cache TTL controls how long a published entry remains available. They are three different clocks.
A waiting caller that times out must not immediately bypass coordination and hit the origin. That would recreate the stampede precisely when the refresh is slow. This implementation propagates the timeout. Map that outcome to an explicit application policy: a bounded error, an eligible stale response, or a separately capacity-limited fallback.
The owner performs its load inside the request task. Cancelling that owner can cancel the refresh; cancelling a different waiter does not. Cleanup checks the token and waits at most an additional 200 milliseconds under normal cooperative scheduling. Expiry remains the recovery path if cleanup cannot reach Redis.
Python’s timeout mechanism uses cancellation. A loader that blocks the event loop or suppresses cancellation can exceed the intended budget. The example does not turn an arbitrary dependency into a forcibly interruptible operation.
If refreshes should survive client disconnects, give them a service-owned task manager with bounded admission, observed errors, and an orderly shutdown policy. Do not turn every miss into an untracked background task. Likewise, refreshing a lease requires a separate ownership-checked renewal design; the sample intentionally does not renew leases.
Use jitter to spread keys, not to elect an owner
The implementation chooses a TTL between ttl_ms - jitter_ms and ttl_ms, inclusive. With the defaults, a successful fill lives for between nine and ten seconds. Subtracting jitter keeps the configured storage TTL as an upper bound.
This helps when many different entries would otherwise be written together and expire together. The draw happens once per successful fill, not independently for each reader. All readers of one hot entry still share its expiry, so that key still needs a refresh policy.
A wider jitter window spreads expirations further but shortens average residence time and can increase steady-state refresh traffic. Choose the range against origin capacity and acceptable cache age. It cannot solve a full cache flush, a newly popular uncached key, or an origin outage.
Storage TTL is also not a proof of source freshness. A slow loader may return a snapshot that was already old before publication. If freshness is defined against source observation time or a database version, preserve that information and account for load duration, replication lag, and the chosen time basis.
Stale data needs two boundaries and one refresh policy
Serving an older value can keep refresh latency away from users, but “stale is allowed” is incomplete. Define a fresh interval and a final serving deadline. For example, a value might be fresh for five minutes and serveable for another thirty seconds while a refresh is attempted.
| Age in this example | Response policy | Refresh policy |
|---|---|---|
| Less than 300 seconds | Serve as fresh. | No expiry-driven refresh required. |
| 300 to less than 330 seconds | Serve the eligible stale value. | Admit a coordinated refresh. |
| 330 seconds or more, or no value | Do not serve under this stale window. | Wait within a deadline or apply an explicit failure policy. |
The HTTP concepts are described by RFC 5861: stale-while-revalidate permits serving stale content during refresh, while stale-if-error permits it under specified error conditions. An application cache must implement its own corresponding policy; storing an HTTP-like label in Redis does not activate that behavior.
Keep the payload physically available through its allowed stale window. A Redis key that disappears at the fresh boundary leaves nothing to serve stale. Store freshness metadata with the value and set physical expiration to cover the final allowed window. Use a consistent time basis across application instances and budget for clock skew.
The code above implements cold-miss waiting and TTL jitter. A stale-serving extension must inspect the envelope, return the old value when eligible, and schedule a refresh through a bounded owner. Its refresh path must bypass the ordinary cache-hit return; wrapping get_or_load around a still-present stale value would simply return that value without refreshing it.
A failed refresh must not silently move the old value’s final deadline forward. Otherwise a nominal thirty-second stale allowance becomes indefinite. Choose this policy per data class: an old product description and an authorization decision do not have the same correctness requirements.
Know what the lease cannot protect
The token check rejects publication by an expired owner, but it cannot stop that owner’s origin request from continuing. Lease expiry can therefore allow overlapping loads. Keep the loader repeatable and avoid using this cache mechanism to coordinate irreversible writes.
It also does not solve an update-versus-fill race. A database update and cache invalidation can occur while a valid owner is loading an earlier snapshot. Preventing that old snapshot from being republished requires coordination with invalidation or source-version checks, not merely the refresh lease.
Redis failure changes the coordination assumptions. A lost lease during failover or eviction can permit another owner. Redis’s locking guidance explains the replication-related limitations. If occasional duplicate reads are acceptable, handle them within an origin capacity budget; if correctness depends on exclusivity, this example is insufficient.
Redis eviction policies can remove keys under memory pressure before their TTL expires. Monitor that separately from ordinary expiry. The sample returns Redis read, acquisition, and publication errors rather than opening an unrestricted path to the origin.
Per-key coordination is not a global concurrency limit. Ten thousand distinct cold keys can still produce ten thousand owners. Bound total origin work, queued refreshes, request admission, and connection usage. Waiting callers also generate Redis reads and lease attempts; local coalescing can reduce that pressure before requests reach the shared lease.
Test failures as carefully as the successful fill
The integration suite ran against a disposable Redis 8.2.9 process and passed 15 tests. Besides the 50-request burst, it checks cache hits containing empty strings, independent keys, a fill between the first read and lease acquisition, origin failure, invalid loader output, caller cancellation, waiter timeout, lease expiry, and both jitter endpoints.
The most important race test holds an old loader until its real Redis lease expires, lets a replacement owner publish, and then releases the old loader. The old caller returns the replacement value instead of overwriting the cache. A separate wrong-token script test verifies that an old token cannot publish or delete a replacement owner’s still-active lease.
A connection-failure test also verifies that loss of Redis connectivity does not call the origin through an implicit fallback. The laboratory does not test Redis Cluster, replica promotion, network partitions, sustained load, or a real database. Those belong in deployment-specific validation.
A useful production report combines refresh amplification, waiter duration, rejected late publications, origin errors, and the age of stale responses. Hit rate can remain a headline metric, but these measurements show whether the cache is protecting the data service when requests arrive together.
What do you think?