Azure, Debugging

When Redis times out in an application but Redis is fine: Lessons from a Real-World investigation

This post is a friendly recap of a real Azure Managed Redis investigation

Some incidents teach you more than any documentation page ever could. This one started with a few intermittent Redis timeouts. This grew into a deep investigation that touched Kubernetes CPU limits, Java client internals, cluster topology, and even the limits of what you can simulate on a managed service. One symptom, three very different stories.

The customer stays anonymous, because the lessons matter more than the logo. What’s left is the part you can reuse: what happened, what we learned, and tips and tricks for testing. Along the way, we try to showcase how you can look beyond the error message. Settle in, because this one has a few twists. Spoiler: it was mostly not Redis itself.

The short version

  • The setup: a busy, spiky, customer-facing Java platform on Azure Kubernetes Service (AKS), talking to Azure Managed Redis (AMR) in OSS Cluster mode through Spring Data Redis and the Lettuce client.
  • The symptom: intermittent client-side Redis command timeouts while Redis itself looked healthy.
  • The findings: three different patterns, not one. The clearest one hit freshly started pods, where lazy connection setup collided with CPU throttling.
  • The surprise: a production incident raised a harder question: how do you check that your application copes with normal platform events, like maintenance and failover, when you can’t trigger them on demand?
  • The testing advice: you can’t simulate AMR failover yourself (not as of writing of this post), so validate your client, test cold starts realistically, instrument everything, and learn to read the signals.
Quick who’s who, because the names and acronyms get crowded
  • Azure Managed Redis (AMR) is the cache service.
  • OSS (open source software) cluster policy is one of the three clustering policies AMR offers (OSS, Enterprise and Non-clustered). With OSS, clients follow the standard open source Redis Cluster protocol and connect directly to each shard, so the client library is the one that makes a connection to every shard.
  • Shard is one slice of the cache’s data, served by its own Redis process. AMR can spread your data across several shards.
  • Lettuce is the Java Redis client library that the application uses to talk to the cache.
  • Spring Data Redis is the Spring layer that applications use, with Lettuce as the driver underneath.
  • Netty is the Java networking library that Lettuce relies on. It comes along as a transitive dependency of Lettuce, so it lives in the application, not in the Redis cache. Any Netty settings you see below (event loops, native transport) are client-side settings.
  • Product group means the Azure Managed Redis team at Microsoft together with Redis’s client engineering team, who both jumped in to help.

What happened

1. Redis was calm, the clients were not

Under load, commands timed out on the client side after a few seconds, yet the Redis server reported very low latency. That mismatch sent the engineers running the application down the rabbit hole, and each stop along the way taught something:

  • Connection pooling didn’t help. It just moved the failure from “command timed out” to “timed out waiting for a connection from the pool.” A side tip from the product group: if you do use a pool, set max-idle equal to max-active so connections stay open, because creating a new one is slow and expensive.
  • Too many client resources. Inside the application, each cache connection created its own Lettuce client resources, which include Netty event loops (Netty is the networking library Lettuce uses under the hood), thread pools and timers. Several of these sets ended up talking to the same Redis endpoint. The Redis client team advised collapsing them into one shared set and called it the biggest, lowest-risk win.
  • A homemade “force reconnect” wrapper. It had been added to escape stalled sockets that can hang for 13 minutes or more. The better fix is the TCP_USER_TIMEOUT setting, after which the wrapper can go. If you keep one, audit its trigger: an over-eager threshold can tear down healthy connections and create reconnect churn of its own.
TCP_USER_TIMEOUT is a Linux network setting that limits how long sent data can go unacknowledged before the connection is closed. Without it, a connection that silently stalls can hang for many minutes before the operating system gives up.

One idea worth taping to your monitor: a “command timed out” error is not always a latency problem. Lettuce starts the clock the moment a command is issued, even if it hasn’t been written to the socket yet. If the client is starved for threads or CPU, the timer can expire before Redis ever sees the command.

2. Measuring the right things

The breakthrough came from measuring more. The engineers running the application exposed Lettuce’s built-in first-response and completion timings and added two custom metrics of their own: end-to-end command time and event-loop lag. All of them rose at the same time, but only for pods running on particular worker nodes, not for particular application instances. That pointed away from any single stage of the Redis client and toward the nodes themselves, which were showing high memory use, heavy CPU throttling, or connection-tracking pressure.

To test the idea, the engineers running the application ran a load test with extra node capacity provisioned ahead of time, and the service they were analyzing came back clean. Steady-state load was clean too once the pods had settled in. The trouble showed up while pods were starting and the nodes were under strain. Pre-scaling the nodes gave the best results in later tests, but it never made the timeouts disappear completely, so the next step was sorting the remaining ones into distinct patterns.

3. Three patterns, not one

With the nodes pre-scaled, the timeouts that remained fell into three distinct patterns.

Pattern A: startup timeouts. Newly created pods timed out during their first CPU spike. That’s the known CPU-throttling mechanism, which is how Linux enforces Kubernetes CPU limits in short 100 ms windows. To understand it better, the Redis team built a minimal lab: a 6-shard TLS OSS cluster, Spring Data Redis with Lettuce 7.x, and a burst of concurrent requests on a freshly started JVM. Here’s what they found in that lab:

  • Spring Data Redis’s eagerInitialization option sets up the cluster topology and one connection at startup, but the remaining per-shard connections are still opened on first use.
  • That’s intentional driver behavior, not a bug, so there is no fix tied to a specific release.
  • On a CPU-throttled pod, the TCP and TLS handshakes to every shard land on the request path and can push commands past their timeout.
  • In that lab test, pre-warming all primary connections cut first-burst p99 latency (the 99th percentile) by roughly 12x and overall burst completion time by about 2.7x compared with a cold start (medians over 30 test runs). Read these as results under those test conditions, a small local reproduction without CPU throttling, and not as a guaranteed production outcome. Your own results will depend on things like your shard count, TLS setup, CPU limits and traffic.
  • The published warm-up recipe works on current Lettuce 7.x, so no upgrade is needed. The warm-up lives in application code.
p99 (99th percentile) means 99 percent of requests were faster than this number, so it shows how slow the slowest 1 percent get.

Pattern B: synchronized bursts. This one looked different. Many pods on different nodes timed out within the same minute, while CPU and memory stayed calm and Redis stayed lightly loaded.

When lots of unrelated pods stumble at the same moment, the cause is usually something they all share, not a problem inside each pod. Two candidates were on the list to rule out: something shared on the Redis side, like a proxy or shard event, or a change to the cluster topology (the map of which shard holds which data). The suggestion was to line the timeouts up with Lettuce’s topology refresh entries, the moments it re-checked that map.

A follow-up test with topology debug logging turned on showed timeouts at just two moments:

  • One lined up with a scale-down, when the cluster removes pods and they shut down.
  • The other hit pods that were up and running, with no throttling, no scaling and nothing unusual in the Lettuce logs.

The product group also checked for Azure maintenance on the client nodes around those times and found none, so that explanation was ruled out. Rule of thumb: when many pods fail together, look for what they share.

Pattern C: memory at 100%. Nearly all the timeouts in one test landed in the final minutes, exactly when Redis memory hit 100%. This was the first genuinely server-side signal of the whole investigation. At full memory, Redis starts evicting keys (or blocking writes, depending on the policy), and the resulting latency is easy to miss if you only watch CPU. The suggestion was to correlate timeouts with the evicted_keys metric near the memory limit.

What to take from the three patterns:

  • Start with Pattern A. It has a clear cause and practical fixes: warm up your connections, hold traffic back until they’re ready, and ease CPU limits during startup.
  • Look for Patterns B and C in your own environment. For B, ask whether many pods on different nodes failed in the same minute, then look at scale-downs, topology refreshes and maintenance around that time. For C, watch memory and the evicted keys metric, and set an alert so you hear about it before the cache reaches 100 percent.
  • Keep looking past the first fix. One error message can hide several causes, so fixing the first one you find doesn’t mean you’re done. Measure again after each change.

4. A production surprise and the bigger lesson

At one point, a production incident came up, and Azure support was notified. The incident itself matters less than what it exposed: a testing gap.

Platform events such as OS updates, failovers, freeze events, live migration and Redis maintenance can’t be triggered on demand in a pre-production environment today. That meant the engineers running the application couldn’t reproduce them, and so couldn’t check how their client would cope. You can’t validate what you can’t reproduce, and that question shapes the testing recommendations later in this post.

5. Two small upgrade surprises

  • What you think you run isn’t always what you run. Netty isn’t part of the Redis cache. It arrives in your application as a transitive dependency of Lettuce, so your build decides which version you get. In this case, an engineer found that a newer Netty and Lettuce pairing had been improperly forced into a service build. When the Redis team asked whether Lettuce 7.6 was actually in use when the issue occurred, the answer was no: the service was running an older Lettuce 6.x. Verify what’s really on your runtime classpath, including the Netty version Lettuce pulls in. A dependency report from your build tool (for example mvn dependency:tree or gradle dependencies) shows exactly what resolved.
  • Native transport matters. Netty’s native transport (epoll on Linux) is a client-side choice, and it’s what makes TCP_USER_TIMEOUT work. After enforcing it, the engineers running the application learned during testing that Lettuce 7.6.0 refuses to start their service without the right transport. That’s a friendly fail-fast. A Lettuce maintainer at Redis called removing the custom reconnect logic and adding TCP_USER_TIMEOUT a sound direction, with one caveat: depending on your operating system and Java version, the default NIO transport won’t give you TCP_USER_TIMEOUT, so check your logs for messages about which transport was picked. They also noted that recent Lettuce versions work best with Netty’s io_uring transport, which has graduated from incubation and shows considerable latency improvements.
  • Native transport is Netty’s optional, Linux-specific network code. It talks to the operating system more directly than Java’s standard library does, and it is what unlocks settings such as TCP_USER_TIMEOUT.
  • epoll is the Linux feature that Netty’s native transport uses to watch many network connections at once.
  • io_uring is a newer Linux interface for fast, fully asynchronous input and output. Netty offers a transport built on it.
  • NIO (New I/O) is Java’s built-in library for non-blocking network and file input and output. It is Netty’s default transport, and depending on your operating system and Java version it may not support Linux-only settings such as TCP_USER_TIMEOUT.

Testing recommendations

1. Know what you can (and can’t) simulate today

At the time of writing, this investigation did not identify a supported customer-initiated way to trigger Azure Managed Redis failover or maintenance events in a test environment. Check the latest product documentation before designing a test plan around this limitation.

2. Validate your client

Focus on the following behaviors. Redis’s Lettuce production usage guide covers most of them:

  • Cluster topology refresh (a failover should trigger a topology update on the Lettuce side)
  • Reconnect behavior and connection recreation
  • TCP timeout handling (TCP_USER_TIMEOUT with the right native transport)
  • Connection warm-up
  • Kubernetes readiness gating

Then build in the standard resilient patterns: retry with backoff, circuit breakers, graceful degradation, and connection recreation when needed. A small failure rate from transient disruptions is expected, which is exactly why retries are recommended.

Set your expectations from the platform side too. Maintenance events are routine and often come with client connection resets. Your client library is responsible for recreating those connections automatically, and Redis endpoints are typically available again within a few seconds.

If Lettuce ever surprises you in a way the documentation doesn’t explain, the Lettuce issue tracker is the most direct way to reach its maintainers.

3. Test cold starts and scale-outs like production does

  • Warm up all shards, for every cache, in every service. Make the warm-up complete and readiness-gated, so Kubernetes only routes traffic once it finishes. A partial or ungated warm-up leaves the gap open.
  • Use the documented warm-up approach instead of custom code. It’s described in Warm up cluster connections (Redis docs) and Warming up connections (Lettuce docs). In short:
  • Burst-test a fresh JVM across all shards, with and without warm-up, and compare first-burst p99. That’s how the Redis team measured the difference in their lab, so measure your own setup rather than expecting the same numbers. Autoscaler scale-ups should behave much like startups.
  • Use production-like CPU limits. The Redis lab test ran without throttling, so the connection cost showed up as extra latency rather than timeouts. Their working assumption is that a CPU-throttled starting pod makes those handshakes much slower, which is where timeouts come from. My takeaway: give your test pods the same CPU limits as production so you can see it for yourself.
  • Ease CPU limits during startup, by raising or removing them for that window.
  • Check service mesh startup ordering. If you use Istio, make sure the sidecar is ready before your app opens Redis connections, otherwise your warm-up connections race the mesh.
PING is a tiny Redis command that simply asks the server whether it is there. In the warm-up recipe, its only job is to make the client open a connection to each shard up front.

4. Instrument your tests

You can’t fix what you can’t see, so switch these on before your next test run. Each item says what to watch and how to enable it. The logging switches and one code snippet follow the list.

Quick note on terms used below
  • CFS (Completely Fair Scheduler) is the Linux process scheduler introduced in kernel 2.6.23, designed to share CPU time fairly among tasks. Kubernetes CPU limits are enforced through its “bandwidth control” feature, which gives each container a fixed slice of CPU time every 100 milliseconds and pauses (throttles) it once the slice is used up. Newer kernels have moved to a successor scheduler, but this feature keeps the CFS name.
  • cgroup (control group) is the Linux feature Kubernetes uses to cap and measure the CPU and memory a container gets. Version 2 (v2) is the newer design intended to replace version 1 (v1).
  • QoS (Quality of Service) class is a label Kubernetes gives every pod (Guaranteed, Burstable or BestEffort) based on its CPU and memory settings. It decides which pods are evicted, meaning removed, first when a node runs short of resources.
  • Check for CFS throttling first. Throttling happens in 100 ms windows, so a comfortable 50 to 60 percent CPU average can hide it. Also consider moving toward requests equal to limits (Guaranteed QoS), or at least raising requests.
  • Turn on ConnectionWatchdog debug logging and count reconnects and “reset by peer” errors around the timestamps of your timeouts. Set the io.lettuce.core.protocol.ConnectionWatchdog logger to DEBUG (see the logging block below). For a cleaner count, Lettuce also publishes connection events (connected, disconnected, reconnect failed) that you can subscribe to. See Lettuce events.
  • Compare Lettuce’s first-response latency with server-side latency. If first-response time is far above the server’s roughly 100 microseconds, you’re looking at client-side queueing.
  • Enable topology refresh debug logging when you’re chasing synchronized bursts. Set the io.lettuce.core.cluster.topology and io.lettuce.core.cluster.RedisClusterClient loggers to DEBUG (logging block below), then line up refresh events with your timeout timestamps. Lettuce also publishes cluster topology events on the same event bus. For the client settings that control refreshes, see the Cluster topology refresh section of Redis’s Lettuce production usage guide.
  • Watch memory and evicted_keys whenever Redis gets close to 100 percent. In the Azure portal, open your cache, go to Metrics, and chart Used Memory Percentage and Evicted Keys next to your timeout count. Then add an alert on Used Memory Percentage so you hear about it before the cache fills up. Microsoft’s memory management guidance recommends alerting on that metric.
  • Log every error and exception from Redis components. Multiple maintenance events routinely touch every Azure Redis resource, so detailed logs let you confirm healthy resilience or diagnose trouble after the fact. Don’t silence the Lettuce and Spring Data Redis loggers (INFO is a good baseline for test runs). Log each exception with its full cause chain, because Spring Data Redis wraps Lettuce errors in its own, which puts the original at the bottom. Add the cache name, command and elapsed time to the message, and count errors with a Micrometer counter tagged by exception type so the same events show up on a chart.
  • Know exactly which changes are live in each test run. It helps you separate client-side gains from infrastructure-side gains. At startup, log the versions that actually resolved (Lettuce, Netty, Spring Data Redis) and the effective client options such as timeouts and thread pool sizes. Keep a small run sheet with CPU requests and limits, node counts and whether nodes were pre-scaled, and tag your metrics with a run ID.
  • Keep a fallback comparison in mind. If Lettuce can’t be tuned to the reliability you need, the product group suggested considering Jedis, which uses its own network stack instead of Netty. A side-by-side run under the same workload is a natural test. Spring Data Redis supports both connectors behind the same APIs, so the swap mostly lives in your connection factory configuration (JedisConnectionFactory instead of LettuceConnectionFactory). See Spring Data Redis drivers.

Logging switches in one place. These are Spring Boot application.properties lines. Use the same logger names in any other logging setup, and see Spring Boot log levels for the details. Debug logging is chatty, so switch it on for test runs rather than leaving it on.

logging.level.io.lettuce.core.protocol.ConnectionWatchdog=DEBUG
logging.level.io.lettuce.core.cluster.topology=DEBUG
logging.level.io.lettuce.core.cluster.RedisClusterClient=DEBUG

Lettuce command latency to Micrometer. Based on the Lettuce observability guide. Pass the resulting client resources to your connection factory, for example through LettuceClientConfiguration.

5. Read the signals in production

Azure records maintenance in your cache’s Activity Log (see Failover and patching). A “healthevent” entry is emitted when maintenance begins, and you can set up an alert on the Activity Log to be notified. Your own metrics and clients give you a second view, because you can also infer that Redis was impacted from patterns like these:

What you seeWhat it usually means
A momentary, significant dip in the connectedclients metric (aggregated by minimum)A Redis node reset all its client connections, which happens during many kinds of platform maintenance
One client, or a few closely related clients (for example pods on a single AKS node), affected, with little change in connectedclientsMost likely a client-side issue
Many clients hit at the same time on the same Redis resourceMost likely maintenance on that resource

Two more pointers. The public documentation covers scheduled maintenance windows (a preview feature) and notes that some activities are excluded from them: host OS updates, Azure networking component updates, critical security patches and major Redis version upgrades. And if you suspect Azure maintenance on your own side, check whether it touched your client nodes, which is exactly what the product group did here.

6. Keep noise from hiding real problems

  • Shut down gracefully. Use a preStop hook and call shutdown() on your Lettuce client resources. It can eliminate noisy shutdown errors and gives you a clearer view of the real ones.
  • Check your dashboards after upgrading Lettuce. Lettuce 7.6.0 sends a CLIENT MAINT_NOTIFICATIONS command on new connections to subscribe to maintenance events. As far as the Redis team knows, AMR doesn’t support that command, so it is rejected. There’s no functional impact on your traffic, but it can look scary in APM tools. Redis suggests opting out with .maintNotificationsConfig(MaintNotificationsConfig.disabled()) rather than downgrading the protocol to RESP2. The Microsoft product team’s view: you can leave it alone or filter that harmless error out of your dashboards.
RESP (Redis Serialization Protocol) is the language Redis clients and servers use to talk to each other. RESP3 is the newer version and RESP2 is the older one.

Final thoughts

Remember the spoiler at the top? Redis was mostly innocent. The bigger lesson is that a timeout is a symptom, not a diagnosis.

If you only remember three things, make it these:

  1. Measure before you blame. The client, the cluster and the server can all produce the same error, so check each one before pointing fingers.
  2. Test the awkward moments. Startup, scale-out, shutdown and maintenance are where systems get shy. Steady-state load is the easy exam.
  3. Aim for a great houseguest. A good client arrives prepared (warm-up), recovers gracefully when something goes wrong (reconnects), tells you what it’s up to (logging), and leaves quietly (graceful shutdown).

Managed services keep improving, so keep the official documentation bookmarked. Now go put your own client through its paces, and may your timeouts be few and your logs be clear.

This post reflects the guidance available at the time of writing. Features and guidance can change, so double-check the official documentation before you plan around them.