Azure, C#, Debugging

When Redis times out in an application but Redis is fine: The .NET Edition

Did you read When Redis times out in an application but Redis is fine: Lessons from a Real-World investigation and think, “Great story, but we’re a .NET shop”? Fair enough! This is the companion piece. The investigation happened on the Java side, with the Lettuce client, but the lessons travel well, and most of them map neatly to StackExchange.Redis, the main Redis client library for .NET.

A quick recap, in case you haven’t read it. An application running on Kubernetes saw intermittent client-side Redis timeouts while Redis itself looked healthy. The clearest cause was freshly started pods, where connections to each shard were opened on first use just as CPU throttling slowed everything down. Two other patterns showed up too, and checking how a client copes with maintenance and failover turned out to be hard to do in a test environment.

One honest note before we dive in: the investigation itself was Java. Everything below draws on the public documentation for StackExchange.Redis, .NET and Azure Managed Redis, so treat it as guidance to check against your own setup, not as a record of what happened.

The lessons at a glance

Here is how each part of the Java post maps to .NET.

In the Java postFor .NET
LettuceStackExchange.Redis
Spring Data RedisRoughly, Microsoft.Extensions.Caching.StackExchangeRedis (ASP.NET Core’s distributed cache)
Netty and its event loopsNot used. Watch the .NET thread pool instead
One shared set of client resourcesOne shared ConnectionMultiplexer
Warm-up, then readiness gatingConnect at startup, then a readiness health check
TCP_USER_TIMEOUTThe Linux tcp_retries2 setting
Custom force-reconnect wrapperThe ForceReconnect pattern, with generous thresholds
Lettuce metrics and eventsTimeout message details, connection events, OpenTelemetry traces
Checking versions on the classpathdotnet list package –include-transitive

1. Share one connection

StackExchange.Redis shares a small number of connections between all your requests through one ConnectionMultiplexer. Its own docs are direct about it: a multiplexer is designed to be shared and reused, and you should not create one per operation. Create it once at startup and reuse it, as Microsoft’s ASP.NET Core sample for Azure Managed Redis does. It’s the .NET version of the Java post’s lesson about sharing client resources. Instead of a connection pool, the multiplexer pipelines requests from many callers over a shared connection.

With the OSS cluster policy you connect the same way. The multiplexer discovers the topology and routes commands to the right shard, so there’s no separate cluster client like Lettuce’s RedisClusterClient (details).

Microsoft’s connection resilience guidance recommends both settings. abortConnect=false lets the multiplexer handle reconnection, and a five-second connect timeout keeps a slow start from turning into a connect, fail, retry loop. For authentication, see the .NET quickstart.

2. Read the timeout message before blaming Redis

Quick note on terms
  • Thread pool is the set of worker threads .NET keeps ready to run your code. When they are all busy, new work waits in line.
  • IOCP (I/O completion port) threads are the ones that handle finished network reads and writes.
  • Sync over async means blocking a thread while you wait for an async call to finish, for example with .Result or .Wait().

A RedisTimeoutException is the .NET cousin of Lettuce’s “command timed out”, and it doesn’t mean Redis was slow. Log the whole message and read these fields first (see the timeout docs and the sync over async page):

  • qs: commands sent and still waiting for a reply.
  • in: bytes that have arrived but haven’t been read. If it’s large, Redis answered and nothing was free to pick up the reply.
  • IOCP, WORKER and POOL: the thread pool’s statistics. When Busy is at or above Min, the pool adds threads at only about one or two per second.

The Java post’s startup problem has a .NET twist. The thread pool’s default minimum equals the processor count, and in a container with a CPU limit, .NET rounds that limit up to a whole number of processors (see ThreadPool.SetMinThreads and Environment.ProcessorCount). A pod limited to 500m of CPU therefore starts with a minimum of one thread, and a startup burst can queue behind the slow trickle of new ones. So when you test cold starts, give your test pods the same CPU limit as production, as the Java post recommends.

What to do, in order:

  1. Stop blocking on async calls. Blocking with .Result, .Wait() or .GetAwaiter().GetResult() is, per the StackExchange.Redis docs, the most common cause by far of a saturated thread pool, and the synchronous API isn’t an escape hatch. Use await all the way up (ASP.NET Core best practices). Current versions of the package ship a build-time analyzer (rule SER307) that flags these calls.
  2. Treat a higher thread pool minimum as a stopgap. It helps a genuine burst but not an app that keeps blocking, and Microsoft warns it can hurt performance elsewhere. Raise it in small steps and measure, in runtime configuration or with ThreadPool.SetMinThreads.
  3. Ease the CPU limit during startup, as in the Kubernetes advice in the Java post.

To watch the pool live, the dotnet-counters tool shows threadpool-thread-count and threadpool-queue-length (see well-known EventCounters).

3. Connect at startup, prove it, then take traffic

The Java lesson was to warm up connections and hold traffic back until they’re ready. Whether a client opens its per-shard connections up front or on first use is an implementation detail, so rather than assume, make the readiness check prove it. In .NET, connect when the app starts instead of on the first request, then hold traffic back until Redis answers, using a health check that your readiness probe calls (health checks in ASP.NET Core). This one pings every primary node once and remembers the result:

It only has to succeed once, so a brief Redis blip later won’t pull every pod out of rotation at the same time. The advice in the Java post about CPU limits applies unchanged.

4. Reconnects and retries

  • Mind the Linux TCP settings. Microsoft’s guidance says the default TCP settings in some Linux versions can leave a broken connection undetected for 13 minutes or more, and recommends setting net.ipv4.tcp_retries2 to 5 for Linux-hosted clients. It’s the same stalled-socket problem that TCP_USER_TIMEOUT solves in the Java post. On Kubernetes, check with your platform team where this setting can be applied.
  • Be careful with ForceReconnect. Microsoft’s connection resilience guidance describes a pattern that rebuilds the multiplexer in the rare case it doesn’t recover. Use generous ReconnectMinInterval and ReconnectErrorThreshold values, especially if you trigger it on timeouts, because eager reconnects can cascade into an already overloaded server. Same caution as the homemade reconnect wrapper in the Java post: audit the trigger.
  • Stagger reconnects. When many pods reconnect at the same moment after a failover, new connections are rate limited and recovery takes longer. A little jitter in your retry policy helps.
  • Retry with backoff and add a circuit breaker. Polly and Microsoft.Extensions.Resilience give you both.

5. Watch and verify

  • Log the multiplexer’s events. ConnectionFailed, ConnectionRestored, ErrorMessage, HashSlotMoved (a cluster moved a hash slot between nodes) and ConfigurationChanged are listed in the Events docs. Line them up against your timeout timestamps, the job Lettuce’s connection and topology events do. Setting LoggerFactory on ConfigurationOptions sends connection logs to your normal logging.
  • Add traces. The OpenTelemetry instrumentation for StackExchange.Redis records Redis calls as traces. It’s still in beta, so expect changes, and it only covers the multiplexer you register with it.
  • Check what you actually run. StackExchange.Redis often arrives indirectly: the Microsoft.Azure.StackExchangeRedis package pulls it in, much like Netty arrives with Lettuce. dotnet list package --include-transitive (or dotnet package list --include-transitive on the .NET 10 SDK) shows the versions that really resolved (details).

The server-side signals from the Java post, like the connectedclients dip during maintenance, memory pressure and evicted keys, work the same for .NET clients.

Final thoughts

Remember the spoiler from the Java post? Redis was mostly innocent. The .NET edition of that lesson is the same: a timeout is a symptom, not a diagnosis, and the diagnosis is usually sitting right there in the message.

If you only remember three things, make it these:

  1. Read the whole timeout message. The queue size, the bytes waiting and the thread pool numbers tell you whether Redis, the network or your own threads were the bottleneck.
  2. Await all the way down. Blocking on an async Redis call holds a thread hostage while the reply waits for a free thread of its own. Nobody wins that standoff.
  3. Be a great houseguest. Connect at startup, prove you’re ready before taking traffic, reconnect without drama, and ease off the CPU limit while you start up.

If you want the whole story, from the three patterns to the full set of testing recommendations, the Java post has it: When Redis times out in an application but Redis is fine: Lessons from a Real-World investigation. Managed services keep improving, so keep the official documentation bookmarked. Now go put your own client through its paces, and may your threads be plentiful and your timeouts few.

This post reflects the guidance available at the time of writing. Features and guidance can change, so double-check the official documentation before you plan around them.