Tag

resilience testing

Azure, C#, Debugging

When Redis times out in an application but Redis is fine: The .NET Edition

Did you read When Redis times out in an application but Redis is fine: Lessons from a Real-World investigation and think, “Great story, but we’re a .NET shop”? Fair enough! This is the companion piece. The investigation happened on the Java side, with the Lettuce client, but the lessons travel well, and most of them map neatly to StackExchange.Redis, the main Redis client library for .NET. A quick recap, in case you haven’t read it. An application running on Kubernetes saw intermittent client-side Redis timeouts while Redis itself looked healthy. The clearest cause was freshly started pods, where connections to each shard were opened on first use just as CPU throttling slowed everything down. Two other patterns showed up too,…

Read more
Azure, Debugging

When Redis times out in an application but Redis is fine: Lessons from a Real-World investigation

This post is a friendly recap of a real Azure Managed Redis investigation Some incidents teach you more than any documentation page ever could. This one started with a few intermittent Redis timeouts. This grew into a deep investigation that touched Kubernetes CPU limits, Java client internals, cluster topology, and even the limits of what you can simulate on a managed service. One symptom, three very different stories. The customer stays anonymous, because the lessons matter more than the logo. What’s left is the part you can reuse: what happened, what we learned, and tips and tricks for testing. Along the way, we try to showcase how you can look beyond the error message. Settle in, because this one has…

Read more