Case study · Production · Khumbu Systems
Redis + Lettuce to Valkey Serverless + GLIDE
Our restaurant-integration platform ran its cache on provisioned ElastiCache Redis, sized for peak and paid for around the clock, while traffic peaks at breakfast, lunch and dinner and falls away in between. I led the move to ElastiCache Serverless on Valkey with AWS’s GLIDE client. What looked like a client swap turned into lessons about serverless economics, client behaviour and Spring’s lifecycle.
Outcome
| Metric | Before | After |
|---|---|---|
| Cache infrastructure cost | Provisioned, 24×7 | ~75% lower |
| ECPU consumed per day | 177M | 100M (−44%) |
| New connections opened | 1,000+ | ~0 |
| Cache hit rate (cache-aside) | — | 96–99% |
Why move at all
Licensing started the conversation: in 2024 Redis moved from BSD to RSALv2 and SSPL, and Valkey continued under BSD-3 with full protocol compatibility. Cost made the decision. With provisioned nodes the bill is identical at 500 requests per second at noon and 2 at 3 AM; Serverless charges for what is consumed, and scales without capacity reviews or node resizing. For a workload with sharp daily peaks and valleys, that was worth about 75% of the cache bill.
GLIDE replaced Lettuce not because Lettuce was failing, but for alignment: AWS builds it for Valkey and ElastiCache, its core is Rust on Tokio, and it removes Netty from the stack.
Spring does not know GLIDE
Spring Data Redis is built around RedisConnectionFactory, and GLIDE does not implement it, so there
is no drop-in replacement for @Cacheable. I built a cache on
AbstractValueAdaptingCache with its own cache manager. The design rule that mattered most: a cache
failure is a cache miss. If the cache is unavailable, the request reads from its source of truth and carries
on.
@Override
protected Object lookup(Object key) {
try {
String json = glideClient.get(buildRedisKey(key)).get();
return json == null ? null : objectMapper.readValue(json, Object.class);
} catch (Exception e) {
log.warn("Cache lookup failed for '{}' key='{}': {}", name, key, e.getMessage());
return null; // exception = cache miss. Never propagate.
}
}
The next surprise: reads came back as LinkedHashMap, not domain objects, because deserialising into
Object.class without type information falls back to maps. The fix was a dedicated
ObjectMapper with default typing, isolated from the application’s main mapper. Enabled globally, it
would have written @class fields into every REST response.
The bug that only appeared during deploys
Everything worked until shutdown. During a deploy, Spring drains in-flight SQS messages, and those handlers still
use the cache. But GLIDE registers its own JVM shutdown hook, which marks the client as closing the moment the JVM
begins to stop. Every cache call made during the drain failed with ClosingException. It looked like
a connection failure. It was lifecycle ordering.
The fix needs three coordinated changes. Remove any one and it breaks:
// 1. Spring must not close the client on its own
@Bean(destroyMethod = "")
public GlideClusterClient glideClient() throws Exception {
// 2. No GLIDE shutdown hook racing Spring's graceful shutdown
System.setProperty("glide.autoShutdownHook", "false");
// ... client configuration ...
}
@Bean
public SmartLifecycle glideClientLifecycle(GlideClusterClient glideClient) {
return new SmartLifecycle() {
// ...
// 3. We own the timing: close last, after the SQS listeners drain
@Override public void stop() { glideClient.close(); }
@Override public int getPhase() { return Integer.MIN_VALUE; }
};
}
SmartLifecycle stops higher phases first, so a phase of Integer.MIN_VALUE closes the
client after the SQS listeners have drained. After the change, the deploy-time exceptions disappeared.
Tuning for how Serverless bills
- ECPU down 44%, 177M → 100M a day, by gzipping large menu payloads before they reach the cache. Serverless bills processing by data transferred, so payload size is cost.
-
New connections 1,000+ → ~0. GLIDE’s multiplexed client replaced our custom connection pooling.
I traced the remaining intermittent
ClosingExceptionerrors to the Serverless proxy recycling connections. - Write-through → cache-aside, with invalidation driven by DynamoDB Streams and the same “exception = miss” contract: 96–99% hit rate.
- A limit worth knowing before production: ElastiCache Serverless allows at most 3,999 arguments per command, and exceeding it returns a bare TCP reset, not an error. Every call was audited before go-live.
The full write-up, with the complete configuration: From Redis + Lettuce to Valkey GLIDE, on Medium →