All work

Case study · Production · Khumbu Systems

Redis + Lettuce to Valkey Serverless + GLIDE

Our restaurant-integration platform ran its cache on provisioned ElastiCache Redis, sized for peak and paid for around the clock, while traffic peaks at breakfast, lunch and dinner and falls away in between. I led the move to ElastiCache Serverless on Valkey with AWS’s GLIDE client. What looked like a client swap turned into lessons about serverless economics, client behaviour and Spring’s lifecycle.

Role
Led the migration
Stack
Spring Boot 3, ElastiCache Serverless (Valkey), GLIDE
Also
SQS, DynamoDB Streams, S3
Write-up
Medium, June 2026

Outcome

MetricBeforeAfter
Cache infrastructure cost Provisioned, 24×7 ~75% lower
ECPU consumed per day 177M 100M (−44%)
New connections opened 1,000+ ~0
Cache hit rate (cache-aside) — 96–99%

Why move at all

Licensing started the conversation: in 2024 Redis moved from BSD to RSALv2 and SSPL, and Valkey continued under BSD-3 with full protocol compatibility. Cost made the decision. With provisioned nodes the bill is identical at 500 requests per second at noon and 2 at 3 AM; Serverless charges for what is consumed, and scales without capacity reviews or node resizing. For a workload with sharp daily peaks and valleys, that was worth about 75% of the cache bill.

GLIDE replaced Lettuce not because Lettuce was failing, but for alignment: AWS builds it for Valkey and ElastiCache, its core is Rust on Tokio, and it removes Netty from the stack.

Spring does not know GLIDE

Spring Data Redis is built around RedisConnectionFactory, and GLIDE does not implement it, so there is no drop-in replacement for @Cacheable. I built a cache on AbstractValueAdaptingCache with its own cache manager. The design rule that mattered most: a cache failure is a cache miss. If the cache is unavailable, the request reads from its source of truth and carries on.

@Override
protected Object lookup(Object key) {
    try {
        String json = glideClient.get(buildRedisKey(key)).get();
        return json == null ? null : objectMapper.readValue(json, Object.class);
    } catch (Exception e) {
        log.warn("Cache lookup failed for '{}' key='{}': {}", name, key, e.getMessage());
        return null; // exception = cache miss. Never propagate.
    }
}

The next surprise: reads came back as LinkedHashMap, not domain objects, because deserialising into Object.class without type information falls back to maps. The fix was a dedicated ObjectMapper with default typing, isolated from the application’s main mapper. Enabled globally, it would have written @class fields into every REST response.

The bug that only appeared during deploys

Everything worked until shutdown. During a deploy, Spring drains in-flight SQS messages, and those handlers still use the cache. But GLIDE registers its own JVM shutdown hook, which marks the client as closing the moment the JVM begins to stop. Every cache call made during the drain failed with ClosingException. It looked like a connection failure. It was lifecycle ordering.

The fix needs three coordinated changes. Remove any one and it breaks:

// 1. Spring must not close the client on its own
@Bean(destroyMethod = "")
public GlideClusterClient glideClient() throws Exception {
    // 2. No GLIDE shutdown hook racing Spring's graceful shutdown
    System.setProperty("glide.autoShutdownHook", "false");
    // ... client configuration ...
}

@Bean
public SmartLifecycle glideClientLifecycle(GlideClusterClient glideClient) {
    return new SmartLifecycle() {
        // ...
        // 3. We own the timing: close last, after the SQS listeners drain
        @Override public void stop() { glideClient.close(); }
        @Override public int getPhase() { return Integer.MIN_VALUE; }
    };
}

SmartLifecycle stops higher phases first, so a phase of Integer.MIN_VALUE closes the client after the SQS listeners have drained. After the change, the deploy-time exceptions disappeared.

Tuning for how Serverless bills

  • ECPU down 44%, 177M → 100M a day, by gzipping large menu payloads before they reach the cache. Serverless bills processing by data transferred, so payload size is cost.
  • New connections 1,000+ → ~0. GLIDE’s multiplexed client replaced our custom connection pooling. I traced the remaining intermittent ClosingException errors to the Serverless proxy recycling connections.
  • Write-through → cache-aside, with invalidation driven by DynamoDB Streams and the same “exception = miss” contract: 96–99% hit rate.
  • A limit worth knowing before production: ElastiCache Serverless allows at most 3,999 arguments per command, and exceeding it returns a bare TCP reset, not an error. Every call was audited before go-live.

The full write-up, with the complete configuration: From Redis + Lettuce to Valkey GLIDE, on Medium →