All work

Case study · Side project · 2026

Real-time multiplayer chess platform

The interesting engineering here is not chess. It is keeping a turn-based mutation of shared state correct under concurrency, network failure, pod death and rolling deploys, with a clock that decides outcomes and so cannot be wrong. Then measuring whether it is, rather than claiming it.

Role
Sole engineer: design, build, test, deploy
Backend
Java 25, Spring Boot 4, PostgreSQL, Valkey, SQS
Infrastructure
AWS ECS Fargate, Terraform, Kubernetes (kind), Docker
Source
github.com/Abhinav2777/chess-platform
Two real browsers against the real server: seek, pair through the Valkey queue, scholar’s mate. The rating change (±16) reaches both players asynchronously, via outbox → SQS → rating worker → WebSocket.

What I built

A modular monolith: one image with three roles (API, worker, migrate) and module boundaries enforced by ArchUnit tests. API instances are stateless: any pod can serve any game, and a player can reconnect to a different pod mid-game. Moves arrive over a raw WebSocket protocol and each one is a single PostgreSQL transaction. Valkey carries the cross-pod work (fanning moves out to the right sockets, presence and the matchmaking queue), and an outbox relay hands finished games to a worker that computes ratings.

Chess platform architecture Browsers connect over REST and WebSocket through a load balancer to N API instances. Each move is one PostgreSQL transaction. API instances share fanout, presence and the seek queue through Valkey. A worker relays the transactional outbox to SQS, consumes rating events, and publishes rating updates through Valkey. REST · WS 1 txn / move pub/sub outbox RATING_UPDATED Browser React SPA ALB / ingress load balancer API × N game · realtime matchmaking · identity stateless PostgreSQL the only source of truth Worker outbox relay · ratings Valkey fanout · presence · seeks SQS + DLQ
One deployable, three roles. PostgreSQL is the source of truth; everything in Valkey can be rebuilt. The same image runs on AWS (ALB, ECS Fargate, RDS, ElastiCache, SQS) and on Kubernetes (ingress-nginx, in-cluster data).

Measured

Every figure comes from a report in the repository stating its date, environment and method. After every load test, each game is checked move by move against the server: every move a player saw acknowledged must be stored, at its position, unchanged, and nothing else.

WhatResultReport
Concurrency invariant 16 requests for the same move, 100 rounds: exactly one winner every time Project state
1,000 WebSocket connections (kind) 500 live games; server p99 6.7 ms, flat from 100 to 1,000 sockets Baseline
Rolling deploy during 40 live games 40/40 games consistent, 0 abnormal closes, reconnect p99 498 ms Rolling deploy
Stress that crashed both pods Memory sized from measurement: OOM-killed 4× per pod → 0 restarts Optimisation 01
Sign-up burst, AWS Fargate (0.5 vCPU) 30/50 games → 50/50; 265 error log lines → 0 Optimisation 02
PostgreSQL frozen 30 s, 20 live games 20/20 games consistent; longest wait 33.5 s → 2.9 s; ERROR lines 127 → 1 Failure drills
Queue or worker down 60 s Games unaffected; ratings delivered within 3 s / 14 s of recovery Failure drills
Unauthenticated 100 MB request bodies Heap 58 MB → 2.4 GB (enough to end an AWS task) → now 413 at 16 KB Security review

Measured on AWS up to about 100 concurrent sockets; I stopped the larger run on cost. The 1,000-socket figure is from kind on one laptop, and is labelled that way wherever it appears.

Decisions

Each of these is an Architecture Decision Record in the repository, with the alternatives I rejected and what would make me revisit it.

  1. The clock does not tick.

    Remaining time is computed from three persisted columns and PostgreSQL’s now(). It is identical from any pod, survives reconnecting to a different instance, and cannot drift, because there is no timer to drift. ADR-006

  2. No distributed lock.

    Move safety comes from an idempotency key, an optimistic version column and a composite primary key, all inside one database transaction. A Redis lock would add a TTL-expiry failure mode without removing any of those. ADR-005

  3. PostgreSQL is the only source of truth.

    Valkey holds nothing that cannot be rebuilt, so it can be restarted mid-game. Fanout, presence and the seek queue degrade; games do not. ADR-004

  4. Exactly-once effects on at-least-once delivery.

    A transactional outbox, SQS Standard, and a consumer that records each event in the same transaction as its effect. Ratings are never applied twice, and a game never waits for them. ADR-008

  5. A server can leave mid-game.

    On shutdown a pod turns readiness off and closes every socket with a reconnect hint; the client resumes on another pod from a snapshot. ADR-024

A rolling restart, mid-game

A real kubectl rollout restart replacing both API pods during a game. Each old pod drains (readiness off, every socket closed with 1001 GOING_AWAY); the browsers reconnect to a new pod and resume from a snapshot. In this take they reconnected in 453 ms and 275 ms. Waits are shown at 4×, marked on screen.

Under load the same scenario ran with 40 live games and 80 sockets: 116 sockets closed with GOING_AWAY, 0 abnormal closes, 0 moves rejected, all 40 games consistent, reconnect p99 498 ms.

What measuring found

The most useful part of the project was not building it but trying to break it. Each of these was found by a test, a drill or a recording, then fixed with before-and-after numbers.

1. A burst of sign-ups stopped the whole server

On Fargate at 0.5 vCPU, a 50-game burst failed 20 games with connection-pool timeouts, though no password hash held a connection any more. The clue was in the timeout itself: a thread allowed to wait 3,000 ms for a connection woke after 4,795 ms. Threads were not being scheduled.

Requests run on virtual threads, which share a few carrier threads and give one up only when they block. bcrypt never blocks: it is hundreds of milliseconds of pure CPU. While hashes occupied every carrier, a request that had already sent its SQL waited behind them, with its connection still checked out. A local experiment pinned to one CPU showed it directly: a 5 ms request finished 1,854 ms late (p50) when hashes shared its carriers, and 0 ms late when they ran on a platform thread. On my 16-core laptop it never appeared, which is why the kind baseline missed it.

The fix routes hashing through a bounded pool of platform threads, which the OS preempts, so request carriers keep getting CPU. When the queue is full, the request is refused at once with 503 SERVER_BUSY and Retry-After. The refusal depends only on the queue, not on whether the account exists, so it does not reopen a user-enumeration timing side channel. Result: 50/50 games, 265 error lines → 0, with the CPU still at 100%. Nothing else fails any more. The startup log also corrected one of my own predictions: the JVM sees 2 processors on a 0.5 vCPU task, not 1.

2. A frozen database hung every move for the whole outage

I froze PostgreSQL for 30 seconds under 20 live games. My first attempt sent SIGSTOP from inside the container, where the postmaster is PID 1, and Linux ignores that signal sent to a namespace’s init from inside it. Only the backends froze, and the postmaster kept accepting connections. I noticed because the sampler’s own psql kept answering. Sent from the node, it worked.

The drill found two problems. pgjdbc’s socket timeout defaults to forever, so a move waited 33.5 s, the whole freeze; a silent network partition would have hung for TCP’s retransmission timeout, about fifteen minutes. And every refused move reached players as INTERNAL and was logged at ERROR, 127 lines in 30 seconds. After a 10 s socket timeout and a typed SERVICE_UNAVAILABLE answer on both REST and WebSocket: longest wait 2.9 s, ERROR lines 127 → 1, and still 20/20 games consistent.

One trade-off I kept on purpose: the database is part of readiness, so an outage makes every pod unready and new visitors see a 503. In exchange, a new version that cannot reach the database never receives traffic, and the rollout stalls instead of replacing healthy pods.

3. A move vanished in the instant a pod drained

While recording the rolling-restart demo, a move made at the exact moment a pod drained was lost with its socket: the server never received it, and the player’s move silently disappeared. The browser now keeps its last unacknowledged move and re-sends it after reconnecting, with the same clientMoveId. The server’s idempotency key makes the re-send safe even if the original did land. An end-to-end test pins it.

4. One unauthenticated request could end an AWS task

In an OWASP API Top 10 review: Spring parses a request body before the controller and its rate limit run, and nothing limited the size. Four concurrent 100 MB bodies took the heap from 58 MB to 2,463 MB, against a 512 MB heap on AWS. A filter first in the chain now refuses any body over 16 KB with 413, in 2 ms, without reading it.

Limits, stated

  • Fargate is measured to about 100 concurrent sockets. 1,000 sockets is measured only on kind, on one machine.
  • Single-AZ RDS, by cost decision; a Multi-AZ failover has not been drilled.
  • The AWS environment is applied for test sessions and then destroyed. It is not running permanently.
  • Sign-up capacity on that deployment is about one per second (an estimate from the measurements). bcrypt, not chess, sets that limit.

The repository has the full write-ups: 26 ADRs, load-test reports with their raw method, the failure drills, the security review, diagrams and a design Q&A. Read it on GitHub →