Case study · Side project · 2026
Real-time multiplayer chess platform
The interesting engineering here is not chess. It is keeping a turn-based mutation of shared state correct under concurrency, network failure, pod death and rolling deploys, with a clock that decides outcomes and so cannot be wrong. Then measuring whether it is, rather than claiming it.
What I built
A modular monolith: one image with three roles (API, worker, migrate) and module boundaries enforced by ArchUnit tests. API instances are stateless: any pod can serve any game, and a player can reconnect to a different pod mid-game. Moves arrive over a raw WebSocket protocol and each one is a single PostgreSQL transaction. Valkey carries the cross-pod work (fanning moves out to the right sockets, presence and the matchmaking queue), and an outbox relay hands finished games to a worker that computes ratings.
Measured
Every figure comes from a report in the repository stating its date, environment and method. After every load test, each game is checked move by move against the server: every move a player saw acknowledged must be stored, at its position, unchanged, and nothing else.
| What | Result | Report |
|---|---|---|
| Concurrency invariant | 16 requests for the same move, 100 rounds: exactly one winner every time | Project state |
| 1,000 WebSocket connections (kind) | 500 live games; server p99 6.7 ms, flat from 100 to 1,000 sockets | Baseline |
| Rolling deploy during 40 live games | 40/40 games consistent, 0 abnormal closes, reconnect p99 498 ms | Rolling deploy |
| Stress that crashed both pods | Memory sized from measurement: OOM-killed 4× per pod → 0 restarts | Optimisation 01 |
| Sign-up burst, AWS Fargate (0.5 vCPU) | 30/50 games → 50/50; 265 error log lines → 0 | Optimisation 02 |
| PostgreSQL frozen 30 s, 20 live games | 20/20 games consistent; longest wait 33.5 s → 2.9 s; ERROR lines 127 → 1 | Failure drills |
| Queue or worker down 60 s | Games unaffected; ratings delivered within 3 s / 14 s of recovery | Failure drills |
| Unauthenticated 100 MB request bodies | Heap 58 MB → 2.4 GB (enough to end an AWS task) → now 413 at 16 KB | Security review |
Measured on AWS up to about 100 concurrent sockets; I stopped the larger run on cost. The 1,000-socket figure is from kind on one laptop, and is labelled that way wherever it appears.
Decisions
Each of these is an Architecture Decision Record in the repository, with the alternatives I rejected and what would make me revisit it.
-
The clock does not tick.
Remaining time is computed from three persisted columns and PostgreSQL’s now(). It is identical from any pod, survives reconnecting to a different instance, and cannot drift, because there is no timer to drift. ADR-006
-
No distributed lock.
Move safety comes from an idempotency key, an optimistic version column and a composite primary key, all inside one database transaction. A Redis lock would add a TTL-expiry failure mode without removing any of those. ADR-005
-
PostgreSQL is the only source of truth.
Valkey holds nothing that cannot be rebuilt, so it can be restarted mid-game. Fanout, presence and the seek queue degrade; games do not. ADR-004
-
Exactly-once effects on at-least-once delivery.
A transactional outbox, SQS Standard, and a consumer that records each event in the same transaction as its effect. Ratings are never applied twice, and a game never waits for them. ADR-008
-
A server can leave mid-game.
On shutdown a pod turns readiness off and closes every socket with a reconnect hint; the client resumes on another pod from a snapshot. ADR-024
A rolling restart, mid-game
kubectl rollout restart replacing both API pods during a game. Each old pod drains (readiness
off, every socket closed with 1001 GOING_AWAY); the browsers reconnect to a new pod and resume from a
snapshot. In this take they reconnected in 453 ms and 275 ms. Waits are shown at 4×, marked on screen.
Under load the same scenario ran with 40 live games and 80 sockets: 116 sockets closed with GOING_AWAY, 0 abnormal closes, 0 moves rejected, all 40 games consistent, reconnect p99 498 ms.
What measuring found
The most useful part of the project was not building it but trying to break it. Each of these was found by a test, a drill or a recording, then fixed with before-and-after numbers.
1. A burst of sign-ups stopped the whole server
On Fargate at 0.5 vCPU, a 50-game burst failed 20 games with connection-pool timeouts, though no password hash held a connection any more. The clue was in the timeout itself: a thread allowed to wait 3,000 ms for a connection woke after 4,795 ms. Threads were not being scheduled.
Requests run on virtual threads, which share a few carrier threads and give one up only when they block. bcrypt never blocks: it is hundreds of milliseconds of pure CPU. While hashes occupied every carrier, a request that had already sent its SQL waited behind them, with its connection still checked out. A local experiment pinned to one CPU showed it directly: a 5 ms request finished 1,854 ms late (p50) when hashes shared its carriers, and 0 ms late when they ran on a platform thread. On my 16-core laptop it never appeared, which is why the kind baseline missed it.
The fix routes hashing through a bounded pool of platform threads, which the OS preempts, so request carriers
keep getting CPU. When the queue is full, the request is refused at once with 503 SERVER_BUSY and
Retry-After. The refusal depends only on the queue, not on whether the account exists, so it does
not reopen a user-enumeration timing side channel. Result: 50/50 games, 265 error lines → 0,
with the CPU still at 100%. Nothing else fails any more. The startup log also corrected one of my own
predictions: the JVM sees 2 processors on a 0.5 vCPU task, not 1.
2. A frozen database hung every move for the whole outage
I froze PostgreSQL for 30 seconds under 20 live games. My first attempt sent SIGSTOP from inside
the container, where the postmaster is PID 1, and Linux ignores that signal sent to a namespace’s init from
inside it. Only the backends froze, and the postmaster kept accepting connections. I noticed because the
sampler’s own psql kept answering. Sent from the node, it worked.
The drill found two problems. pgjdbc’s socket timeout defaults to forever, so a move waited 33.5 s, the whole
freeze; a silent network partition would have hung for TCP’s retransmission timeout, about fifteen minutes. And
every refused move reached players as INTERNAL and was logged at ERROR, 127 lines in 30 seconds.
After a 10 s socket timeout and a typed SERVICE_UNAVAILABLE answer on both REST and WebSocket:
longest wait 2.9 s, ERROR lines 127 → 1, and still 20/20 games consistent.
One trade-off I kept on purpose: the database is part of readiness, so an outage makes every pod unready and new visitors see a 503. In exchange, a new version that cannot reach the database never receives traffic, and the rollout stalls instead of replacing healthy pods.
3. A move vanished in the instant a pod drained
While recording the rolling-restart demo, a move made at the exact moment a pod drained was lost with its socket:
the server never received it, and the player’s move silently disappeared. The browser now keeps its last
unacknowledged move and re-sends it after reconnecting, with the same clientMoveId. The server’s
idempotency key makes the re-send safe even if the original did land. An end-to-end test pins it.
4. One unauthenticated request could end an AWS task
In an OWASP API Top 10 review: Spring parses a request body before the controller and its rate limit run, and
nothing limited the size. Four concurrent 100 MB bodies took the heap from 58 MB to 2,463 MB, against a 512 MB
heap on AWS. A filter first in the chain now refuses any body over 16 KB with 413, in 2 ms, without
reading it.
Limits, stated
- Fargate is measured to about 100 concurrent sockets. 1,000 sockets is measured only on kind, on one machine.
- Single-AZ RDS, by cost decision; a Multi-AZ failover has not been drilled.
- The AWS environment is applied for test sessions and then destroyed. It is not running permanently.
- Sign-up capacity on that deployment is about one per second (an estimate from the measurements). bcrypt, not chess, sets that limit.
The repository has the full write-ups: 26 ADRs, load-test reports with their raw method, the failure drills, the security review, diagrams and a design Q&A. Read it on GitHub →