Project

Testing a live poker table through a hostile network

A real-time poker table fails in ways a unit test rarely sees: a call that arrives twice, a phone that drops mid-hand and comes back a minute later, an action timer that fires in the same millisecond as the player's click, a server that dies halfway through writing a hand. I built a small No-Limit Hold'em table server and a test harness aimed at exactly those failures, then seeded seventeen bugs into it to check that the harness notices.

System under test
A Hold'em table server over WebSockets, in TypeScript
Checks
20 in six layers, plus tests over real sockets
Chaos
Latency, reordering, duplicates, disconnects and server crashes, reproducible by seed
Seeded bugs
17 of 17 caught; the real code passes every check
A terminal showing a 100-hand chaos run with 24 server crashes in which every client converged, and a table of seventeen seeded bugs with the kinds of checks that caught each one

The problem

The rules of poker are the easy part to test. The hard part of a live table is everything around them: several clients holding their own copy of the game, a network that delivers late, twice or never, and a server that has to stay the single source of truth while players reconnect, retry and race each other and the clock.

Those bugs rarely show up in a test that sends one message and checks one reply. They show up in the hundredth hand, when a retry lands after a timeout. So the harness plays whole games through a network built to misbehave, and checks what must stay true at every step.

The system under test

  • The rules as a pure function: a command goes in, and either a rejection or the next state and its events come out. The deck is shuffled from a seed, so a log of commands replays to identical events.
  • One service owns the table. It deduplicates commands by sender and command id, applies them one at a time, stores each change before publishing it, arms the action timers and sends every viewer a numbered event stream with other players' hole cards removed.
  • Clients keep their own copy of the table from those events. They apply them strictly in sequence order, drop duplicates, resume from their last event after a reconnect, and re-send unanswered commands under the same id so a retry can't act twice.
  • Every action names the hand and turn it was decided on, so a decision made on a stale view is rejected instead of landing on whatever is current.
  • If the server dies, a new process rebuilds the table from its log: it replays every command through the rules, refuses a log that doesn't match, restores the same event numbers and the results of commands already applied, and re-arms the timers. Clients carry on as after any reconnect.

Tech: TypeScript, Node.js, WebSockets (ws), Vitest, fast-check and GitHub Actions.

Six layers of checks

Reference
Outcomes written by hand from the rules: heads-up blind order, minimum raises and short all-ins, a three-way all-in split into main pot, side pot and returned bet, the odd chip of a split pot.
Property
fast-check generates 300 sessions of up to 400 commands, legal and not, and every invariant must hold after each one: chips conserved, one seat per player, busted players never dealt in, turns and streets only moving forward, nobody winning more than they could have lost.
Protocol
A retried command gets its first result and applies once; resuming after event k sends exactly k+1 onwards; hole cards reach only their owner; a timer that fires after its turn does nothing.
Concurrency
Two players racing for one seat, and a player's action landing together with their timeout, while storage takes a few milliseconds: exactly one wins each time.
Chaos
Whole games on a virtual clock through a network that delays, reorders, duplicates and drops. At the end every client, spectator included, must see exactly the server's table, and the command log must replay to the same events.
Recovery
The server is killed at each stage of a write: before it reaches the disk, on disk but not yet published, published but not yet delivered. The restored table must match exactly, a retry must get its original result and apply once, the turn timer must come back, and a tampered log must be refused.

The same seed gives the same run, message for message, so any failure comes with a seed that reproduces it. Across 800 simulated games with a server crash every few seconds, every client converged on the server's table and the log replayed exactly. The key guarantees are also tested over real sockets and real timers, including killing the server mid-hand and restarting it from its log file on the same port.

Testing the tests

Seventeen deliberate bugs can be switched on, one at a time, across the rules, the server, the client and crash recovery. Every check runs against every one of them. All seventeen are caught, and the table they produce says more about the checks than the pass count does.

  • One layer can hide another's bug. When the server resends one event too many on resume, the client's deduplication absorbs it and whole games look perfect. Only the protocol check, which holds the server to its own contract, sees it.
  • Invariants keep the state sane but don't know the rules. Accepting a raise below the minimum breaks no invariant, and the random sessions never notice. A reference case written from the rules does.
  • Races need a window. With instant storage, validating commands concurrently never goes wrong. The concurrency checks use storage that takes a moment, and the simulator randomizes it, so the lost update actually happens.
  • Whole-game simulation is the widest net, including both client bugs, but it can be blind to a specific timing. Publishing events before storing them was caught in only one of ten simulated seeds; a check that crashes the server at exactly that moment catches it every time.
  • Five of the bugs only exist after a restart. The most telling one: when the server forgets which commands it had already applied, the first symptom is a player being told their action failed when it had actually happened.

The simulator also found bugs of mine while I was building it. On a quiet table the bots never sat down, because the client didn't wake them after connecting. The automatic start of the next hand reused one id per hand number, so once a start failed for lack of players, the idempotency cache kept answering every later start with that same rejection and the table froze for good. And running the crash games against one of the seeded bugs showed that an error thrown inside a timer's command went unhandled, which in Node ends the whole process; with recovery, it would have become a crash loop.

What comes next

This is one table in one process, with play chips and no authentication, and a restart replays the whole log without snapshots. Next: a multi-table tournament with blind levels and players moved between tables without losing an event, and a load run with many tables at once under the same convergence checks.

Contact

I'm Pedro Morago, a Senior QA Engineer with a Mathematics degree and 10 years of poker, working remotely from Spain. If you'd like to talk about testing real-time or poker systems, email me at pedro@pedromorago.com.