Project
Testing poker math when there is no perfect oracle
A solver or a trainer gives you numbers: equities, frequencies, expected values. For a handful of spots you can check them by hand; for the rest there is no answer key, only relationships that must hold. This lab applies that idea at two levels: a small poker equity engine with a suite of checks that each name their oracle, and a validator for exported solver strategies with a regression diff between versions. Seeded bugs and corrupted files show what each check catches.
- Layer 1
- An all-in equity engine in pure Python, with 15 checks across six kinds of oracle
- Layer 2
- A validator for exported strategies: 13 rules, suit isomorphism and a regression diff
- Tests
- 233, including Hypothesis-generated strategy files
- Seeded defects
- 9 engine bugs and 17 file corruptions, all caught

The problem
Most tests compare an output with an expected value. For a poker engine that only works for the easy cases: a royal flush beats a full house, two identical hands split. Ask for the equity of a flush draw against an overpair on a given flop and nobody has the answer written down. The same is true, at a much bigger scale, of a solver's frequencies or a model's recommendations.
So the question becomes what must be true of any correct answer, and how to check that without knowing the answer.
Layer 1: the equity engine
A pure-Python engine: a seven-card hand evaluator, exact equity that enumerates every remaining board, and Monte Carlo equity that samples boards and reports a standard error with every estimate. Exact equity is accumulated in integer shares of the pot, so exactness can be asserted without floating-point tolerance.
Six kinds of oracle
Every check is written against an interchangeable engine and names the kind of oracle it relies on.
- Reference
- Hand-written expected values, only where they are beyond argument: textbook hand categories, kickers, identical hands splitting.
- Differential
- 3,000 random seven-card matchups must be ordered the same way as by an independent evaluator.
- Invariant
- Card order doesn't matter, shares add up to the pot exactly, a complete board has no uncertainty, sampled boards only use legal cards.
- Metamorphic
- Renaming suits changes nothing, swapping players swaps equities, and flop equity is exactly the average of the equities after each turn card.
- Statistical
- Monte Carlo estimates fall within 4.5 standard errors of the exact value, and quadrupling the samples roughly halves the error.
- Regression
- Exact equities for 11 spots, recorded only after the independent evaluator agreed with every one.
On top of them, Hypothesis generates random hands and spots for the same properties, plus one more: adding a seventh card can never make the best hand worse.
Testing the engine’s tests
A suite for a system without an answer key has to answer a question about itself: if the engine were wrong in this way, would anything notice? Nine versions of the engine each carry one deliberate bug, and every check runs against every one of them. All nine are caught.
- Sampling the board with replacement moves the estimates so little that they still land inside their error bars. The invariant on the sampled cards catches it directly.
- Giving every tie to the first player keeps the shares adding up to the pot, so that invariant passes. Swapping the players exposes it.
- Never dealing the last card of the deck produces plausible equities everywhere, but breaks the law of total probability: flop equity stops being the average of the turn equities.
- Differential testing catches every evaluator bug, and it is only as good as the other implementation. That's why the baselines are cross-checked against it once, rather than trusted on their own.
Layer 2: validating exported strategies
A solver doesn't hand you one number. It hands you a strategy: for every hand in a range, how often to take each action at a decision, and what each action is worth. The second layer checks files like that against a documented JSON format and a set of rules any correct export must satisfy, and points at the exact entry that breaks one, such as nodes[0].combos['2h2c'].freqs: sums to 1.07.
- Contract
- Strict parsing into typed data, with errors located down to the combo and field. Duplicate combos are kept so a rule can report them.
- 13 rules
- Frequencies in range and summing to one, finite EVs, legal actions and bet sizes, valid and unique combos, no combo using a board card, overall EV consistent with the action EVs, range weights consistent with the actions that led to the node, and no action left at 0% while its EV beats everything played.
- Suit isomorphism
- When two suits play the same role on the board, hands that differ only by swapping them must play the same way. Across two files, a strategy for a board and one for its suit-relabelled twin must map onto each other.
- Regression diff
- Baseline against candidate: per-hand frequency and EV shifts, range-weighted shifts per action, hands added or removed, and tree changes, against configurable thresholds, with a Markdown report and a failing exit code for CI.
The strategy files in the repository are synthetic, generated or written by hand for the lab. None comes from a real solver, and the lab isn't one.
Tech: Python, pytest, Hypothesis and GitHub Actions.
Testing the validator
Seventeen corruptions of a valid strategy file, from frequencies that sum to 1.07 to one hand drifting away from its suit-swapped twin, each run through validate, diff and the isomorphism check. All seventeen are caught, and the table says which tool is needed for what.
- A hand that uses a board card, and the same hand written twice in a different card order, are each caught by one rule only. The diff can't see either, so validation has to run before comparing versions.
- A valid file with a large strategy change passes every rule. Only the regression diff flags it.
- Hypothesis found a case I had missed: two hands that swap strategies leave every range-weighted total unchanged. Only the per-hand threshold catches it, and there is now a test for it.
What comes next
Next in the same repository: a heads-up push/fold equilibrium for short stacks, checked by exploitability, which would give the strategy validator a real solver to point at and a stronger oracle than its best-response rule; and range against range equity, checked against its hand-against-hand parts.

Contact
I'm Pedro Morago, a Senior QA Engineer with a Mathematics degree and 10 years of poker, working remotely from Spain. If you'd like to talk about testing poker or quantitative systems, email me at pedro@pedromorago.com.