What a game integrity audit actually measures

The short answer: "integrity audit" covers three genuinely different questions, and they need different data, different methods and different evidence standards. Confusing them is why reviews so often produce a number nobody can act on. The three are: is the equipment fair, are these results explainable by chance, and is someone betting as though they know what is coming.

The three questions

1. Is the equipment fair?

The subject is a physical object: a wheel, a shoe, a shuffler, a die. The question is whether its outcomes match the distribution its design implies.

This is the most self-contained of the three. It needs only outcomes and the identity of the equipment, no player information at all, and the expected distribution is known exactly in advance: a single-zero wheel should produce each of 37 pockets equally often. Comparing observed against expected is a goodness-of-fit problem.

It is also the slowest. Equipment defects are small effects, so the samples required run to thousands of outcomes, accumulated over weeks. The upside is that the data is cheap to collect and keeps its value. A wheel's history is as useful as its present.

Evidence standard: a whole-equipment test, plus a view of where the departure sits, mapped to the physical object rather than the felt layout. A finding here directs a maintenance inspection; it does not replace one.

2. Are these results explainable by chance?

The subject is a player, or a table, or a session. The question is whether the outcome sits inside the range the game's own probabilities allow.

Every game has a known house edge and a known variance. From those you can say what range of results a given volume of play should produce. A player far outside that range is not proof of anything, but it is a well-defined reason to look. Importantly, a player who feels extraordinary to the pit is very often inside the expected range once the volume is accounted for.

The volume is the part that gets skipped. A player up a large sum over forty hands and a player up the same sum over four thousand are completely different observations. Without wagers and hand counts, a win figure alone cannot be assessed at all.

Evidence standard: results measured against the game's own probabilities, with the wagered volume in the denominator, and a confidence interval attached. A result should always be reported with the uncertainty around it, because a point estimate on its own invites over-reading.

3. Is someone betting as though they know what is coming?

The subject is the relationship between a player's wagers and the state of the game. The question is whether bet sizing moves with the edge.

This is the strongest of the three when it is available, because it does not depend on the player winning. Variance is loud; over any realistic sample an informed player can easily be down. But a player who consistently raises their stake when the odds have quietly shifted in their favour, and lowers it when they have not, is displaying something that luck does not produce. That relationship is the signal, not the win.

It also needs the most data: outcomes, wagers, and enough of the card or spin sequence to establish what the edge was at the moment each bet was placed. On shoe-dealt games that means the deal stream, not just results.

Evidence standard: a persistent relationship across a sustained run, not a handful of well-timed bets. Anyone will look prescient over ten hands.

What counts as evidence

The distinction that matters to a surveillance team is between a result and a finding.

A result is a number. A finding is a number with its assumptions, its sample size and its uncertainty attached, produced by a method that was chosen before the data was examined. The difference determines whether it survives contact with a compliance officer, a regulator, or a lawyer.

Three things make the difference in practice:

The method was fixed in advance. Choosing an analysis after seeing which one produces an interesting answer is how false positives are manufactured. If a threshold was set after the fact, say so.

The sample size is stated. Not as a footnote, but as part of the result. "No significant deviation across 400 spins" and "no significant deviation across 40,000 spins" are very different statements, and only one of them is reassuring.

Uncertainty is reported. An interval, and a plain-English sentence about what it means. A tool that never says "not enough data yet" is not doing statistics, and a report that never says it is not being honest with its reader.

The sample sizes that make a conclusion safe

There is no single number, because it depends on how large an effect you are trying to detect. But the shape of the answer is consistent, and it is worth internalising:

QuestionTypical working sampleWhy
Equipment fairnessThousands of outcomesThe effect is small relative to natural spread
Outcome significanceHundreds of rounds, with wagersDepends on the game's variance and the size of the anomaly
Wager–edge relationshipA sustained run across sessionsShort runs of good timing are common by chance

The general rule: the smaller the effect you want to detect, the more data you need, and the effects that matter in game protection are mostly small. Anything large enough to see in a hundred rounds was probably visible without statistics.

What this is not

An integrity audit does not establish intent, and it does not establish what happened at the table. Reviewing footage does that. The audit establishes whether what happened is consistent with chance, which is a different job, and the one that converts a suspicion into something defensible.

Nor does it replace judgement. A finding is an input to a decision made by people who know the floor, the game and the context. Its value is that it makes the basis of that decision explicit and reviewable afterwards.

Getting started without a project

The data these questions need is data most surveillance rooms already keep, usually in a spreadsheet. Winning numbers per spin. Outcomes per coup. Hands with the wagers attached. Players identified by whatever code your team already uses.

That is worth saying plainly, because "integrity analytics" tends to imply an integration project. The equipment question in particular needs nothing but outcomes and an equipment identifier, which means a wheel audit can begin from a logbook.


EdgeTrack Lab runs all three questions from the records your team already keeps: outcome significance against the game's own probabilities, wager–edge correlation on composition-dependent side bets, and equipment integrity from outcome data alone, all with confidence intervals on the readouts and an audit-grade PDF at the end. Manual entry, batch paste or file import, with any equipment. Book a demo.

Sources

  • The statistical methods underlying all three questions are standard: goodness-of-fit testing for distributions, confidence intervals for proportions, and correlation measures for the wager–edge relationship, all covered in any general statistics text.
  • House-edge and variance figures for individual games and side bets are published and independently checkable. Wizard of Odds is the most widely cited public reference.
  • Game rules and permitted bets come from published regulator and government sources for each jurisdiction, which is also what makes an expected distribution defensible in a regulator conversation.

Common questions

What is a casino game integrity audit?
A structured statistical review asking whether a game's outcomes, a player's results, or a piece of equipment behave the way the published rules and probabilities say they should, with the working documented so a finding can support an incident file or a regulator conversation.
What data does an integrity audit need?
Whatever the game already emits: winning numbers per spin, outcomes per coup, hands with the wagers attached. Most surveillance teams already keep this in a spreadsheet or logbook. Equipment audits need only outcomes and the equipment identity.
How is this different from what a surveillance team already does?
Reviewing footage establishes what happened. An integrity audit establishes whether what happened is consistent with chance. Those answer different questions, and the second one is what turns a suspicion into something defensible.
What if the sample is too small to conclude anything?
Then that is the finding, and it should be stated. A great deal of harm comes from acting on results that look dramatic but sit inside normal variation. A tool that never reports 'not enough data yet' is not doing statistics.

Get started

See what your tables actually did.

Hardware-agnostic · manual entry, paste or file import · no player PII required