Explainer
Can a slot demo sample measure RTP? Sample size explained

The short answer
A sample measures observed return for those recorded outcomes. It does not directly reveal theoretical RTP. Interpretation also needs the distribution, a stable configuration and a clearly defined observation unit.
In this article
A video or spreadsheet might report that a demo returned 93% over a short sequence. That can be an accurate description of the recorded sequence and still be poor evidence for the underlying theoretical percentage. There are two separate questions: was the ratio calculated correctly, and what can the sample tell us about the model?
Define one observation before counting it
For a simple fixed-input model, one observation can be the total amount returned by a complete game cycle divided by its input. A feature with several animations may still belong to that one cycle. Counting each animation as an independent observation would change the sample size without necessarily adding independent information.
The same issue arises if inputs differ. An unweighted average of per-event return percentages need not equal total returned divided by total input. Keep the raw numerator and denominator so the aggregation can be reproduced. For this explanation, every hypothetical observation has one unit of input and the same distribution.
Sampling error shrinks slowly
For independent, identically distributed observations with finite variance, the standard error of the mean is the population standard deviation divided by the square root of the sample size. OpenStax’s chapter review gives the statistical background. Standard error describes variation across repeated samples; it is not a guaranteed bound on one result.
Illustrative example
| Independent observations | Assumed standard deviation | Standard error of the mean |
|---|---|---|
| 100 | 5 return units | 0.50 units |
| 400 | 5 return units | 0.25 units |
| 10,000 | 5 return units | 0.05 units |
With one unit of input, 0.50 units of uncertainty in the mean corresponds to 50 percentage points in the return ratio. Quadrupling the number of observations halves the standard error in this example. The arithmetic does not tell us how many observations are sufficient for an unknown game: its distribution and the accuracy required have not been specified.
Rare outcomes can dominate the average
Imagine a model that returns 0 units with probability 99% and 90 units with probability 1%. Its expected return is 0.90 units. In 100 independent observations, the probability of seeing no 90-unit outcome is 0.99 to the power 100, approximately 36.6%. A zero-return sample is therefore compatible with this model even though its expected return is 90% of the one-unit input.
That original example is intentionally extreme. It demonstrates why an apparently uneventful record may omit a material part of an expected value. It also explains why assuming a convenient bell-shaped interval around a small sample can be misleading for a strongly skewed distribution.
A larger record does not fix a different configuration
- Record the exact build, mathematical configuration and rules.
- Define the complete cycle and identify any linked feature states.
- Retain totals, observation count and the sampling method.
- Do not silently combine sequences from different versions.
Even a well-recorded demo sample establishes facts about the tested environment. It does not prove that another environment uses the same configuration. A defensible report states what was observed, how it was counted and which configuration was documented. It leaves theoretical certification and cross-version equivalence as separate claims requiring separate evidence.
