4.4 Hypothesis testing and significance
p-values, t-tests, confidence intervals and sample size.
Key points
A hypothesis test asks whether an observed difference could easily happen by chance. The p-value is that chance under the 'no difference' assumption; a small p-value means significance.
What NVIDIA says (1)
“The biggest result is that, across all attempts, both of the lower-latency conditions (25 ms and 55 ms) improved the number of targets eliminated (Figure 3), a difference that was found to be statistically significant in pairwise t-tests ( p-value << 0.001).”
A confidence interval is a range that likely contains the true value. More trials shrink it; too few trials leave comparisons inconclusive.
What NVIDIA says (2)
“A single success rate on N rollouts tells you almost nothing about how confident you should be in a policy’s true performance.”
“Most published benchmarks do not run a sufficient number of rollouts to achieve statistical significance when comparing the performance of two policies.”
The width of a confidence interval shrinks roughly with the square root of the sample size. So precision gets expensive.
What NVIDIA says (1)
“Narrowing the confidence interval from 10 to 2 percentage points requires roughly 15x more rollouts (70 to 1,030).”
Key terms
- p-value: The chance of seeing a result this extreme if there were really no effect; small values suggest a real effect.
- Confidence interval: A range that likely contains the true value; more samples make it narrower.
Sample question
A study compares players under lower and higher latency and reports a pairwise t-test with p-value << 0.001. What does that mean?
Show the answer
Answer: The difference is very unlikely to be due to chance alone, so it is statistically significant
A hypothesis test asks whether an observed difference could easily happen by chance. The p-value is that chance under the 'no difference' assumption; a small p-value means significance.
What NVIDIA says (1)
“The biggest result is that, across all attempts, both of the lower-latency conditions (25 ms and 55 ms) improved the number of targets eliminated (Figure 3), a difference that was found to be statistically significant in pairwise t-tests ( p-value << 0.001).”
Practice 4.4 (3 questions) Full Descriptive Analysis and Visualization guide
← 4.3 Choosing the right plot · 4.5 Patterns, trends and relationships →