4.4 Hypothesis testing and significance

NCA-ADS · Descriptive Analysis and Visualization (13% of the exam) · Official objective: “Hypothesis testing and statistical significance evaluation”

p-values, t-tests, confidence intervals and sample size.

Key points

  1. A hypothesis test asks whether an observed difference could easily happen by chance. The p-value is that chance under the 'no difference' assumption; a small p-value means significance.

    What NVIDIA says (1)

    “The biggest result is that, across all attempts, both of the lower-latency conditions (25 ms and 55 ms) improved the number of targets eliminated (Figure 3), a difference that was found to be statistically significant in pairwise t-tests ( p-value << 0.001).”

    — Improving Player Performance with Low Latency as Evident from FPS Aim Trainer Experiments

  2. A confidence interval is a range that likely contains the true value. More trials shrink it; too few trials leave comparisons inconclusive.

    What NVIDIA says (2)

    “A single success rate on N rollouts tells you almost nothing about how confident you should be in a policy’s true performance.”

    — How to Evaluate General-Purpose Robot Policies for Real-World Deployment

    “Most published benchmarks do not run a sufficient number of rollouts to achieve statistical significance when comparing the performance of two policies.”

    — How to Evaluate General-Purpose Robot Policies for Real-World Deployment

  3. The width of a confidence interval shrinks roughly with the square root of the sample size. So precision gets expensive.

    What NVIDIA says (1)

    “Narrowing the confidence interval from 10 to 2 percentage points requires roughly 15x more rollouts (70 to 1,030).”

    — How to Evaluate General-Purpose Robot Policies for Real-World Deployment

Key terms

Sample question

A study compares players under lower and higher latency and reports a pairwise t-test with p-value << 0.001. What does that mean?

Show the answer

Answer: The difference is very unlikely to be due to chance alone, so it is statistically significant

A hypothesis test asks whether an observed difference could easily happen by chance. The p-value is that chance under the 'no difference' assumption; a small p-value means significance.

What NVIDIA says (1)

“The biggest result is that, across all attempts, both of the lower-latency conditions (25 ms and 55 ms) improved the number of targets eliminated (Figure 3), a difference that was found to be statistically significant in pairwise t-tests ( p-value << 0.001).”

— Improving Player Performance with Low Latency as Evident from FPS Aim Trainer Experiments

Practice 4.4 (3 questions) Full Descriptive Analysis and Visualization guide

← 4.3 Choosing the right plot · 4.5 Patterns, trends and relationships →