Protocol
Statistical Uncertainty Before Model Claims
Audit a synthetic A/B measurement study with repeated users. Choose the estimand and sampling unit before computing a result, then defend the conclusion under dependence.
45–75 active hours60–90 min sessionsBeginner
Mastery contract
Core protocol · Standard 1.0 · Audit a synthetic A/B measurement study with repeated users. Choose the estimand and sampling unit before computing a result, then defend the conclusion under dependence.
Loading your learning records…
Reasoning — not yet passed
The variance derivation states independence and identifies the missing covariance terms for dependent observations.
Reliable implementation — not yet passed
Bootstrap code resamples the correct units and records seeds, repetitions and the statistic definition.
Reproducible experiment — not yet passed
Report empirical coverage and its Monte Carlo uncertainty; do not require exactly 95% success in a finite simulation.
Defense and handoff — not yet passed
Write an uncertainty analysis citing the saved simulation results, an effect-size calculation and a multiplicity counterexample; distinguish statistical confidence from practical relevance.
All four criteria must pass. Activity completion does not satisfy them.
An independent reviewer reruns your work and varies something you did not rehearse: a fresh input, a different working directory, or a declared edge case.
Pilot decisions are provisional, not external credentials. For an appeal or sensitive artifact, contact your designated pilot operator with the submission ID. Export your assessment history.
What you will do
6 total
Derive the sample mean variance under independent sampling and show why correlated repeats violate that formula; explain frequentist interval coverage.
Implement a bootstrap interval for a mean and a paired difference with explicit resampling units; compare with an analytic interval on Gaussian data.
Run 200 simulated datasets to measure nominal 95% interval coverage; repeat with clustered observations and compare row-level with cluster-level resampling.
Debug a test that treats repeated measurements from one subject as independent and a multiple-testing result selected after inspecting 20 outcomes.
Write a 500–900 word technical report linking the reasoning, code, measurements, and limitations; include the project decision below.
Assess all 4 mastery criteria against saved artifacts, request an independent review, and repeat each failed criterion on a fresh example.
Essentials
A reading session alone never satisfies a mastery criterion.
Store source references, assumptions, and artifact paths beside every result.
Keep an untouched check case that differs from the worked example.
Record all attempts, including failures and results that contradict your prediction.
Use the stated comparison conditions; document every deviation before drawing a conclusion.
Protocol guardrails
Do
- +State the expected result before running the comparison.
- +Keep one minimal reproducible failing case when debugging.
- +Record environment versions and the exact command used.
- +Compare explanations with saved intermediate values.
- +Ask a reviewer to challenge the weakest assumption.
Don't
- ×Do not copy a worked solution and present it as an independent implementation.
- ×Do not tune against held-out evaluation outcomes.
- ×Do not report only the best seed or discard inconvenient runs.
- ×Do not equate elapsed hours or a completed run with a passed assessment.
- ×Do not conceal reduced-scale experiments behind claims about the original full-scale result.
Protocol authorship
Written by NuthinButta as instructional design. The teaching sources are the official references linked in each milestone; the exercises, workload and pass thresholds are ours, not their authors'.