One quiz per session, then read the weakness dashboard.
One session: exam sheet → concept cards → interactive graphs.
Vider la file des cards due chaque jour. Se rappeler bat relire.
Timed mock exam with the marking scheme, then fix the gaps.
Statistics turns a sample into a statement about a population, and the whole course is about the error between the two. Session 2 supplies the machinery: model the data as random variables, and two limit theorems follow — the law of large numbers says the sample mean converges, the central limit theorem says the error is Normal and of size \(1/\sqrt n\). Session 3 insists that no mathematics repairs a badly collected sample, and turns the CLT into a rule for how large \(n\) must be. Sessions 4 and 6 convert that error into confidence intervals and then into hypothesis tests — of a mean, of a difference, of a distribution, of independence in a contingency table. Sessions 5, 7 and 8 apply the same logic to learning: clustering without a response variable, regression to predict a continuous one, and logistic regression to classify a binary one — with the same warning repeated throughout, that association is not causation.
S1 Introduction · S2 Random variables and visualisation · S3 Survey design · S4 Confidence intervals · S5 Clustering · S6 Hypothesis testing · S7 Prediction with regression · S8 Statistical models and classification.
This course defines the empirical variance with divisor \(n\): \(S_n^2=\frac1n\sum_i (x_i-\bar x_n)^2\). Statistical software usually divides by \(n-1\) — always state which you are using.
The eight decks slides N *.pdf, the French summaries resumeN.pdf, and the full BusinessStatistics-lecturenotes.pdf. The extracted slide images appear in each session page.
This file works with no internet connection. Formulas fall back to plain text if the KaTeX CDN is unreachable; everything else is self-contained.
All progress lives in this browser’s local storage. Export before clearing site data or switching machine.
Checks that every navigation target resolves, that every graph draws, and reports content counts. Also available from the console as runSmokeTest().
Set a timer for the stated minutes on each part. Do not reveal anything until it ends. Then mark yourself against each rubric line — that is where the examiner actually puts the points, and it is usually the interpretation sentence rather than the algebra.
The CLT does not only say the error is Normal — it says it shrinks like \(1/\sqrt n\). Everything below is that one line, rewritten.
Interval width \(\le 4\hat\sigma_n/\sqrt n\) and \(\hat\sigma_n\le 1/2\) give \(n\ge 4/w^2\). Halving the width costs four times the sample.
\(S_n/\sqrt n\) is the half-width divided by the quantile. Every interval in the course is centre ± quantile × standard error.
Test statistics carry a \(\sqrt n\) factor, so power rises with \(\sqrt n\) too: a trivial effect becomes significant once \(n\) is large enough.
Precision is expensive and gets more expensive. Going from ±3 points to ±1.5 points quadruples the fieldwork budget; going to ±0.75 multiplies it by sixteen.
The rate applies to sampling error only. Selection bias, question bias and confounding are unaffected by \(n\) — a biased frame stays biased at any size.
\(H_0\) is what you try to reject; \(H_1\) is its complement. Exactly one is true. Fix the direction before looking at the data.
This step carries all the assumptions. Mean ⇒ \(T=\sqrt n(\bar X_n-\mu)/S_n\sim\mathcal N(0,1)\) (or \(t_{n-1}\) if Normal). Distribution ⇒ \(\chi^2_{C-1}\). Table ⇒ \(\chi^2_{(I-1)(J-1)}\).
Usually 5%. It is the probability of a type I error you are willing to accept — a decision, not a fact.
Two-sided for a mean: \(|T|>1.96\). One-sided: \(T>1.64\). χ² tests are always upper-tailed.
Reject if it lands in \(R_\alpha\); otherwise do not reject. Then write the interpretation in the units of the problem — that sentence is where the marks are.
Goodness of fit \(C-1\) · independence \((I-1)(J-1)\) · Student \(n-1\). One off-by-one loses the whole question.
Type I: reject a true \(H_0\), controlled at \(\alpha\). Type II: fail to reject a false one. Power \(=1-P(\text{type II})\), rising with \(n\) and effect size.
Failing to reject means the data are compatible with \(H_0\). With small \(n\) the test has almost no power, so this is close to saying nothing.
Rejecting when \(\mu\) falls outside the \(1-\alpha\) interval is exactly the test at level \(\alpha\). Two languages, one computation.
A relationship within each group can reverse when groups are pooled. Aggregation alone can flip a sign, without anything causal changing.
The course states it explicitly: one can reject \(H_0\) without any causal link between the two variables. A confounder produces dependence just as well.
Whatever the predictive performance, regression rests on potential associations. These are not causal links, and additional assumptions are always necessary.
A controlled experiment with random assignment. Everything else — cross-sections, panels, opportunistic data — needs assumptions you must state.
Diameter predicts tree volume without causing it. Prediction needs a stable association; intervention needs causality. Confusing the two is how models break after a policy change.
“This measures the association between \(x\) and \(y\) in this sample. Interpreting it as the effect of an intervention on \(x\) would require assumptions such as random assignment or the absence of omitted confounders.”