due 0 exam in —
Business Statistics & Analytics — interactive revision
ESSEC META-IDSS2103 · 8 sessions, 13 interactive graphs, spaced-repetition flashcards, timed mock exam with marking schemes
0
cards due
0
mastered
0
day streak
—
to the exam
🚦Start here — what to do right now▼
1 · Diagnose
One quiz per session, then read the weakness dashboard.
2 · Learn
One session: exam sheet → concept cards → interactive graphs.
3 · Retain
Vider la file des cards due chaque jour. Se rappeler bat relire.
4 · Perform
Timed mock exam with the marking scheme, then fix the gaps.
🗺️The whole course in one paragraph▼

Statistics turns a sample into a statement about a population, and the whole course is about the error between the two. Session 2 supplies the machinery: model the data as random variables, and two limit theorems follow — the law of large numbers says the sample mean converges, the central limit theorem says the error is Normal and of size \(1/\sqrt n\). Session 3 insists that no mathematics repairs a badly collected sample, and turns the CLT into a rule for how large \(n\) must be. Sessions 4 and 6 convert that error into confidence intervals and then into hypothesis tests — of a mean, of a difference, of a distribution, of independence in a contingency table. Sessions 5, 7 and 8 apply the same logic to learning: clustering without a response variable, regression to predict a continuous one, and logistic regression to classify a binary one — with the same warning repeated throughout, that association is not causation.

📋Course organisation & resources▼
Progression

S1 Introduction · S2 Random variables and visualisation · S3 Survey design · S4 Confidence intervals · S5 Clustering · S6 Hypothesis testing · S7 Prediction with regression · S8 Statistical models and classification.

Notation convention

This course defines the empirical variance with divisor \(n\): \(S_n^2=\frac1n\sum_i (x_i-\bar x_n)^2\). Statistical software usually divides by \(n-1\) — always state which you are using.

Files in this folder

The eight decks slides N *.pdf, the French summaries resumeN.pdf, and the full BusinessStatistics-lecturenotes.pdf. The extracted slide images appear in each session page.

Offline

This file works with no internet connection. Formulas fall back to plain text if the KaTeX CDN is unreachable; everything else is self-contained.

The quantiles to memorise. \(q_{2.5\%}=-1.96\), \(q_{5\%}=-1.64\), \(q_{95\%}=1.64\), \(q_{97.5\%}=1.96\), and \(q_\alpha=-q_{1-\alpha}\). Two-sided 95% uses 1.96; one-sided 5% uses 1.64. Getting this wrong is the cheapest mark to lose in the whole paper.
🎯 Path to 20/20
A readiness estimate from your own activity, plus the six levers that actually move the grade
⏰ Due today
Everything the spaced-repetition scheduler has queued for you right now
⏰ Review session
All due cards, shuffled across the eight sessions
—
Remaining
0
Known
0
Again
DUE—
Loading…
Click to reveal the answer
⚡ Quick revision — 5 minutes
Read top to bottom. If any line surprises you, that is where to spend the next hour.
📅 Three-week plan
Click a day to tick it off. Progress is stored locally in this browser.
⭐ Starred cards
The cards you flagged while revising
⚙️ Settings
Backup, restore, reset, and a self-test of the whole app
💾Progress backup▼

All progress lives in this browser’s local storage. Export before clearing site data or switching machine.

🧪Self-test▼

Checks that every navigation target resolves, that every graph draws, and reports content counts. Also available from the console as runSmokeTest().

🧮 All formulas
Grouped by session. Cover the right-hand side and reproduce each formula from its name.
🔤 Notation glossary
Every symbol used in the course, with the ambiguities flagged
📈 Concept drills
A situation is described. Name the model, the assumption, the statistic and the conclusion — out loud — before revealing.
🧮 Numeric drills
Compute on paper first. Only mark a drill mastered if you got there without peeking.
🃏 Flashcards (spaced repetition)
The full deck across all eight sessions. Rating a card schedules its next appearance.
—
Remaining
0
Known
0
Again
CARD—
Loading…
Click to reveal the answer
✅ Quiz
Multiple choice with an explanation on every answer. Wrong answers feed the weakness dashboard.
🧪 Exam simulator
Twenty random questions against the clock, with a projected grade band at the end
📝 Mock exam
Five parts · 30 marks · 75 minutes. Write full answers on paper before revealing the model solution and its marking scheme.
How to use this properly

Set a timer for the stated minutes on each part. Do not reveal anything until it ends. Then mark yourself against each rubric line — that is where the examiner actually puts the points, and it is usually the interpretation sentence rather than the algebra.

⚠️ Top exam traps
Twenty-one mistakes that cost marks every year
🎯 Weakness dashboard
Ranked by your own accuracy across cards, quiz and drills — weakest first
🧩 Coverage map
What exists per session and how much of it you have actually mastered
🗺️ The course thread
How the eight sessions chain into a single argument
🔗 Themes ↔ sessions index
Find the session, the diagram and the typical question for any theme
⏱️ Revision 30 / 60 / 90
Pick the time you actually have and follow the plan literally
🕯️ Exam eve
The night before — consolidate, do not accumulate
📏 The 1/√n thread
One rate governs surveys, intervals, tests and the cost of every experiment you will ever run
Where it comes from
$$\sqrt n\,\big(\bar X_n-E[X]\big)\ \xrightarrow[n\to\infty]{}\ \mathcal N\big(0,V[X]\big) \quad\Longleftrightarrow\quad \bar X_n - E[X] = O\!\left(\frac{1}{\sqrt n}\right)$$
S2 · The theorem

The CLT does not only say the error is Normal — it says it shrinks like \(1/\sqrt n\). Everything below is that one line, rewritten.

S3 · Sample size

Interval width \(\le 4\hat\sigma_n/\sqrt n\) and \(\hat\sigma_n\le 1/2\) give \(n\ge 4/w^2\). Halving the width costs four times the sample.

S4 · Standard error

\(S_n/\sqrt n\) is the half-width divided by the quantile. Every interval in the course is centre ± quantile × standard error.

S6 · Power

Test statistics carry a \(\sqrt n\) factor, so power rises with \(\sqrt n\) too: a trivial effect becomes significant once \(n\) is large enough.

The practical consequence

Precision is expensive and gets more expensive. Going from ±3 points to ±1.5 points quadruples the fieldwork budget; going to ±0.75 multiplies it by sixteen.

⚠️ What it never fixes

The rate applies to sampling error only. Selection bias, question bias and confounding are unaffected by \(n\) — a biased frame stays biased at any size.

🧭 The testing recipe
Five steps that answer every test question in the course — means, distributions, contingency tables
1
Define \(H_0\) and \(H_1\)
\(H_0\) is what you try to reject; \(H_1\) is its complement. Exactly one is true. Fix the direction before looking at the data.
2
Find a statistic \(T\) whose law is known under \(H_0\)
This step carries all the assumptions. Mean ⇒ \(T=\sqrt n(\bar X_n-\mu)/S_n\sim\mathcal N(0,1)\) (or \(t_{n-1}\) if Normal). Distribution ⇒ \(\chi^2_{C-1}\). Table ⇒ \(\chi^2_{(I-1)(J-1)}\).
3
Choose a level \(\alpha\)
Usually 5%. It is the probability of a type I error you are willing to accept — a decision, not a fact.
4
Build the rejection region \(R_\alpha\)
Two-sided for a mean: \(|T|>1.96\). One-sided: \(T>1.64\). χ² tests are always upper-tailed.
5
Compute \(T^{obs}\) and conclude
Reject if it lands in \(R_\alpha\); otherwise do not reject. Then write the interpretation in the units of the problem — that sentence is where the marks are.
Degrees of freedom

Goodness of fit \(C-1\) · independence \((I-1)(J-1)\) · Student \(n-1\). One off-by-one loses the whole question.

The two errors

Type I: reject a true \(H_0\), controlled at \(\alpha\). Type II: fail to reject a false one. Power \(=1-P(\text{type II})\), rising with \(n\) and effect size.

⚠️ Not rejecting is not accepting

Failing to reject means the data are compatible with \(H_0\). With small \(n\) the test has almost no power, so this is close to saying nothing.

Test ⇔ interval

Rejecting when \(\mu\) falls outside the \(1-\alpha\) interval is exactly the test at level \(\alpha\). Two languages, one computation.

⚠️ Association vs causation
The warning the course repeats in Sessions 2, 6 and 7 — and the sentence that earns the mark
S2 · Simpson’s paradox

A relationship within each group can reverse when groups are pooled. Aggregation alone can flip a sign, without anything causal changing.

S6 · The χ² warning

The course states it explicitly: one can reject \(H_0\) without any causal link between the two variables. A confounder produces dependence just as well.

S7 · The regression warning

Whatever the predictive performance, regression rests on potential associations. These are not causal links, and additional assumptions are always necessary.

What does license causality

A controlled experiment with random assignment. Everything else — cross-sections, panels, opportunistic data — needs assumptions you must state.

Prediction still works

Diameter predicts tree volume without causing it. Prediction needs a stable association; intervention needs causality. Confusing the two is how models break after a policy change.

🎓 The sentence to write

“This measures the association between \(x\) and \(y\) in this sample. Interpreting it as the effect of an intervention on \(x\) would require assumptions such as random assignment or the absence of omitted confounders.”

Write that sentence unprompted whenever you report a coefficient, a correlation or a rejected independence test. It costs one line and signals that you understood the deepest point of the course.