Probability & Statistics · Central limit theorem

Central limit theorem

Average n independent draws from a population with a finite, positive variance: magnified n\sqrt{n} times about the mean, the distribution of the average approaches a normal curve as n grows, whatever the population’s shape.

01

The central limit theorem in this visualization

Let X1,…,XnX_1, \dots, X_n be independent draws from one population (independent and identically distributed) with mean μ\mu and variance 0<σ2<∞0 < \sigma^2 < \infty, and let Xˉ\bar X be their average. The central limit theorem, in its Lindeberg–Lévy form, says that the standardized average converges in distribution to the standard normal: its distribution function tends to Φ\Phi at every point. Informally, Xˉ≈N(μ,σ2/n)\bar X \approx N(\mu, \sigma^2/n). The page shows how the shape of the distribution of Xˉ\bar X changes with n.

The theorem
P ⁣(n (Xˉ−μ)σ≤z)→Φ(z)for every zP\!\left(\frac{\sqrt{n}\,(\bar X - \mu)}{\sigma} \le z\right) \to \Phi(z) \quad \text{for every } z

The upper layer is the population. The lower layer shows the distribution of the average Xˉ\bar X of nn draws. Repeating this NN times gives the yellow histogram: it fluctuates with sampling and can be compared with the theoretical curve. A fixed seed reproduces the same samples. Magnified by √n, the lower axis is Y=μ+n (xˉ−μ)Y = \mu + \sqrt{n}\,(\bar x - \mu): its mean μ and standard deviation σ match the population, making the shapes easier to compare.

Curves marked “numeric” are computed numerically; the others use exact distribution formulas. Skewness measures asymmetry. Excess kurtosis compares the standardized fourth moment with that of a normal distribution, for which it is zero. When the corresponding population moments exist, the average has skewness γ/n\gamma/\sqrt{n} and excess kurtosis κ/n\kappa/n. Neither measure alone guarantees a close normal approximation.

02

Things to notice

  • The theorem is about the distribution of the average, not about the data: the upper layer never changes. Its narrowing like σ/n\sigma/\sqrt{n} at actual width is the √n law, a separate fact.
  • “n ≥ 30 is enough” is a rule of thumb. For an event of probability 1%, the average of 30 trials is still exactly 0 with probability 0.9930≈0.7400.99^{30} \approx 0.740, and its skewness γ/n\gamma/\sqrt{n} drops to 0.2 only at n = 2426. Skewness 0 is not enough either: the symmetric version (±10, each with probability 0.5%) has skewness 0 for every n and still a large atom at 0.
  • Bars of a lattice law are one lattice step wide, so their area is a probability: compare shapes, not heights point by point. The Cauchy distribution has no mean and no variance: the average of n draws has exactly the distribution of one draw, and the theorem does not apply.
  • The normal limit goes back to de Moivre (1733, for the binomial) and Laplace (1812); the name is Pólya’s (1920).
03

Controls

  • Draw turns the upper layer into a drawing board: what you draw is scaled to area 1 when you let go. The red key replays 10 000 averages computed in advance; pressed while running it stops, pressed again it goes on with the same seed, and at the end it starts the next seed. Share stores the seed, the algorithm’s version and how far the replay went.
04

Related

  • H05Normal distributionPlanned
  • H18Sampling distributionPlanned
  • H12Law of large numbers
  • H16Binomial distributionPlanned
  • H20Confidence intervalPlanned
  • H26Probability density functionPlanned

Further reading: Wikipedia: Central limit theorem; OpenStax, Introductory Statistics 2e, 7 The Central Limit Theorem.