6m left·0%
Reading Time: 6 min
Last Updated: September 9, 2026
Main Ideas: 4
Reading Time: 6 min
Last Updated: September 9, 2026
Main Ideas: 4

Topic 4.1 Notes – Sampling Distributions for Sample Means

Verified for 2027 AP® Statistics Exam
Read aloud
The whole point of this topic is that one sample mean is not the whole story. If you kept taking random samples of the same size and computed xˉ\bar{x} each time, those sample means would form their own distribution, and that distribution is what AP Stats calls the sampling distribution of xˉ\bar{x}.

What the Sampling Distribution of x̄ Is

For one random sample of size nn, the sample mean is

xˉ=x1+x2+⋯+xnn \bar{x}=\frac{x_1+x_2+\cdots+x_n}{n}

That xˉ\bar{x} is just one sample’s average. If you imagined taking all possible random samples of that same size and averaging each one, the distribution of those averages is the sampling distribution of xˉ\bar{x}.

Three distributions show up here, and mixing them up is one of the most common mistakes:

  • Population distribution
    All individual values in the population. Mean =μ=\mu, standard deviation =σ=\sigma.
  • One sample’s data distribution
    The actual values in the sample you collected. Mean =xˉ=\bar{x}, standard deviation =s=s.
  • Sampling distribution of xˉ\bar{x}
    All possible sample means from repeated samples of size nn. Mean =μxˉ=\mu_{\bar{x}}, standard deviation =σxˉ=\sigma_{\bar{x}}.

The picture below compares the population distribution with sampling distributions of xˉ\bar{x} for two sample sizes. Keep your focus on the repeated-sample distributions centered at μ\mu.

Study guide illustration

Population distribution and sampling distributions of xˉ\bar{x}

This is a long-run theoretical distribution. You usually do not list all possible samples. A simulation can approximate it by repeatedly sampling and recording means.

The key idea is simple. Sample means vary, but averaging several values makes them less variable than individual observations.

Center and Spread of the Sampling Distribution

The center comes first:

μxˉ=μ \mu_{\bar{x}}=\mu

That means the sample mean is an unbiased estimator of the population mean. Over many random samples of size nn, the average of the sample means will equal the true population mean.

Be careful here. This does not mean your one sample mean must equal μ\mu.

The spread is

σxˉ=σn \sigma_{\bar{x}}=\frac{\sigma}{\sqrt{n}}

as long as the sampled values are independent.

What this means in words:

  • σxˉ\sigma_{\bar{x}} is the typical distance of sample means from μ\mu
  • it uses the same units as the original variable
  • for sample-mean problems, use σ/n\sigma/\sqrt{n}, not σ\sigma

Larger samples keep the same center but shrink the spread. In the graph, all three sampling distributions are centered at μ=50\mu=50, but the curves get narrower as nn increases.

  • Multiply nn by 4 →\rightarrow spread is cut in half
  • Multiply nn by 9 →\rightarrow spread is cut by 3
  • Double nn →\rightarrow spread is divided by 2\sqrt{2}

Sampling distributions of xˉ\bar{x} for different sample sizes

A bigger sample improves precision. It does not fix bad sampling.

Conditions and Shape

Two separate questions matter here.

Independence

These conditions justify the formulas for center and spread:

  • Randomization means the data come from a random sample or another valid random process.
  • 10% condition matters when sampling without replacement from a finite population. Check n≤0.10Nn \le 0.10N.

If sampling is with replacement, you do not need the 10% condition.

Shape

These conditions justify using normal probability calculations:

  • If the population is normal, then xˉ\bar{x} is normal for any sample size.
  • If the population is not normal, the Central Limit Theorem says xˉ\bar{x} is approximately normal for large enough nn.
  • In AP Stats, n≥30n \ge 30 is the usual rule of thumb.
  • If the population is extremely skewed or has extreme outliers, you may need a much larger sample.

The CLT changes the shape of the sampling distribution of means, not the population itself.

Describing and Using the Sampling Distribution

A complete description gives shape, center, and spread in context.

A standard AP-style response usually sounds like this:

  1. Define the variable as xˉ\bar{x}, the sample mean.
  2. Check randomization and the 10% condition if sampling without replacement.
  3. Justify the shape using population normality or the CLT.
  4. State μxˉ=μ\mu_{\bar{x}}=\mu.
  5. State σxˉ=σ/n\sigma_{\bar{x}}=\sigma/\sqrt{n}.

For probabilities, once the sampling distribution is normal or approximately normal, use

z=xˉ−μσ/n z=\frac{\bar{x}-\mu}{\sigma/\sqrt{n}}

Example. Suppose household electricity use has μ=31\mu=31, σ=18\sigma=18, and n=100n=100.

μxˉ=31,σxˉ=18100=1.8 \mu_{\bar{x}}=31,\qquad \sigma_{\bar{x}}=\frac{18}{\sqrt{100}}=1.8

If you want P(xˉ>34)P(\bar{x}>34),

z=34−311.8=1.67 z=\frac{34-31}{1.8}=1.67

So P(xˉ>34)≈0.048P(\bar{x}>34)\approx 0.048.

In context, about 4.8% of random samples of 100 households would have a mean electricity use above 34 kilowatt-hours.

That probability is about sample means, not individual households.

Key Takeaways

The sampling distribution of xˉ\bar{x} is a distribution of sample means, not raw data values.
For sample means, the center is μxˉ=μ\mu_{\bar{x}}=\mu, but one particular xˉ\bar{x} does not have to equal μ\mu.
The spread for sample means is σxˉ=σ/n\sigma_{\bar{x}}=\sigma/\sqrt{n}, and using σ\sigma instead is a classic mistake.
The 10% condition checks approximate independence for sampling without replacement, and it has nothing to do with normality.
If the population is normal, xˉ\bar{x} is normal for any nn.
The rule n≥30n \ge 30 is a guideline for the CLT, not a requirement in every problem.
Probability statements must name xˉ\bar{x} when the question is about a sample mean.

AP® is a trademark registered by the College Board, which is not affiliated with, and does not endorse this website.

Notes

1 credit used · 5/5 remaining