5m left·0%
Reading Time: 5 min
Last Updated: September 11, 2026
Main Ideas: 5
Reading Time: 5 min
Last Updated: September 11, 2026
Main Ideas: 5

Topic 4.6 Notes – Sampling Distributions for the Difference Between Two Sample Means

Verified for 2027 AP® Statistics Exam
Read aloud
This topic is about the sampling distribution of the difference between two sample means, written as xˉ1−xˉ2\bar{x}_1-\bar{x}_2. You use it when comparing two independent groups on one quantitative variable and want to understand the center, spread, shape, and probability behavior of that sample-mean difference over many repeated samples.

What the Sampling Distribution of xˉ1−xˉ2\bar{x}_1 - \bar{x}_2 Is

You have two independent groups or populations, one quantitative variable, and two sample means, xˉ1\bar{x}_1 and xˉ2\bar{x}_2. The statistic is their difference:

xˉ1−xˉ2 \bar{x}_1-\bar{x}_2

That statistic estimates the population difference:

μ1−μ2 \mu_1-\mu_2

The subtraction order matters the whole time. If population 1 is “students using Method A” and population 2 is “students using Method B,” then keep that order for both sample means and population means.

A sampling distribution is the distribution of xˉ1−xˉ2\bar{x}_1-\bar{x}_2 from many repeated independent samples. It is about sample means, not individual data values.

Study guide illustration

Two independent samples from two populations

The diagram shows the setup. You draw one independent sample from population 1 and another from population 2, then compare their sample means.

The center is

μxˉ1−xˉ2=μ1−μ2 \mu_{\bar{x}_1-\bar{x}_2}=\mu_1-\mu_2

So xˉ1−xˉ2\bar{x}_1-\bar{x}_2 is an unbiased estimator of μ1−μ2\mu_1-\mu_2. If μ1=μ2\mu_1=\mu_2, the sampling distribution is centered at 0, even though actual sample differences will bounce above and below 0.

Spread and Shape of the Sampling Distribution

The standard deviation of the sampling distribution is

σxˉ1−xˉ2=σ12n1+σ22n2 \sigma_{\bar{x}_1-\bar{x}_2}=\sqrt{\frac{\sigma_1^2}{n_1}+\frac{\sigma_2^2}{n_2}}

Here’s the part students mix up a lot. Variances add for independent random variables, even when you subtract the variables. That’s why the formula adds σ12n1\frac{\sigma_1^2}{n_1} and σ22n2\frac{\sigma_2^2}{n_2}.

  • You do not add or subtract standard deviations directly.
  • The spread is in the same units as the original variable.
  • Bigger n1n_1 or n2n_2 makes the sampling distribution tighter.
  • The population with larger σ2/n\sigma^2/n contributes more to the total spread.

For the shape, xˉ1−xˉ2\bar{x}_1-\bar{x}_2 is:

  • Exactly normal if both population distributions are normal
  • Approximately normal if both sample sizes are large, meaning n1≥30n_1 \ge 30 and n2≥30n_2 \ge 30

Both samples have to meet the large-sample condition. One large sample does not rescue one small sample.

Full model:

xˉ1−xˉ2∼N(μ1−μ2,σ12n1+σ22n2) \bar{x}_1-\bar{x}_2 \sim N\left(\mu_1-\mu_2,\sqrt{\frac{\sigma_1^2}{n_1}+\frac{\sigma_2^2}{n_2}}\right)

Conditions to Check Before Using the Model

You need to justify independence and normal shape.

Independence conditions

  • For sampling, you need two independent random samples
  • For experiments, you need random assignment to independent groups
  • This is not for matched pairs or repeated measures. Those use paired-data methods.

The 10% condition

This only matters for sampling without replacement from populations.

Check each sample separately:

  • n1≤0.10N1n_1 \le 0.10N_1
  • n2≤0.10N2n_2 \le 0.10N_2

You compare each sample to its own population, not a combined sample to a combined population. Randomized experiments do not need the 10% condition.

What to say on an exam

Name the randomization source from the prompt, say the two samples or groups are independent, check 10% separately if sampling, and justify normality with either normal populations or both sample sizes at least 30.

Finding and Interpreting Probabilities

Only do probability calculations after the normal model is justified.

Standardize with

z=(xˉ1−xˉ2)−(μ1−μ2)σ12n1+σ22n2 z=\frac{(\bar{x}_1-\bar{x}_2)-(\mu_1-\mu_2)}{\sqrt{\frac{\sigma_1^2}{n_1}+\frac{\sigma_2^2}{n_2}}}

For a cutoff cc,

z=c−(μ1−μ2)σ12n1+σ22n2 z=\frac{c-(\mu_1-\mu_2)}{\sqrt{\frac{\sigma_1^2}{n_1}+\frac{\sigma_2^2}{n_2}}}

Example. Suppose μ1=78\mu_1=78, μ2=74\mu_2=74, σ1=12\sigma_1=12, σ2=10\sigma_2=10, n1=64n_1=64, n2=100n_2=100.

Then

μxˉ1−xˉ2=78−74=4 \mu_{\bar{x}_1-\bar{x}_2}=78-74=4

σxˉ1−xˉ2=12264+102100=2.25+1=3.25≈1.803 \sigma_{\bar{x}_1-\bar{x}_2}=\sqrt{\frac{12^2}{64}+\frac{10^2}{100}}=\sqrt{2.25+1}=\sqrt{3.25}\approx1.803

For P(xˉ1−xˉ2>7)P(\bar{x}_1-\bar{x}_2>7),

z=7−41.803≈1.66 z=\frac{7-4}{1.803}\approx1.66

So P(Z>1.66)≈0.048P(Z>1.66)\approx0.048.

In context, you’d say there is about a 0.048 probability that repeated independent samples of 64 and 100 students would produce a sample mean score difference greater than 7 points.

What Students Mix Up

  • Using this for paired data instead of independent groups
  • Flipping the subtraction order halfway through
  • Describing the distribution as if it were about individual observations
  • Forgetting both samples need normal populations or both need n≥30n \ge 30
  • Using the 10% condition for experiments
  • Using s1s_1 and s2s_2 here. This topic uses population SDs, σ1\sigma_1 and σ2\sigma_2
  • Treating probability as about μ1−μ2\mu_1-\mu_2. The parameter is fixed and the statistic varies
  • Thinking large samples fix bad sampling or confounding

Key Takeaways

Keep the subtraction order consistent in xˉ1−xˉ2\bar{x}_1-\bar{x}_2 and μ1−μ2\mu_1-\mu_2 or your sign will be wrong.
The mean of the sampling distribution is μ1−μ2\mu_1-\mu_2, so xˉ1−xˉ2\bar{x}_1-\bar{x}_2 is an unbiased estimator.
The spread uses σxˉ1−xˉ2=σ12/n1+σ22/n2\sigma_{\bar{x}_1-\bar{x}_2}=\sqrt{\sigma_1^2/n_1+\sigma_2^2/n_2}, and variances add even when means are subtracted.
Both samples must satisfy the normal condition route you use, either both populations normal or both sample sizes at least 30.
The 10% condition is checked separately for each population and is only for sampling without replacement, not experiments.
Probability statements here are about what repeated samples might produce, not about the chance that μ1−μ2\mu_1-\mu_2 has some value.

AP® is a trademark registered by the College Board, which is not affiliated with, and does not endorse this website.

Notes

1 credit used · 5/5 remaining