7m left·0%
Reading Time: 7 min
Last Updated: September 14, 2026
Main Ideas: 5
Reading Time: 7 min
Last Updated: September 14, 2026
Main Ideas: 5

Topic 4.10 Notes – Carrying Out a Test for the Difference Between Two Population Means

Verified for 2027 AP® Statistics Exam
Read aloud
A two-sample t-test for a difference in means checks whether two independent groups seem to come from populations with different average values on a quantitative variable. In this topic, you’re carrying out the test, checking when it’s valid, computing the test statistic and p-value, and writing the conclusion in correct AP Stats language.

What a Two-Sample t-Test for a Difference in Means Does

This test compares two independent groups. The groups might be two populations from random samples, or two treatments from a randomized experiment.

  • The variable must be quantitative.
  • The parameter you care about is the difference in population means, μ1−μ2 \mu_1 - \mu_2 .
  • The sample statistic is the difference in sample means, xˉ1−xˉ2 \bar{x}_1 - \bar{x}_2 .

A usual null hypothesis says there is no difference:

  • H0:μ1−μ2=0H_0:\mu_1-\mu_2=0
  • same idea as H0:μ1=μ2H_0:\mu_1=\mu_2

The alternative depends on the question:

  • Ha:μ1−μ2≠0H_a:\mu_1-\mu_2\ne 0
  • Ha:μ1−μ2>0H_a:\mu_1-\mu_2>0
  • Ha:μ1−μ2<0H_a:\mu_1-\mu_2<0

Because the population standard deviations are unknown, this uses sample standard deviations s1s_1 and s2s_2, so it is a t-test.

One fast warning that shows up all the time on tests: this is only for independent groups. If the data are paired, like before/after or twins, you turn each pair into a difference and do a one-sample t-test on those differences.

Also, the subtraction order must stay fixed. If you define μ1−μ2 \mu_1-\mu_2, every statistic and conclusion must match that order.

When the Test Is Appropriate

Both groups have to meet the conditions. One solid group does not rescue a bad one.

Randomization

You need either:

  • two independent random samples, or
  • a randomized experiment with random assignment to treatments

On the AP exam, saying “the samples are independent” is too vague by itself. Say how the data were collected.

Independence

For random samples taken without replacement, check the 10% condition for both groups:

  • n1≤0.10N1n_1 \le 0.10N_1
  • n2≤0.10N2n_2 \le 0.10N_2

That condition is not needed in a randomized experiment.

Sample data condition

The t-procedure is okay if:

  • both populations are approximately normal, or
  • both sample sizes are at least 30, or
  • with small samples, both groups show no strong skewness and no outliers on separate graphs

That is why side-by-side boxplots are so helpful here. You are checking each group for overall shape and any clear outliers before using the test.

Study guide illustration

Boxplots showing skewness and outliers

Carrying Out the Test

A full test has a standard flow.

  1. Define the parameters in context
    Example: μ1 \mu_1 = the mean battery life for all phones using Brand A, and μ2 \mu_2 = the mean battery life for all phones using Brand B.

  2. Write hypotheses using population means, never sample means.

  3. Name the procedure
    Two-sample t-test for a difference between population means.

  4. Compute the standard error

    SE=s12n1+s22n2 SE=\sqrt{\frac{s_1^2}{n_1}+\frac{s_2^2}{n_2}}

    The variances add because the samples are independent. This is unpooled, so on a calculator use pooled = No.

  5. Compute the test statistic

    t=(xˉ1−xˉ2)−0s12n1+s22n2 t=\frac{(\bar{x}_1-\bar{x}_2)-0}{\sqrt{\frac{s_1^2}{n_1}+\frac{s_2^2}{n_2}}}

    Example with xˉ1=18.6,s1=4.2,n1=24 \bar{x}_1=18.6, s_1=4.2, n_1=24 and xˉ2=15.1,s2=3.6,n2=21 \bar{x}_2=15.1, s_2=3.6, n_2=21:

    SE=4.2224+3.6221≈1.163 SE=\sqrt{\frac{4.2^2}{24}+\frac{3.6^2}{21}}\approx 1.163

    t=18.6−15.11.163≈3.01 t=\frac{18.6-15.1}{1.163}\approx 3.01

    A positive tt means group 1’s sample mean is above group 2’s.

  6. Get df from technology
    In AP Stats, df is usually not n1+n2−2n_1+n_2-2.

  7. Find the p-value from the correct tail(s).

    This depends on your alternative hypothesis. A one-sided test puts all the area in one tail. A two-sided test splits the area between both tails, like the comparison shown here.

Study guide illustration

One-tailed vs. two-tailed t-tests

Interpreting the p-Value and Making the Decision

The p-value is the probability of getting a sample difference at least as extreme as the one observed, in the direction of HaH_a, assuming H0H_0 is true.

In context, that means assuming the two population means are equal.

If p-value≤αp\text{-value} \le \alpha, reject H0H_0.
If p-value>αp\text{-value} > \alpha, fail to reject H0H_0.

Your conclusion must:

  • mention population means
  • match the direction of HaH_a
  • use AP wording like “convincing statistical evidence”

Example conclusion:
“There is convincing statistical evidence that the population mean battery life for phones using Brand A differs from the population mean battery life for phones using Brand B.”

Random samples let you generalize to populations. Random assignment lets you talk about cause and effect. Statistical significance does not automatically mean the difference is important in real life.

Common Mistakes to Catch Fast

  • Using this test for matched pairs data
  • Switching from μ1−μ2 \mu_1-\mu_2 to xˉ2−xˉ1 \bar{x}_2-\bar{x}_1
  • Writing hypotheses with xˉ1 \bar{x}_1 and xˉ2 \bar{x}_2 instead of μ1 \mu_1 and μ2 \mu_2
  • Using pooled procedures or automatically using df=n1+n2−2df=n_1+n_2-2
  • Using the wrong standard error formula
  • Picking the wrong tail for the p-value
  • Saying “accept H0H_0”
  • Saying a nonsignificant result proves the means are equal
  • Writing a conclusion about the samples instead of the population means

Key Takeaways

This test is for two independent groups with one quantitative variable.
Matched pairs data belong in a one-sample t-test on the sample of differences.
Keep the subtraction order consistent from μ1−μ2 \mu_1-\mu_2 to xˉ1−xˉ2 \bar{x}_1-\bar{x}_2 to your conclusion.
The standard error is s12n1+s22n2 \sqrt{\frac{s_1^2}{n_1}+\frac{s_2^2}{n_2}} , not s1+s2n \frac{s_1+s_2}{\sqrt{n}} and not a pooled formula.
On AP Stats, use the unpooled procedure, so calculator setting is usually pooled = No.
The p-value is computed assuming the population means are equal in context.
A small p-value gives evidence against H0H_0, not proof that the difference is practically important.
“Fail to reject H0H_0” means the data did not give convincing evidence of a difference, not that the means are equal.

AP® is a trademark registered by the College Board, which is not affiliated with, and does not endorse this website.

Notes

1 credit used · 5/5 remaining