6m left·0%
Reading Time: 6 min
Last Updated: September 15, 2026
Main Ideas: 5
Reading Time: 6 min
Last Updated: September 15, 2026
Main Ideas: 5

Topic 5.3 Notes – Linear Regression Models

Verified for 2027 AP® Statistics Exam
Read aloud
A linear regression model gives you a straight-line equation for predicting one quantitative variable from another. In this topic, the main job is using the equation correctly, reading what its parts mean, and deciding whether a prediction is supported by the data or is a shakier extrapolation.

What a Linear Regression Model Is

You use a linear regression model when a scatterplot of two quantitative variables looks approximately straight-line. The model uses an explanatory variable xx to predict a response variable yy.

y^=a+bx \hat y = a + bx

A few pieces here matter a lot:

  • y^\hat y means the predicted response, not the actual observed response yy.
  • Simple linear regression means one explanatory variable and a straight-line model.
  • The regression line is made of predicted points (x,y^)(x,\hat y).
  • Your actual data points are observed points (x,y)(x,y), and they do not all have to lie on the line.

This kind of graph usually shows all of that at once, including the observed point, the predicted point on the line, and the residual between them.

Study guide illustration

Regression line with predicted value and residual

Two common AP Stats reminders belong here:

  • A model that predicts yy from xx does not automatically prove that xx causes yy.
  • You also cannot just reverse the equation to predict xx from yy. Variable roles matter.

Reading the Equation

The equation tells you two things about the line: its slope and intercept.

Slope

The slope bb is how much the predicted response changes when xx increases by 1 unit.

  • If b>0b>0, predicted yy goes up as xx goes up.
  • If b<0b<0, predicted yy goes down as xx goes up.
  • Units are response units per explanatory-variable unit.

Example. If y^=12+3.5x\hat y = 12 + 3.5x, then for each 1-unit increase in xx, the predicted yy increases by 3.5 units.

Intercept

The intercept aa is the predicted response when x=0x=0.

  • It always has a clear math meaning.
  • It may have no useful real-world meaning.

If x=0x=0 is impossible, unrealistic, or outside the data range, say that. On the exam, don’t force an interpretation that makes no sense.

Calculating and Reporting a Predicted Response

This is the calculation you’ll do most often.

Suppose the model is

height^=8.7+0.63(diameter) \widehat{\text{height}}=8.7+0.63(\text{diameter})

and diameter is in centimeters.

For a tree with diameter 3030 cm,

height^=8.7+0.63(30)=8.7+18.9=27.6 \widehat{\text{height}}=8.7+0.63(30)=8.7+18.9=27.6

So the prediction is 27.6 meters.

Say it like this in context:

  • “According to the regression model, the predicted height of a mature red maple tree with diameter 30 cm is 27.6 meters.”

That wording matters. Don’t say “the height is 27.6 meters.” The model gives a prediction, not an observed fact.

Graphically, this means you go to x=30x=30, move up to the regression line, and read the corresponding y^\hat y.

Interpolation and Extrapolation

A prediction is only as trustworthy as where it falls relative to the observed xx-values.

Interpolation

Interpolation means predicting for an xx-value inside the observed range.

  • If the original diameters were 12 to 46 cm, then x=30x=30 is interpolation.
  • This is usually more defensible because the model is being used where data exist.

Extrapolation

Extrapolation means predicting for an xx-value outside the observed range.

  • If the same data used 12 to 46 cm, then x=70x=70 is extrapolation.
  • It is less reliable because it assumes the linear pattern continues beyond the data.
  • The farther outside the range, the less reliable it usually is.

A predicted value can look reasonable and still be weakly supported. On the AP exam, say that clearly.

When the Model Is Reasonable and Where Students Go Wrong

A linear model makes sense only if the scatterplot looks roughly linear. A strong correlation alone is not enough if the pattern is curved.

Also check whether the prediction makes sense in context:

  • impossible values like negative height
  • values beyond a fixed scale
  • predictions for individuals very different from those in the original data

Big mistakes students make:

  • confusing yy with y^\hat y
  • leaving out units or context
  • treating a prediction as an observed value
  • ignoring interpolation vs extrapolation
  • reversing explanatory and response variables
  • using causal language from regression alone

Key Takeaways

In y^=a+bx\hat y=a+bx, y^\hat y is the predicted response, not the actual observed response.
The slope describes change in the predicted response for each 1-unit increase in xx.
The intercept is the predicted response at x=0x=0, but that may not be meaningful in context.
A complete prediction answer must include context, units, and the word “predicted” or “estimated.”
Interpolation uses an xx-value inside the observed range and is usually more trustworthy than extrapolation.
Extrapolation is less reliable because the model is being used beyond the data that created it.
A regression equation for predicting yy from xx cannot just be reversed to predict xx from yy.
Regression shows association and prediction, not causation by itself.

AP® is a trademark registered by the College Board, which is not affiliated with, and does not endorse this website.

Notes

1 credit used · 5/5 remaining