7m left·0%
Reading Time: 7 min
Last Updated: August 18, 2026
Main Ideas: 5
Reading Time: 7 min
Last Updated: August 18, 2026
Main Ideas: 5

Topic 1.10 Notes – The Investigative Question Revisited and Data Collection

Verified for 2027 AP® Statistics Exam
Read aloud
An investigative question does more than ask what you’re curious about. In AP Stats, it tells you what data to collect, what analysis fits, and what kind of conclusion the study can actually support. This topic ties the wording of the question to study design, randomness, and the claims you’re allowed to make.

What an Investigative Question Must Specify

A good investigative question locks in three things before data are collected.

  1. Variables
    • You need the variable or variables measured on each observational unit. That just means the individual or item being measured.
    • If it’s a relationship question, identify the explanatory variable and response variable.
    • The variables must be operationalized clearly. “Success” is vague. “Final exam score” or “whether the student passed” is usable.
  2. Parameter or relationship
    • This tells you what analysis makes sense.
    • A hypothesis test asks whether the data support a claim.
    • A confidence interval asks for an estimate of a parameter with a range of plausible values.
  3. Population and intended conclusion
    • The question should name who the conclusion is about.
    • If it asks a causal question, the design must be an experiment with random assignment or that causal wording is not justified.

Wording that signals the analysis

For hypothesis tests, the question must show the parameter and the direction of the alternative:

  • not equal to
  • greater than
  • less than
  • associated
  • not independent

Example

“Do students using planner reminders submit more assignments than students without reminders?”
That points to a one-sided alternative.

For confidence intervals, the question names the parameter being estimated.
Example
“What is the difference in the mean number of assignments submitted...?”

A common mistake is choosing the direction from the sample results. The direction has to come from the question, not from what the data happen to show.

Census, Sample, and What the Data Actually Cover

A population is the full group of interest. A sample is the part you actually collect data from.

A census records information from every member of the population. If some people never respond, it was only an attempted census.

  • “Sent to everyone” does not mean census
  • You need actual data from everyone
  • A census can still have measurement error, missing data, or inaccurate responses

This also trips people up on tests. Census vs. sample is separate from experiment vs. observational study. You can have a sample in either kind of study.

Experiment or Observational Study

The split is simple.

  • Experiment means the researcher imposes treatments
  • Observational study means the researcher only records what naturally happens

Experiments

A randomized experiment usually looks like this.

Study guide illustration

Randomized controlled trial

In an experiment:

  • The experimental unit gets the treatment
  • The explanatory variable is a factor
  • Its categories are levels
  • With one factor, the levels are the treatments
  • With multiple factors, treatments are combinations of levels
  • The response variable is measured after treatment

If students are assigned to daily reminders or no reminders, that’s an experiment.

Observational studies

If students choose for themselves whether to use reminders, that’s observational. You can talk about association, not causation.

Types you should recognize:

  • Survey: standard set of questions given to people
  • Prospective study: select units now, collect data now and in the future
  • Retrospective study: select units now, gather past data

Confounding

A confounding variable must be related to both:

  • the explanatory variable
  • the response variable

It gives another explanation for the association. If students who drink more energy drinks also have heavier workloads, and workload also affects sleep, workload is a confounder.

Random Selection, Random Assignment, and What Conclusions Are Justified

These are different ideas and AP loves to test that.

Study featureWhat it supports
Random selectionGeneralizing to the population
Random assignmentCause-and-effect conclusion

The four combinations:

  • Both random selection and random assignment
    You can generalize to the population and make a causal claim.
  • Random selection only
    You can generalize, but only about association.
  • Random assignment only
    You can make a causal claim for the experimental units or similar individuals, but not broadly generalize.
  • Neither
    No broad generalization and no causation.

A convenience sample or voluntary response sample is not random. A huge sample size does not fix selection bias.

How to Classify a Study Fast

When you read a study description, work through it in this order:

  1. Identify the units and variables
  2. Decide whether it’s a census or sample
  3. Ask whether the researcher imposed treatments
  4. If yes, name the experimental units, factor(s), levels/treatments, and response
  5. If no, call it observational and note survey, prospective, or retrospective if it fits
  6. Check for confounding
  7. Ask two separate questions
    • Were units randomly selected?
    • Were treatments randomly assigned?
  8. Match the conclusion to the design and write it in context

Key Takeaways

The investigative question must specify the variables, the parameter or relationship, and the population plus intended conclusion.
A hypothesis-test question must include the direction of the alternative, and that direction comes from the question, not the sample data.
A census requires data from every member of the population, not just that everyone was contacted.
An experiment has imposed treatments; a group comparison is not automatically an experiment.
Observational studies can show association, but AP Stats does not let you claim causation from them.
A confounder must be associated with both the explanatory and response variables.
Random selection supports generalization to the population, and random assignment supports cause-and-effect.
Large nonrandom samples still have selection bias.
Population language belongs only when random selection supports it.
Causal language belongs only when random assignment in an experiment supports it.

AP® is a trademark registered by the College Board, which is not affiliated with, and does not endorse this website.

Notes

1 credit used · 5/5 remaining