7.5 Statistical Questions, Sampling, and Inference
Key Takeaways
- A statistical question anticipates variability and is answered with a distribution of data, unlike a question with a single deterministic answer.
- Inference uses a sample statistic (such as a sample proportion) to estimate a population parameter, with the understanding that samples vary.
- Simple random sampling gives each population member an equal chance of selection; convenience and voluntary response methods commonly introduce bias.
- Large sample size does not fix selection bias; survey wording and nonresponse can still distort results even when selection is random.
Statistical Questions, Sampling, and Inference
ETS Category V.A asks whether candidates can distinguish statistical questions from non-statistical ones, recognize how samples support inferences about populations, and identify sampling methods that are random versus biased. These ideas underpin middle school statistics and appear frequently in Praxis teaching scenarios about survey design.
What Makes a Question Statistical?
A statistical question anticipates variability in the data and is answered by collecting and analyzing a distribution of values—not by a single deterministic fact.
| Question | Statistical? | Why |
|---|---|---|
| How tall is the school's flagpole? | No | One measurable height; no anticipated variability in answers |
| How tall are the students in grade 7? | Yes | Heights vary; a distribution is expected |
| What is 15% of 80? | No | A single calculated value |
| What percent of grade 7 students eat breakfast daily? | Yes | Individual habits vary; a sample proportion estimates the population |
Key test: if every correct data collection would produce essentially the same single answer, the question is not statistical. If answers will differ across individuals or trials, it is statistical.
Populations, Samples, and Parameters
- A population is the entire group of interest (all middle school students in a district).
- A sample is a subset from which you collect data.
- A parameter is a numerical summary of the population (true mean height, true proportion who own a phone).
- A statistic is the corresponding summary computed from the sample (sample mean, sample proportion).
Inference means using a sample statistic to estimate or draw a conclusion about a population parameter. Middle school Praxis items typically involve informal inference: if about 40% of a random sample of students prefer soccer, we estimate that about 40% of the population prefers soccer—with the understanding that different samples will vary.
Worked example: estimating a population count. A school has 900 students. In a random sample of 50 students, 18 walk to school. Estimate how many students in the school walk to school.
Sample proportion: 18/50 = 0.36.
Estimated population count: 0.36 × 900 = 324 students.
This is an estimate, not a guarantee; another random sample might yield a slightly different proportion.
Random Sampling Versus Biased Sampling
A simple random sample gives every member of the population an equal chance of selection (and typically every sample of a given size an equal chance). Random sampling supports unbiased estimation in the long run.
Common problematic methods:
- Convenience sampling: survey whoever is easiest to reach (students in the cafeteria at 11:00). Easy, but may miss subgroups.
- Voluntary response sampling: post a survey and let people choose to respond. Those with strong opinions often respond more, skewing results.
- Systematic but non-random shortcuts: surveying only one advisory class, only athletes, or only students who stay after school.
Bias is a systematic tendency for the sample to differ from the population in a particular direction.
| Bias type | Example |
|---|---|
| Selection / undercoverage | Surveying only bus riders about transportation preferences |
| Response bias | Wording: "Don't you agree that homework is excessive?" |
| Nonresponse bias | Many selected students ignore the survey; responders differ from nonresponders |
| Voluntary response bias | Online poll about phone bans filled mostly by students who feel strongly |
Random selection does not eliminate all bias (poor wording can still distort answers), but it is the primary defense against selection bias.
Classroom Teaching Scenario: Evaluating a Student's Survey Design
A student wants to know: "How many hours per week do students at our school spend on homework?" The student designs this plan:
- Stand at the entrance to the gym after school and ask the first 30 students who walk by.
- Ask: "You don't spend more than 2 hours a night on homework, right?"
- Report the average of those 30 responses as the exact average for the whole school.
Evaluate the design.
- The question itself is statistical—hours vary—so the topic is appropriate.
- The sample is a convenience sample of students near the gym after school (likely athletes or club members), so it is not representative of all students.
- The wording is leading (response bias), pushing respondents toward "no more than 2 hours."
- Reporting the sample mean as the exact population mean overstates certainty; it should be framed as an estimate.
A stronger design would select a simple random sample from the student roster (or stratified samples by grade), use neutral wording ("About how many hours per week do you spend on homework?"), and describe the result as an estimate of the schoolwide average.
Making Reasonable Inferences
Praxis items often ask which conclusion is justified.
- From a random sample of 40 out of 500 club members, if 10 prefer online meetings, estimating that about 25% of club members prefer online meetings is reasonable.
- Concluding that exactly 125 members prefer online meetings overstates precision.
- Extending the result to all teenagers nationwide from one school's sample is unjustified (wrong population).
- A convenience sample of friends cannot support a schoolwide claim.
Praxis Traps
- Calling any question with a number in it "statistical" (calculations are not statistical questions).
- Assuming a large convenience sample is automatically representative—size does not cure selection bias.
- Confusing a sample statistic with a population parameter, or treating an estimate as exact.
- Ignoring question wording that introduces response bias even when selection was random.
- Generalizing beyond the population from which the sample was drawn.
For Praxis 5164, practice classifying questions, naming sampling methods and bias types, computing proportion-based population estimates, and critiquing student survey plans with precise statistical language.
Which of the following is a statistical question?
A school has 1,200 students. In a random sample of 60 students, 21 bring lunch from home. Which is the best estimate of how many students in the school bring lunch from home?
A student posts an optional online survey about school lunch quality and uses only the responses from students who choose to reply. Which sampling problem is most clearly present?
A teacher wants every student in a school of 800 an equal chance of being chosen for a 40-student sample. Which method best achieves that goal?
You've completed this section
Continue exploring other exams