12.4 Binomial, Geometric & Normal Distributions, Experiments and Confidence Intervals
Key Takeaways
A binomial setting has a fixed number of independent trials with the same success probability p, so P(X = k) = C(n, k)p^k(1 − p)^(n − k) with mean np.
A geometric model counts trials until the first success, with P(X = k) = (1 − p)^(k − 1)p and an expected value of 1/p.
In a normal distribution, about 68%, 95%, and 99.7% of values lie within 1, 2, and 3 standard deviations of the mean.
Only a randomized experiment with a control group supports cause-and-effect conclusions; observational studies can show association only.
A 95% confidence interval for a proportion is about p̂ ± 1.96√(p̂(1 − p̂)/n), so quadrupling the sample size halves the margin of error.
12.4 Binomial, Geometric and Normal Distributions, Experiments and Confidence Intervals
Competency 013 asks you to use the binomial, geometric and normal distributions to solve problems. Competency 014 asks you to make inferences about a population with those distributions, to understand the relationship between sample size and confidence intervals, and to design, conduct and interpret statistical experiments. This section builds those tools on the counting and probability rules from sections 12.1–12.3.
Random Variables and Probability Distributions
A random variable assigns a number to each outcome of a chance process. A discrete random variable, such as the number of heads in 10 flips, has a list of possible values. A continuous random variable, such as a height, takes any value in an interval. A probability distribution lists each value and its probability (or, for a continuous variable, gives a density curve). The probabilities must add to 1. The mean of the distribution is its expected value, (section 12.3).
The Binomial Distribution
Use the binomial model when all four BINS conditions hold:
- Binary: each trial is a success or a failure.
- Independent: the trials do not affect one another.
- Number: the number of trials is fixed in advance.
- Same: the probability of success is the same on every trial.
Then the probability of exactly successes is
The factor counts the arrangements of the successes among the trials, which is where the combinations from section 12.1 come in.
- Mean: . Standard deviation: .
Worked example. A student guesses on 10 true/false questions, so and .
- Exactly 7 correct: .
- At least 1 correct: . This uses the complement rule.
- Expected number correct: , with standard deviation .
When the binomial does not apply: drawing 5 cards without replacement changes from draw to draw, which breaks the "Same" and "Independent" conditions. When the sample is less than about 10% of the population, the binomial is still a good approximation.
The Geometric Distribution
The geometric model counts the number of trials needed to get the first success. The trials are independent, stays the same, and there is no fixed .
Example. Roll a fair die until the first 6 appears.
- First 6 on the 3rd roll: .
- Expected number of rolls: .
- Probability that it takes more than 4 rolls: . This is the chance that the first four rolls are all non-sixes.
Binomial or geometric? Ask: "Is the number of trials fixed?" If yes, use the binomial and count successes. If you are waiting for the first success, use the geometric.
The Normal Distribution
Many measurements, such as heights, measurement errors and averages of large samples, follow an approximately normal (bell-shaped, symmetric) distribution with mean and standard deviation .
- Empirical rule (68–95–99.7): about 68% of values lie within of the mean, about 95% within , and about 99.7% within .
- z-score: tells how many standard deviations a value lies from the mean. It lets you compare values from different distributions.
Worked example. Test scores are normal with and .
- A score of 700 has . About 95% of scores lie within 2 SD of the mean, so about 2.5% lie above 700 and a 700 is at about the 97.5th percentile.
- About 68% of scores fall between 400 and 600.
- A score of 350 has , so it lies 1.5 standard deviations below the mean.
Normal approximation to the binomial. When and , the binomial distribution is close to normal, with and .
Using Distributions to Draw Conclusions
A coin is flipped 100 times and lands heads 65 times. Is the coin fair? If it were fair, the number of heads would have and . The observed 65 is standard deviations above the mean. A result that extreme happens by chance in only about 0.3% of samples, so the data give strong evidence that the coin favors heads. This logic compares an observed result with what chance alone would produce, and it is the foundation of statistical significance.
Designing Studies: Observational Studies versus Experiments
| Feature | Observational study | Experiment |
|---|---|---|
| What the researcher does | Records existing behavior or characteristics | Imposes a treatment on subjects |
| Can it show cause and effect? | No. Confounding variables may explain the association. | Yes, when well designed with random assignment |
| Example | Survey students about hours of sleep and grades | Randomly assign classes to a new fraction curriculum or the current one and compare results |
Principles of experimental design:
- Comparison (control): include a control group that gets a placebo, no treatment or the current method.
- Random assignment: assign subjects to groups by chance so that lurking variables balance out.
- Replication: use enough subjects so that chance variation does not dominate.
- Blinding: when possible, subjects (and evaluators) do not know who got which treatment.
Remember the difference: random sampling lets you generalize to a population, and random assignment lets you draw cause-and-effect conclusions.
Sampling Distributions and Confidence Intervals
Different random samples give different sample proportions . The distribution of across all possible samples, called the sampling distribution, is centered at the true . Its spread is the standard error:
An approximate 95% confidence interval for is
Example. In a random sample of 400 Texas eighth graders, 60% prefer online homework. Then . The margin of error is , so the 95% interval is about to .
How sample size and confidence level affect the interval:
- Because the margin of error is proportional to , quadrupling the sample size (to 1,600) cuts the margin of error in half (to about 0.024).
- A higher confidence level (99% instead of 95%) makes the interval wider.
- A larger sample makes the interval narrower. The population size barely matters if it is much larger than the sample.
Correct interpretation: "We are 95% confident that the true proportion lies between 55.2% and 64.8%." This means the method captures the true in about 95% of all samples. It does not mean there is a 95% probability that this particular computed interval contains , because the interval either contains it or does not.
Teaching Notes and Common Errors
- Forgetting the combination factor in binomial problems. is the probability of one particular order of successes. Multiply by to count all the orders.
- Using the normal model for small or skewed data. Check the shape, or the and condition, first.
- Confusing association with causation in observational studies.
- Misreading a confidence interval as covering 95% of the individual data values.
- Classroom simulation: Have students flip coins or use random number generators to build a sampling distribution. Seeing the spread shrink as sample size grows makes the rule concrete.
A student randomly guesses on a 5-question quiz where each question has 4 answer choices. What is the probability that the student gets exactly 2 questions correct?
About 0.026
About 0.063
About 0.264
About 0.400
Heights of adult women in a population are approximately normal with mean 162 cm and standard deviation 7 cm. About what percentage of women are between 148 cm and 176 cm tall?
About 95%
About 68%
About 99.7%
About 47.5%
A random sample of 400 students gives a 95% confidence interval for the proportion who prefer online homework with a margin of error of about 4.8 percentage points. About how large a random sample would be needed to cut the margin of error to about 2.4 percentage points at the same confidence level?
800 students, because doubling the sample halves the margin of error
600 students, because the margin of error falls linearly with sample size
4,000 students, because the margin of error falls by a factor of 10 for every tenfold increase
1,600 students, because the margin of error is proportional to 1/√n
Sections you finish are checked off in the contents.