15.2 Inferential Statistics: Hypothesis Testing, Alpha Levels, Power, and Effect Sizes
Key Takeaways
The Central Limit Theorem dictates that the sampling distribution of the mean approaches normality as N ≥ 30 regardless of the underlying population shape, with a mean equal to the population mean and a standard error of SE = σ / √N.
Null Hypothesis Significance Testing (NHST) evaluates the conditional probability of the observed data given that the null hypothesis is true (p = P(Data | H0)), which must never be confused with the probability that the hypothesis itself is true.
Statistical decisions involve a fundamental 2 × 2 matrix: Type I error (α) is a false positive (rejecting true H0), Type II error (β) is a false negative (failing to reject false H0), and statistical power is defined as 1 - β.
Statistical power (1 - β) is maximized by increasing sample size N, targeting larger true effect sizes, setting a more liberal α, using directional one-tailed tests when justified, and minimizing measurement error.
Because p-values depend directly on sample size and fail to quantify the magnitude of experimental divergence, researchers must report standardized effect sizes such as Cohen's d, r², and η², alongside confidence intervals.
Inferential Statistics: Hypothesis Testing, Alpha Levels, Power, and Effect Sizes
While descriptive statistics summarize observed samples, inferential statistics allow psychologists to draw probabilistic conclusions about broader populations based on sample data. Because researchers rarely have access to entire populations, they must quantify the likelihood that observed sample differences reflect genuine psychological phenomena rather than mere sampling error.
1. Probability, Sampling Distributions, and the Central Limit Theorem
To make valid inferences, researchers rely on theoretical sampling distributions—the probability distribution of a specific statistic (such as the sample mean ) computed from an infinite number of random samples of size drawn from the same population.
The Central Limit Theorem (CLT)
The Central Limit Theorem is the foundational mathematical theorem of inferential statistics. It states that for any population with mean and finite standard deviation , as sample size increases, the sampling distribution of the sample mean approaches a normal distribution, regardless of the shape of the parent population (whether skewed, bimodal, or uniform).
Parent Population (Heavily Skewed) Sampling Distribution of the Mean
/\ (Normal Bell Curve)
/ \ /\
/ \ / \
/ \__________ / \
────────────────────> ────────
Raw Scores (X) Sample Means (X̄)
Approaches Normality as N ≥ 30
Core Properties of the Sampling Distribution of the Mean
- Mean of the Sampling Distribution (Expected Value): The average of all sample means in the sampling distribution is identically equal to the population mean :
This property establishes that the sample mean is an unbiased estimator of .
- Standard Error of the Mean ( or ): The standard deviation of the sampling distribution of the mean is termed the standard error. It quantifies the expected margin of sampling error between a sample mean and the population mean:
The Inverse Square Root Law
Because the denominator of the standard error formula contains , reducing the standard error by half requires quadrupling the sample size ():
- If and , .
- To reduce to , sample size must increase to ().
- To reduce to , sample size must increase to ().
This principle illustrates the law of diminishing returns in empirical sample size planning.
2. The Logic of Null Hypothesis Significance Testing (NHST)
Null Hypothesis Significance Testing (NHST) merges the probabilistic decision frameworks of Ronald Fisher, Jerzy Neyman, and Egon Pearson. In this framework, an empirical claim is tested by evaluating the plausibility of its direct opposite.
Hypotheses Formulation
- Null Hypothesis (): States that in the population, there is no true difference, no experimental effect, or no relationship between variables. Any observed difference in the sample is attributed entirely to random sampling error:
- Alternative Hypothesis ( or ): The experimental or research hypothesis predicting that a genuine difference, effect, or relationship exists in the population:
The Alpha Level () and Critical Regions
Before collecting data, the experimenter specifies the significance level (), which represents the maximum probability of committing a Type I error that the researcher is willing to tolerate. By historical convention in psychology, is typically set at , , or .
- The level demarcates the critical region (rejection region) in the sampling distribution.
- If the calculated test statistic falls within this critical region, the researcher rejects and concludes the effect is statistically significant.
- If the test statistic falls outside the critical region, the researcher fails to reject (never "accepts ", because absence of evidence is not evidence of absence).
Deconstructing the p-Value
The -value is the exact probability of obtaining a test statistic at least as extreme as the one observed in the sample, assuming that the null hypothesis is true:
Warning
The Inverse Fallacy: A widespread misconception among students is believing that the -value represents the probability that the null hypothesis is true, or that represents the probability that the research hypothesis is true. In conditional probability notation, . Calculating requires Bayesian inference incorporating prior probabilities.
3. Decision Matrix, Error Typology, and Statistical Power
Every statistical decision in NHST is subject to uncertainty, yielding four possible outcomes structured across a truth matrix:
| Researcher's Decision | Reality: is Actually True (No Effect) | Reality: is Actually False (Real Effect Exists) |
|---|---|---|
| Reject (Claim effect exists) | Type I Error (False Positive); Probability ; (Incorrectly convicting an innocent person) | Correct Decision (Power / True Positive); Probability ; (Correctly convicting a guilty person) |
| Fail to Reject (Claim no effect detected) | Correct Decision (True Negative); Probability ; (Correctly acquitting an innocent person) | Type II Error (False Negative); Probability ; (Incorrectly acquitting a guilty person) |
The Fundamental Trade-Off Between Type I and Type II Errors
There is an inherent trade-off between and :
- If an experimenter becomes extremely conservative to prevent false positives and lowers from to or , the critical threshold moves further out into the tail of the distribution.
- This shift widens the non-rejection region, making it harder to reject . Consequently, the risk of a false negative (Type II error, ) directly increases, which depresses statistical power ().
Sampling Distribution under H0 Sampling Distribution under H1
│ │
▼ ▼
┌─────┐ ┌─────┐
/ │ \ / │ \
/ │ \ / │ \
/ │ \ Critical Boundary (α) / │ \
/ │ \ │ / │ \
───┴───────┼───────┴──────────────┼───────────────┴───────┼───────┴───
μ0 │ μ1
[ Correct Decision ] │ [ Statistical Power ]
Area = 1 - α │ Area = 1 - β
[ Type I Error α ]
│
[ Type II Error β (under H1 left of boundary) ]
Statistical Power ()
Statistical power () is the probability of correctly rejecting a false null hypothesis—that is, the probability of detecting a true experimental effect when one genuinely exists in the population. Jacob Cohen established that psychological studies should aim for a minimum statistical power of (corresponding to an acceptable Type II error rate of ).
The Determinants of Statistical Power
Statistical power is governed by six interacting factors:
- Sample Size (): Increasing reduces the standard error (), narrowing both sampling distributions and reducing their overlap. This is the primary mechanism researchers use to boost power during study design.
- Magnitude of the Effect Size: Larger true differences between population means produce greater physical separation between the and sampling distributions, increasing the proportion of the curve that clears the critical boundary.
- Significance Level (): Setting a more liberal (e.g., instead of ) shifts the critical boundary toward the center of the distribution, increasing the area under to the right of the cutoff and elevating power.
- Directionality of Test (One-Tailed vs. Two-Tailed): A directional (one-tailed) test places the entire level into a single tail, moving the critical threshold closer to the distribution center and increasing power for effects in that predicted direction.
- Measurement Reliability and Error Variance: Reducing extraneous noise and utilizing highly reliable measurement tools shrinks sample variance (), narrowing the standard error and boosting power.
- Experimental Design (Within-Subjects vs. Between-Subjects): Repeated-measures designs partial out individual baseline differences from the error term, substantially increasing power compared to between-subjects designs with identical sample sizes.
4. Directional (One-Tailed) vs. Non-Directional (Two-Tailed) Tests
When formulating statistical hypotheses, researchers choose between non-directional and directional tests based on their theoretical expectations.
Non-Directional (Two-Tailed) Tests
- Hypotheses: versus .
- Critical Region Allocation: The nominal alpha level is split equally between the two extreme tails of the sampling distribution (e.g., for , in each tail).
- Critical Values: For a standard normal -test with , the critical cutoffs are .
- Methodological Advantage: A two-tailed test detects significant effects regardless of whether the treatment group performs significantly better or significantly worse than the control group. It is the accepted scientific default in psychological research.
Directional (One-Tailed) Tests
- Hypotheses: versus (or vice versa).
- Critical Region Allocation: The entire level is concentrated into the single tail corresponding to the predicted direction (e.g., all in the upper tail).
- Critical Values: For a standard normal -test with , the critical cutoff is .
- Methodological Consideration: Because , a one-tailed test has greater statistical power to detect an effect in the predicted direction. However, if the experimental manipulation produces a massive effect in the opposite direction, the researcher cannot reject the null hypothesis, regardless of how extreme the data are.
5. Effect Size Metrics and Confidence Intervals
A major limitation of Null Hypothesis Significance Testing is that statistical significance () is not equivalent to practical or theoretical importance. Because the standard error shrinks as sample size grows (), an extraordinarily large sample () can render an utterly trivial difference (e.g., a memory enhancement of words) statistically significant at .
To decouple the magnitude of a psychological phenomenon from sample size, researchers report standardized effect sizes and confidence intervals.
Cohen's d (Standardized Mean Difference)
Developed by Jacob Cohen, expresses the difference between two sample means in units of pooled standard deviation:
Unlike the -statistic, Cohen's is independent of sample size. Cohen established widely recognized interpretive benchmarks for psychological research:
- Small Effect (): The difference between means equals standard deviations. The groups overlap by approximately . Cohen's example: the difference in mean height between 15- and 16-year-old girls.
- Medium Effect (): The difference equals half a standard deviation. Overlap is approximately . Noticeable to a trained observer without formal measurement.
- Large Effect (): The difference equals standard deviations. Overlap drops to approximately . Cohen's example: the difference in mean height between 13- and 18-year-old girls.
Variance-Accounted-For Effect Sizes
Other effect size indices quantify the proportion of total variance in the dependent variable that is explained by the independent variable:
- Coefficient of Determination (): In correlational designs, represents the proportion of variance shared between two continuous variables.
- Eta-Squared (): In Analysis of Variance (ANOVA), represents the proportion of total variation attributed to the experimental treatment:
- Partial Eta-Squared (): In factorial ANOVA, isolates the effect of interest by removing variance explained by other factors from the denominator:
Confidence Intervals (CIs)
A confidence interval provides an estimated range of values calculated from sample statistics that is likely to include the unknown population parameter at a specified confidence level (typically or ):
- Frequentist Interpretation: A confidence interval does not mean there is a probability that the true population mean lies within that specific interval. Rather, if an experiment were repeated infinitely using identical methods, of the computed intervals would capture the true population parameter .
- Relation to Hypothesis Testing: Confidence intervals provide a direct test of the null hypothesis. If a confidence interval for the difference between two means contains zero (e.g., ), the difference is not statistically significant at . Conversely, if zero falls outside the interval (e.g., ), the null hypothesis is rejected at .
A psychopharmacology researcher draws a random sample of 25 participants to test a novel anxiolytic compound. The sample yields a standard deviation of 15 on an anxiety inventory. If the researcher wishes to reduce the standard error of the mean by exactly half in a follow-up trial, what total sample size must be recruited?
50 participants
75 participants
100 participants
200 participants
An experimental psychologist testing a cognitive intervention lowers the significance threshold from α = .05 to α = .01 to protect against publishing a false positive. Assuming sample size and true effect size remain constant, what direct statistical consequence does this adjustment produce?
It increases the statistical power (1 - β) of the experiment
It eliminates any possibility of committing a Type II error, because power rises to 1.0
It increases the Type II error rate (β) and therefore lowers statistical power
It increases the probability of committing a Type I error (a false positive result)
A research team evaluates two competing interventions for major depressive disorder across two clinical trials. Study A recruits 10,000 patients and finds a mean difference of 0.5 points on a depression inventory (p < .001). Study B recruits 40 patients and finds a mean difference of 8.0 points on the same inventory (p = .04). Why is reporting Cohen's d necessary when comparing these studies?
Because Cohen's d converts ordinal depression inventory scores into ratio-level measurements
Because a significant p-value is invalid whenever the sample size exceeds 1,000 participants
Because Cohen's d is the only statistical metric that can establish causal relationships in clinical trials
Because p-values are heavily confounded by sample size and do not reflect the magnitude of the experimental effect
A developmental psychologist evaluates the difference in reading comprehension scores between children exposed to phonics instruction versus whole-language instruction. The resulting 95% confidence interval for the difference between the two population means (μ_phonics - μ_whole) is computed as [-1.50, +4.80]. How should this finding be interpreted regarding the null hypothesis of no group difference?
There is a 95% probability that the true difference between group means is exactly equal to zero
The researcher fails to reject the null hypothesis at α = .05 because the 95% confidence interval contains zero
The null hypothesis must be rejected at α = .05 because the interval contains positive values up to +4.80
The experiment is invalid because a confidence interval cannot legitimately span both negative and positive numbers
Sections you finish are checked off in the contents.