15.3 Bivariate and Multivariate Techniques: t-Tests, ANOVA Models, Correlation, and Regression
Key Takeaways
Parametric tests require continuous interval/ratio data, independent observations, bivariate/multivariate normality, and homogeneity of variance across groups (evaluated via Levene's test).
Student's t-test evaluates mean differences across one-sample (df = N - 1), independent-samples (df = n1 + n2 - 2), and paired-samples (df = Npairs - 1) designs, with paired designs achieving greater power by partialing out inter-subject variance.
Analysis of Variance (ANOVA) resolves familywise error rate inflation (α_FW = 1 - (1 - α)^c) via an omnibus F-ratio (MS_between / MS_within); factorial ANOVA enables simultaneous testing of main effects and interaction dynamics (A × B).
Pearson r measures linear bivariate association and Ordinary Least Squares (OLS) regression calculates optimal predictive trajectories (Ŷ = bX + a), while non-parametric tests (Chi-Square, Mann-Whitney U, Wilcoxon, Kruskal-Wallis) analyze categorical or rank-ordered distributions free of distributional assumptions.
Bivariate and Multivariate Techniques: t-Tests, ANOVA Models, Correlation, and Regression
Psychological investigations evaluate hypotheses spanning diverse designs, from comparing treatment groups to modeling complex multivariate associations. Selecting the appropriate statistical test requires assessing the scale of measurement, the number of independent and dependent variables, the research design (between-subjects vs. within-subjects), and adherence to distributional assumptions.
1. Assumptions of Parametric Statistical Tests
Parametric statistical tests (e.g., -tests, ANOVA, Pearson correlation, linear regression) estimate population parameters under specific mathematical assumptions. Violating these assumptions can inflate Type I error rates or diminish statistical power.
The Four Core Parametric Assumptions
- Scale of Measurement: The dependent variable must be measured on a continuous interval or ratio scale.
- Independence of Observations: Each observation within a group must be statistically independent of every other observation. Violations of independence (e.g., testing participants in social groups without accounting for clustering) represent the most severe threat to validity, inflating Type I error rates substantially.
- Normality: The population distribution from which samples are drawn (or the distribution of sample residuals) is normally distributed. Due to the Central Limit Theorem, parametric tests are relatively robust to moderate violations of normality when sample sizes are sufficiently large ( per group).
- Homogeneity of Variance (Homoscedasticity): In multi-group designs, the variance of the populations must be equal across all groups:
- Diagnostic Assessment: Homogeneity of variance is assessed empirically using Levene's test or Hartley's test. A statistically significant Levene's test () indicates that the assumption of equal variances has been violated.
- Remedy: When variances are unequal (heteroscedasticity), particularly with unequal sample sizes, researchers use Welch's -test (which applies Satterthwaite's adjusted degrees of freedom) or the Welch/Brown-Forsythe robust ANOVA.
2. Student's t-Test Family
Developed by William Sealy Gosset under the pseudonym "Student" while working at the Guinness Brewery, the -distribution accounts for the added uncertainty of estimating the population variance from a sample.
| -Test Type | Research Application | Formula | Degrees of Freedom () |
|---|---|---|---|
| One-Sample -Test | Compares a single sample mean () to a known or hypothesized population mean (). | ||
| Independent-Samples -Test | Compares means from two completely separate, unrelated groups of participants (between-subjects design). | ||
| Paired / Dependent-Samples -Test | Compares two means from the same participants tested twice (pretest-posttest) or matched participant pairs (within-subjects design). |
Why Paired t-Tests Yield Greater Statistical Power
In a paired-samples design, researchers compute a difference score for each participant:
By analyzing difference scores, the paired -test eliminates between-subject variability (individual differences in baseline intelligence, personality, or genetics) from the error denominator (). Because between-subject variance is often substantial, removing it shrinks the standard error of difference scores, yielding a larger -value and substantially higher statistical power than an independent-samples test with an equivalent number of observations.
3. Analysis of Variance (ANOVA)
When an experiment involves three or more group means, conducting multiple independent -tests causes exponential inflation of the familywise Type I error rate ().
The Problem of Multiple Comparisons
If an experiment compares groups, the number of unique pairwise comparisons is given by:
The probability of committing at least one Type I error across independent comparisons at is calculated as:
- For groups: comparisons ( error risk).
- For groups: comparisons ( error risk).
- For groups: comparisons ( error risk).
Ronald Fisher developed Analysis of Variance (ANOVA) to solve this problem. ANOVA performs a single omnibus test evaluating whether any differences exist among the group means while holding the overall false-positive rate at .
Partitioning Variance in One-Way Between-Subjects ANOVA
ANOVA decomposes the total variation in the data () into two independent additive components:
- Between-Groups Variance (): Reflects differences between group means and the grand mean. This variation stems from two sources: genuine experimental treatment effects PLUS random individual sampling error.
- Within-Groups Variance ( or Error Variance): Reflects variability among participants within the same treatment condition. Because all participants within a group receive identical treatment, this variability is attributable purely to random individual differences and measurement error.
Total Variance (SStotal, df = N - 1)
│
┌─────────────────────────────┴─────────────────────────────┐
▼ ▼
Between-Groups Variance Within-Groups Variance
(SSbetween, df = k - 1) (SSwithin, df = N - k)
Treatment Effect + Random Error Random Error Alone
│ │
└─────────────────────────────┬─────────────────────────────┘
▼
F-Ratio = MSbetween / MSwithin
The F-Ratio and Expected Mean Squares
Dividing each sum of squares by its respective degrees of freedom produces the Mean Squares ():
The test statistic is the -ratio:
- Under (No Treatment Effect): Treatment Effect . Therefore, . The expected value of under the null hypothesis is approximately .
- Under (Treatment Effect Exists): Treatment Effect . Therefore, exceeds , producing an -ratio substantially greater than . If exceeds the critical value, is rejected.
Post-Hoc Tests vs. Planned Comparisons
A significant omnibus -test confirms that at least one group mean differs from another, but does not indicate which specific pairs differ. Researchers resolve this ambiguity using planned contrasts or post-hoc tests:
- Planned (A Priori) Contrasts: Specified before data collection based on explicit theoretical hypotheses. They do not require a significant omnibus and possess higher statistical power.
- Post-Hoc (A Posteriori) Tests: Conducted after obtaining a significant omnibus to explore all possible pairwise differences while strictly controlling :
- Tukey's Honestly Significant Difference (HSD): The standard post-hoc test in psychology; controls familywise error across all pairwise comparisons while maintaining high power.
- Bonferroni Correction: Extremely versatile but conservative adjustment that divides nominal alpha by the number of comparisons ().
- Scheffé Test: The most conservative post-hoc test; controls familywise error for all possible simple and complex linear combinations of means.
4. Factorial ANOVA and Multivariate Extensions
Factorial ANOVA ( Designs)
In a factorial ANOVA, two or more categorical independent variables (factors) are manipulated simultaneously. In a factorial design, two independent variables each possess two levels, yielding four distinct treatment cells.
A two-way factorial ANOVA yields three independent -tests:
- Main Effect of Factor A: The effect of Factor A averaged across all levels of Factor B.
- Main Effect of Factor B: The effect of Factor B averaged across all levels of Factor A.
- Interaction Effect (): Evaluates whether the effect of Factor A on the dependent variable depends upon the specific level of Factor B.
Factorial ANOVA Interaction Graphs:
No Interaction (Parallel Lines) Significant Interaction (Non-Parallel Lines)
DV DV
High ┌────────────────────── B2 (Caffeine) High ┌─────────\────────────── B2 (Caffeine)
│ / │ \ /
│ / │ \ /
│ / │ \ /
│ / │ \ /
│ / │ \ /
│ / │ \ /
│ / │ \/
Low └───────────────/────── B1 (Placebo) Low └────────────────/\────── B1 (Placebo)
Low Stress High Stress Low Stress High Stress
Important
On the GRE Psychology Subject Test, interactions are frequently presented graphically. If the lines connecting condition means are parallel, there is no interaction. If the lines are non-parallel—whether they converge, diverge, or cross each other completely (crossover interaction)—a statistically significant interaction is present, indicating moderation.
Repeated-Measures ANOVA and Sphericity
When the same participants are tested across three or more experimental conditions, researchers employ a repeated-measures ANOVA. This design partitions inter-subject variation out of the error term, maximizing power.
- The Sphericity Assumption: Repeated-measures ANOVA requires sphericity—the assumption that the variances of the differences between all possible pairs of conditions are equal.
- Evaluation and Correction: Sphericity is evaluated using Mauchly's test of sphericity. If violated (), degrees of freedom are reduced using the Greenhouse-Geisser () or Huynh-Feldt epsilon adjustments to prevent Type I error inflation.
Analysis of Covariance (ANCOVA)
ANCOVA combines ANOVA with linear regression. It evaluates group differences on a continuous dependent variable while statistically controlling for one or more continuous extraneous variables known as covariates. Controlling for the covariate removes extraneous error variance, increasing statistical precision and power.
5. Correlation and Linear Regression
Pearson Product-Moment Correlation ()
Developed by Karl Pearson, quantifies the direction and strength of the linear association between two continuous variables ( and ):
- Properties:
- Values range strictly from to .
- The sign indicates direction: positive (direct) or negative (inverse).
- The absolute value () indicates strength.
- It is scale-invariant; linear transformations of or do not alter .
- Coefficient of Determination (): Squaring the correlation coefficient yields , representing the proportion of variance in that is predictable from or shared with . If , , meaning of the variance is shared and remains unexplained.
Critical Limitations and Nuances of Correlation
- Correlation Does Not Imply Causation: An observed correlation between and may stem from causing , causing (directionality problem), or an unmeasured third variable causing both (third-variable problem).
- Restriction of Range: Artificially truncating the range of either variable dramatically attenuates (depresses) the observed correlation coefficient relative to the true population value (e.g., assessing the correlation between SAT scores and college GPA exclusively among Ivy League students).
- Curvilinear Relationships: Pearson's evaluates strictly linear relationships. If two variables share a strong non-linear relationship—such as the Yerkes-Dodson Law relating physiological arousal to task performance (an inverted-U function)—the Pearson can be approximately despite a powerful, deterministic association.
Specialized Correlation Coefficients
| Coefficient | Symbol | Nature of Variable X | Nature of Variable Y | Research Example |
|---|---|---|---|---|
| Pearson Product-Moment | Continuous (Interval/Ratio) | Continuous (Interval/Ratio) | Neuroticism score and Beck Depression Inventory score | |
| Spearman Rank-Order | or | Ranked / Ordinal | Ranked / Ordinal (or non-linear monotonic) | Class rank in high school and class rank in college |
| Point-Biserial | Truly Dichotomous (Binary) | Continuous (Interval/Ratio) | Biological sex () and spatial rotation reaction time | |
| Biserial | Artificially Dichotomized | Continuous (Interval/Ratio) | Pass/Fail on an exam and continuous IQ score | |
| Phi Coefficient | Truly Dichotomous () | Truly Dichotomous () | Handedness (Left/Right) and Psychiatric diagnosis (Schizophrenia Yes/No) |
Simple and Multiple Linear Regression
While correlation describes association, regression uses scores on one or more predictor variables to predict performance on a criterion variable.
Simple Linear Regression
The trajectory of a simple bivariate regression line is expressed as:
- is the predicted value of the criterion variable.
- is the unstandardized regression slope, reflecting the predicted change in for every one-unit increase in :
- is the -intercept, representing the predicted value of when :
- Ordinary Least Squares (OLS) Criterion: The regression line is fitted by minimizing the sum of squared vertical residuals (prediction errors) between observed data points and the line:
Multiple Linear Regression
When multiple continuous predictors () predict a single continuous criterion ():
- Standardized Beta Weights (): When predictors are converted to standard scores (-scores), their coefficients are expressed as . A standardized beta weight represents the unique contribution of that specific predictor to the criterion, holding all other predictors constant.
- Multicollinearity: Occurs when two or more predictor variables in a multiple regression model are highly correlated with each other (). Multicollinearity inflates the standard errors of the regression coefficients, rendering individual estimates unstable and erratic.
6. Nonparametric Statistical Methods
When research data violate the assumptions of parametric tests—such as analyzing categorical counts, ordinal ranks, or small samples with extreme skewness and heteroscedasticity—researchers turn to nonparametric tests (distribution-free tests).
Chi-Square () Tests for Categorical Data
Chi-square tests evaluate frequencies of observations across discrete categorical bins.
1. Chi-Square Goodness-of-Fit Test
Evaluates whether an observed sample frequency distribution across a single nominal variable conforms to a theoretical or hypothesized population distribution:
- is the observed frequency in each category.
- is the expected frequency under the null hypothesis.
- is the number of categories.
2. Chi-Square Test of Independence
Evaluates whether two categorical variables are statistically independent or contingency-associated in a two-way contingency table:
- is the number of rows and is the number of columns.
- The expected frequency for each cell is computed as:
Nonparametric Rank-Order Tests and Their Parametric Counterparts
| Parametric Test | Nonparametric Equivalent | Primary Data Level | Test Mechanism |
|---|---|---|---|
| Independent-Samples -Test | Mann-Whitney Test (or Wilcoxon Rank-Sum) | Ordinal ranks (Between-subjects, 2 groups) | Converts continuous scores to ranks; evaluates whether the sum of ranks differs significantly between two independent groups. |
| Paired-Samples -Test | Wilcoxon Signed-Rank Test | Ordinal ranks (Within-subjects, 2 conditions) | Ranks the absolute differences between paired scores, assigns positive/negative signs, and sums ranks. |
| One-Way Between-Subjects ANOVA | Kruskal-Wallis -Test | Ordinal ranks (Between-subjects, groups) | Nonparametric omnibus test evaluating rank sum distributions across three or more independent groups. |
| Repeated-Measures ANOVA | Friedman Test | Ordinal ranks (Within-subjects, conditions) | Ranks scores across conditions within each individual participant; evaluates condition rank distributions. |
| Pearson Correlation () | Spearman Rank Correlation () | Ordinal ranks / Monotonic association | Computes Pearson correlation on the rank-transformed variables. |
A psycholinguistics researcher compares vocabulary acquisition rates across four independent instructional groups (k = 4). Rather than running an omnibus Analysis of Variance (ANOVA), the researcher conducts all possible pairwise independent t-tests at an alpha level of α = .05. What is the approximate familywise Type I error rate (α_FW) incurred by this testing strategy?
Approximately 5%
Approximately 15%
Approximately 26%
Approximately 50%
A 2 × 2 factorial experiment examines the effects of Task Difficulty (Easy vs. Difficult) and Environmental Noise (Quiet vs. Loud) on cognitive problem-solving accuracy. When the cell means are plotted with Task Difficulty on the x-axis, the line representing the Quiet condition displays a steep negative slope, while the line representing the Loud condition is completely horizontal and flat across both difficulty levels. What statistical conclusion is unambiguously demonstrated by this graph?
A statistically significant interaction effect is present between Task Difficulty and Environmental Noise
The experiment is invalid because regression lines in factorial ANOVA must always remain parallel
There are significant main effects for both variables, but no interaction effect
A ceiling effect has eliminated the within-group error variance across all four experimental cells
An organizational psychologist investigates the relationship between employee biological sex (coded categorically as Male = 0, Female = 1) and continuous scores on an objective standardized emotional intelligence inventory. Which correlation coefficient is mathematically designed to quantify this association?
Biserial correlation coefficient
Phi coefficient
Spearman rank correlation coefficient
Point-biserial correlation coefficient
A clinical neuropsychologist measures recovery time (in days) following traumatic brain injury across three distinct rehabilitation protocols. The data exhibit extreme positive skewness and severe heteroscedasticity across groups, and the sample size is modest (n = 8 per group). Which statistical test should the neuropsychologist employ to evaluate whether recovery times differ across the three protocols?
Wilcoxon signed-rank test
Kruskal-Wallis H-test
One-way between-subjects Analysis of Variance (ANOVA)
Mann-Whitney U test
Sections you finish are checked off in the contents.