15.3 Bivariate and Multivariate Techniques: t-Tests, ANOVA Models, Correlation, and Regression

Key Takeaways

  • Parametric tests require continuous interval/ratio data, independent observations, bivariate/multivariate normality, and homogeneity of variance across groups (evaluated via Levene's test).

  • Student's t-test evaluates mean differences across one-sample (df = N - 1), independent-samples (df = n1 + n2 - 2), and paired-samples (df = Npairs - 1) designs, with paired designs achieving greater power by partialing out inter-subject variance.

  • Analysis of Variance (ANOVA) resolves familywise error rate inflation (α_FW = 1 - (1 - α)^c) via an omnibus F-ratio (MS_between / MS_within); factorial ANOVA enables simultaneous testing of main effects and interaction dynamics (A × B).

  • Pearson r measures linear bivariate association and Ordinary Least Squares (OLS) regression calculates optimal predictive trajectories (Ŷ = bX + a), while non-parametric tests (Chi-Square, Mann-Whitney U, Wilcoxon, Kruskal-Wallis) analyze categorical or rank-ordered distributions free of distributional assumptions.

Last updated: October 2026

Bivariate and Multivariate Techniques: t-Tests, ANOVA Models, Correlation, and Regression

Psychological investigations evaluate hypotheses spanning diverse designs, from comparing treatment groups to modeling complex multivariate associations. Selecting the appropriate statistical test requires assessing the scale of measurement, the number of independent and dependent variables, the research design (between-subjects vs. within-subjects), and adherence to distributional assumptions.

1. Assumptions of Parametric Statistical Tests

Parametric statistical tests (e.g., tt-tests, ANOVA, Pearson correlation, linear regression) estimate population parameters under specific mathematical assumptions. Violating these assumptions can inflate Type I error rates or diminish statistical power.

The Four Core Parametric Assumptions

  1. Scale of Measurement: The dependent variable must be measured on a continuous interval or ratio scale.
  2. Independence of Observations: Each observation within a group must be statistically independent of every other observation. Violations of independence (e.g., testing participants in social groups without accounting for clustering) represent the most severe threat to validity, inflating Type I error rates substantially.
  3. Normality: The population distribution from which samples are drawn (or the distribution of sample residuals) is normally distributed. Due to the Central Limit Theorem, parametric tests are relatively robust to moderate violations of normality when sample sizes are sufficiently large (N≥30N \ge 30 per group).
  4. Homogeneity of Variance (Homoscedasticity): In multi-group designs, the variance of the populations must be equal across all groups:

σ12=σ22=⋯=σk2\sigma_1^2 = \sigma_2^2 = \dots = \sigma_k^2

  • Diagnostic Assessment: Homogeneity of variance is assessed empirically using Levene's test or Hartley's Fmax⁡F_{\max} test. A statistically significant Levene's test (p<.05p < .05) indicates that the assumption of equal variances has been violated.
  • Remedy: When variances are unequal (heteroscedasticity), particularly with unequal sample sizes, researchers use Welch's tt-test (which applies Satterthwaite's adjusted degrees of freedom) or the Welch/Brown-Forsythe robust ANOVA.

2. Student's t-Test Family

Developed by William Sealy Gosset under the pseudonym "Student" while working at the Guinness Brewery, the tt-distribution accounts for the added uncertainty of estimating the population variance from a sample.

tt-Test TypeResearch ApplicationFormulaDegrees of Freedom (dfdf)
One-Sample tt-TestCompares a single sample mean (Xˉ\bar{X}) to a known or hypothesized population mean (μ0\mu_0).t=Xˉ−μ0s/Nt = \frac{\bar{X} - \mu_0}{s / \sqrt{N}}df=N−1df = N - 1
Independent-Samples tt-TestCompares means from two completely separate, unrelated groups of participants (between-subjects design).t=Xˉ1−Xˉ2sp2n1+sp2n2t = \frac{\bar{X}_1 - \bar{X}_2}{\sqrt{\frac{s_p^2}{n_1} + \frac{s_p^2}{n_2}}}df=n1+n2−2df = n_1 + n_2 - 2
Paired / Dependent-Samples tt-TestCompares two means from the same participants tested twice (pretest-posttest) or matched participant pairs (within-subjects design).t=DˉsD/Nt = \frac{\bar{D}}{s_D / \sqrt{N}}df=Npairs−1df = N_{\text{pairs}} - 1

Why Paired t-Tests Yield Greater Statistical Power

In a paired-samples design, researchers compute a difference score for each participant:

Di=X1i−X2iD_i = X_{1i} - X_{2i}

By analyzing difference scores, the paired tt-test eliminates between-subject variability (individual differences in baseline intelligence, personality, or genetics) from the error denominator (sDs_D). Because between-subject variance is often substantial, removing it shrinks the standard error of difference scores, yielding a larger tt-value and substantially higher statistical power than an independent-samples test with an equivalent number of observations.

3. Analysis of Variance (ANOVA)

When an experiment involves three or more group means, conducting multiple independent tt-tests causes exponential inflation of the familywise Type I error rate (αFW\alpha_{\text{FW}}).

The Problem of Multiple Comparisons

If an experiment compares kk groups, the number of unique pairwise comparisons is given by:

c=k(k−1)2c = \frac{k(k - 1)}{2}

The probability of committing at least one Type I error across cc independent comparisons at α=.05\alpha = .05 is calculated as:

αFW=1−(1−α)c\alpha_{\text{FW}} = 1 - (1 - \alpha)^c

  • For k=3k = 3 groups: c=3c = 3 comparisons →αFW=1−(0.95)3≈.143\rightarrow \alpha_{\text{FW}} = 1 - (0.95)^3 \approx .143 (14.3%14.3\% error risk).
  • For k=4k = 4 groups: c=6c = 6 comparisons →αFW=1−(0.95)6≈.265\rightarrow \alpha_{\text{FW}} = 1 - (0.95)^6 \approx .265 (26.5%26.5\% error risk).
  • For k=5k = 5 groups: c=10c = 10 comparisons →αFW=1−(0.95)10≈.401\rightarrow \alpha_{\text{FW}} = 1 - (0.95)^{10} \approx .401 (40.1%40.1\% error risk).

Ronald Fisher developed Analysis of Variance (ANOVA) to solve this problem. ANOVA performs a single omnibus test evaluating whether any differences exist among the kk group means while holding the overall false-positive rate at α=.05\alpha = .05.

Partitioning Variance in One-Way Between-Subjects ANOVA

ANOVA decomposes the total variation in the data (SStotalSS_{\text{total}}) into two independent additive components:

SStotal=SSbetween+SSwithinSS_{\text{total}} = SS_{\text{between}} + SS_{\text{within}}

  1. Between-Groups Variance (SSbetweenSS_{\text{between}}): Reflects differences between group means and the grand mean. This variation stems from two sources: genuine experimental treatment effects PLUS random individual sampling error.
  2. Within-Groups Variance (SSwithinSS_{\text{within}} or Error Variance): Reflects variability among participants within the same treatment condition. Because all participants within a group receive identical treatment, this variability is attributable purely to random individual differences and measurement error.
                               Total Variance (SStotal, df = N - 1)
                                                │
                  ┌─────────────────────────────┴─────────────────────────────┐
                  ▼                                                           ▼
     Between-Groups Variance                                     Within-Groups Variance
    (SSbetween, df = k - 1)                                     (SSwithin, df = N - k)
  Treatment Effect + Random Error                                 Random Error Alone
                  │                                                           │
                  └─────────────────────────────┬─────────────────────────────┘
                                                ▼
                              F-Ratio = MSbetween / MSwithin

The F-Ratio and Expected Mean Squares

Dividing each sum of squares by its respective degrees of freedom produces the Mean Squares (MSMS):

MSbetween=SSbetweenk−1andMSwithin=SSwithinN−kMS_{\text{between}} = \frac{SS_{\text{between}}}{k - 1} \quad \text{and} \quad MS_{\text{within}} = \frac{SS_{\text{within}}}{N - k}

The test statistic is the FF-ratio:

F=MSbetweenMSwithin=Treatment Effect+Random ErrorRandom ErrorF = \frac{MS_{\text{between}}}{MS_{\text{within}}} = \frac{\text{Treatment Effect} + \text{Random Error}}{\text{Random Error}}

  • Under H0H_0 (No Treatment Effect): Treatment Effect =0= 0. Therefore, F=0+Random ErrorRandom Error≈1.0F = \frac{0 + \text{Random Error}}{\text{Random Error}} \approx 1.0. The expected value of FF under the null hypothesis is approximately 1.01.0.
  • Under H1H_1 (Treatment Effect Exists): Treatment Effect >0> 0. Therefore, MSbetweenMS_{\text{between}} exceeds MSwithinMS_{\text{within}}, producing an FF-ratio substantially greater than 1.01.0. If FF exceeds the critical value, H0H_0 is rejected.

Post-Hoc Tests vs. Planned Comparisons

A significant omnibus FF-test confirms that at least one group mean differs from another, but does not indicate which specific pairs differ. Researchers resolve this ambiguity using planned contrasts or post-hoc tests:

  • Planned (A Priori) Contrasts: Specified before data collection based on explicit theoretical hypotheses. They do not require a significant omnibus FF and possess higher statistical power.
  • Post-Hoc (A Posteriori) Tests: Conducted after obtaining a significant omnibus FF to explore all possible pairwise differences while strictly controlling αFW\alpha_{\text{FW}}:
    • Tukey's Honestly Significant Difference (HSD): The standard post-hoc test in psychology; controls familywise error across all pairwise comparisons while maintaining high power.
    • Bonferroni Correction: Extremely versatile but conservative adjustment that divides nominal alpha by the number of comparisons (αadj=α/c\alpha_{\text{adj}} = \alpha / c).
    • Scheffé Test: The most conservative post-hoc test; controls familywise error for all possible simple and complex linear combinations of means.

4. Factorial ANOVA and Multivariate Extensions

Factorial ANOVA (A×BA \times B Designs)

In a factorial ANOVA, two or more categorical independent variables (factors) are manipulated simultaneously. In a 2×22 \times 2 factorial design, two independent variables each possess two levels, yielding four distinct treatment cells.

A two-way factorial ANOVA yields three independent FF-tests:

  1. Main Effect of Factor A: The effect of Factor A averaged across all levels of Factor B.
  2. Main Effect of Factor B: The effect of Factor B averaged across all levels of Factor A.
  3. Interaction Effect (A×BA \times B): Evaluates whether the effect of Factor A on the dependent variable depends upon the specific level of Factor B.
Factorial ANOVA Interaction Graphs:

       No Interaction (Parallel Lines)              Significant Interaction (Non-Parallel Lines)
  DV                                           DV
  High ┌────────────────────── B2 (Caffeine)   High ┌─────────\────────────── B2 (Caffeine)
       │                      /                     │          \            /
       │                     /                      │           \          /
       │                    /                       │            \        /
       │                   /                        │             \      /
       │                  /                         │              \    /
       │                 /                          │               \  /
       │                /                           │                \/
   Low └───────────────/────── B1 (Placebo)     Low └────────────────/\────── B1 (Placebo)
      Low Stress   High Stress                     Low Stress   High Stress

Important

On the GRE Psychology Subject Test, interactions are frequently presented graphically. If the lines connecting condition means are parallel, there is no interaction. If the lines are non-parallel—whether they converge, diverge, or cross each other completely (crossover interaction)—a statistically significant interaction is present, indicating moderation.

Repeated-Measures ANOVA and Sphericity

When the same participants are tested across three or more experimental conditions, researchers employ a repeated-measures ANOVA. This design partitions inter-subject variation out of the error term, maximizing power.

  • The Sphericity Assumption: Repeated-measures ANOVA requires sphericity—the assumption that the variances of the differences between all possible pairs of conditions are equal.
  • Evaluation and Correction: Sphericity is evaluated using Mauchly's test of sphericity. If violated (p<.05p < .05), degrees of freedom are reduced using the Greenhouse-Geisser (ϵ^\hat{\epsilon}) or Huynh-Feldt epsilon adjustments to prevent Type I error inflation.

Analysis of Covariance (ANCOVA)

ANCOVA combines ANOVA with linear regression. It evaluates group differences on a continuous dependent variable while statistically controlling for one or more continuous extraneous variables known as covariates. Controlling for the covariate removes extraneous error variance, increasing statistical precision and power.

5. Correlation and Linear Regression

Pearson Product-Moment Correlation (rr)

Developed by Karl Pearson, rr quantifies the direction and strength of the linear association between two continuous variables (XX and YY):

r=Covariance(X,Y)sXsY=∑(X−Xˉ)(Y−Yˉ)∑(X−Xˉ)2∑(Y−Yˉ)2r = \frac{\text{Covariance}(X, Y)}{s_X s_Y} = \frac{\sum (X - \bar{X})(Y - \bar{Y})}{\sqrt{\sum (X - \bar{X})^2 \sum (Y - \bar{Y})^2}}

  • Properties:
    • Values range strictly from −1.00-1.00 to +1.00+1.00.
    • The sign indicates direction: positive (direct) or negative (inverse).
    • The absolute value (∣r∣|r|) indicates strength.
    • It is scale-invariant; linear transformations of XX or YY do not alter rr.
  • Coefficient of Determination (r2r^2): Squaring the correlation coefficient yields r2r^2, representing the proportion of variance in YY that is predictable from or shared with XX. If r=.50r = .50, r2=.25r^2 = .25, meaning 25%25\% of the variance is shared and 75%75\% remains unexplained.

Critical Limitations and Nuances of Correlation

  1. Correlation Does Not Imply Causation: An observed correlation between XX and YY may stem from XX causing YY, YY causing XX (directionality problem), or an unmeasured third variable ZZ causing both (third-variable problem).
  2. Restriction of Range: Artificially truncating the range of either variable dramatically attenuates (depresses) the observed correlation coefficient relative to the true population value (e.g., assessing the correlation between SAT scores and college GPA exclusively among Ivy League students).
  3. Curvilinear Relationships: Pearson's rr evaluates strictly linear relationships. If two variables share a strong non-linear relationship—such as the Yerkes-Dodson Law relating physiological arousal to task performance (an inverted-U function)—the Pearson rr can be approximately 0.000.00 despite a powerful, deterministic association.

Specialized Correlation Coefficients

CoefficientSymbolNature of Variable XNature of Variable YResearch Example
Pearson Product-MomentrrContinuous (Interval/Ratio)Continuous (Interval/Ratio)Neuroticism score and Beck Depression Inventory score
Spearman Rank-Orderrsr_s or ρ\rhoRanked / OrdinalRanked / Ordinal (or non-linear monotonic)Class rank in high school and class rank in college
Point-Biserialrpbr_{pb}Truly Dichotomous (Binary)Continuous (Interval/Ratio)Biological sex (0/10/1) and spatial rotation reaction time
Biserialrbr_bArtificially DichotomizedContinuous (Interval/Ratio)Pass/Fail on an exam and continuous IQ score
Phi Coefficientϕ\phiTruly Dichotomous (0/10/1)Truly Dichotomous (0/10/1)Handedness (Left/Right) and Psychiatric diagnosis (Schizophrenia Yes/No)

Simple and Multiple Linear Regression

While correlation describes association, regression uses scores on one or more predictor variables to predict performance on a criterion variable.

Simple Linear Regression

The trajectory of a simple bivariate regression line is expressed as:

Y^=bX+a\hat{Y} = bX + a

  • Y^\hat{Y} is the predicted value of the criterion variable.
  • bb is the unstandardized regression slope, reflecting the predicted change in YY for every one-unit increase in XX:

b=r(sYsX)b = r \left( \frac{s_Y}{s_X} \right)

  • aa is the YY-intercept, representing the predicted value of YY when X=0X = 0:

a=Yˉ−bXˉa = \bar{Y} - b\bar{X}

  • Ordinary Least Squares (OLS) Criterion: The regression line is fitted by minimizing the sum of squared vertical residuals (prediction errors) between observed data points and the line:

∑i=1N(Yi−Y^i)2=minimum\sum_{i=1}^N (Y_i - \hat{Y}_i)^2 = \text{minimum}

Multiple Linear Regression

When multiple continuous predictors (X1,X2,…,XkX_1, X_2, \dots, X_k) predict a single continuous criterion (YY):

Y^=b1X1+b2X2+⋯+bkXk+a\hat{Y} = b_1 X_1 + b_2 X_2 + \dots + b_k X_k + a

  • Standardized Beta Weights (β\beta): When predictors are converted to standard scores (zz-scores), their coefficients are expressed as β\beta. A standardized beta weight represents the unique contribution of that specific predictor to the criterion, holding all other predictors constant.
  • Multicollinearity: Occurs when two or more predictor variables in a multiple regression model are highly correlated with each other (r>.80r > .80). Multicollinearity inflates the standard errors of the regression coefficients, rendering individual β\beta estimates unstable and erratic.

6. Nonparametric Statistical Methods

When research data violate the assumptions of parametric tests—such as analyzing categorical counts, ordinal ranks, or small samples with extreme skewness and heteroscedasticity—researchers turn to nonparametric tests (distribution-free tests).

Chi-Square (χ2\chi^2) Tests for Categorical Data

Chi-square tests evaluate frequencies of observations across discrete categorical bins.

1. Chi-Square Goodness-of-Fit Test

Evaluates whether an observed sample frequency distribution across a single nominal variable conforms to a theoretical or hypothesized population distribution:

χ2=∑(O−E)2Ewithdf=k−1\chi^2 = \sum \frac{(O - E)^2}{E} \quad \text{with} \quad df = k - 1

  • OO is the observed frequency in each category.
  • EE is the expected frequency under the null hypothesis.
  • kk is the number of categories.

2. Chi-Square Test of Independence

Evaluates whether two categorical variables are statistically independent or contingency-associated in a two-way contingency table:

χ2=∑(O−E)2Ewithdf=(r−1)(c−1)\chi^2 = \sum \frac{(O - E)^2}{E} \quad \text{with} \quad df = (r - 1)(c - 1)

  • rr is the number of rows and cc is the number of columns.
  • The expected frequency for each cell is computed as:

E=(Row Total)×(Column Total)NtotalE = \frac{(\text{Row Total}) \times (\text{Column Total})}{N_{\text{total}}}

Nonparametric Rank-Order Tests and Their Parametric Counterparts

Parametric TestNonparametric EquivalentPrimary Data LevelTest Mechanism
Independent-Samples tt-TestMann-Whitney UU Test (or Wilcoxon Rank-Sum)Ordinal ranks (Between-subjects, 2 groups)Converts continuous scores to ranks; evaluates whether the sum of ranks differs significantly between two independent groups.
Paired-Samples tt-TestWilcoxon Signed-Rank TestOrdinal ranks (Within-subjects, 2 conditions)Ranks the absolute differences between paired scores, assigns positive/negative signs, and sums ranks.
One-Way Between-Subjects ANOVAKruskal-Wallis HH-TestOrdinal ranks (Between-subjects, ≥3\ge 3 groups)Nonparametric omnibus test evaluating rank sum distributions across three or more independent groups.
Repeated-Measures ANOVAFriedman TestOrdinal ranks (Within-subjects, ≥3\ge 3 conditions)Ranks scores across conditions within each individual participant; evaluates condition rank distributions.
Pearson Correlation (rr)Spearman Rank Correlation (rsr_s)Ordinal ranks / Monotonic associationComputes Pearson correlation on the rank-transformed variables.
Test Your Knowledge

A psycholinguistics researcher compares vocabulary acquisition rates across four independent instructional groups (k = 4). Rather than running an omnibus Analysis of Variance (ANOVA), the researcher conducts all possible pairwise independent t-tests at an alpha level of α = .05. What is the approximate familywise Type I error rate (α_FW) incurred by this testing strategy?

A

Approximately 5%

B

Approximately 15%

C

Approximately 26%

D

Approximately 50%

Test Your Knowledge

A 2 × 2 factorial experiment examines the effects of Task Difficulty (Easy vs. Difficult) and Environmental Noise (Quiet vs. Loud) on cognitive problem-solving accuracy. When the cell means are plotted with Task Difficulty on the x-axis, the line representing the Quiet condition displays a steep negative slope, while the line representing the Loud condition is completely horizontal and flat across both difficulty levels. What statistical conclusion is unambiguously demonstrated by this graph?

A

A statistically significant interaction effect is present between Task Difficulty and Environmental Noise

B

The experiment is invalid because regression lines in factorial ANOVA must always remain parallel

C

There are significant main effects for both variables, but no interaction effect

D

A ceiling effect has eliminated the within-group error variance across all four experimental cells

Test Your Knowledge

An organizational psychologist investigates the relationship between employee biological sex (coded categorically as Male = 0, Female = 1) and continuous scores on an objective standardized emotional intelligence inventory. Which correlation coefficient is mathematically designed to quantify this association?

A

Biserial correlation coefficient

B

Phi coefficient

C

Spearman rank correlation coefficient

D

Point-biserial correlation coefficient

Test Your Knowledge

A clinical neuropsychologist measures recovery time (in days) following traumatic brain injury across three distinct rehabilitation protocols. The data exhibit extreme positive skewness and severe heteroscedasticity across groups, and the sample size is modest (n = 8 per group). Which statistical test should the neuropsychologist employ to evaluate whether recovery times differ across the three protocols?

A

Wilcoxon signed-rank test

B

Kruskal-Wallis H-test

C

One-way between-subjects Analysis of Variance (ANOVA)

D

Mann-Whitney U test

Sections you finish are checked off in the contents.