15.1 Descriptive Statistics, Measures of Central Tendency, Variability, and Normal Distributions
Key Takeaways
Stevens' four measurement scales—nominal, ordinal, interval, and ratio—dictate permissible mathematical operations, with ratio being the only scale featuring a true, non-arbitrary absolute zero.
Central tendency metrics diverge under asymmetry: the mean minimizes squared deviations but is pulled toward the tail in skewed distributions, the median bisects rank-ordered observations and resists outliers, and the mode is the sole metric for nominal data.
Sample variance (s²) utilizes Bessel's correction with N - 1 degrees of freedom in the denominator to correct for systematic underestimation of population variance (σ²), functioning as an unbiased estimator.
The standard normal distribution is unimodal, symmetric, and mesokurtic (Mean = Median = Mode), adhering to the empirical 68–95–99.7 rule and permitting linear transformations into z-scores (μ = 0, σ = 1), T-scores (μ = 50, σ = 10), and deviation IQs (μ = 100, σ = 15).
Descriptive Statistics, Measures of Central Tendency, Variability, and Normal Distributions
Descriptive statistics organize, summarize, and communicate numerical data collected from psychological investigations. Before applying any statistical test, researchers must identify the scale of measurement of their variables, determine the central tendencies and dispersion of their observations, and evaluate the underlying shape of the distribution.
1. S.S. Stevens' Scales of Measurement
In 1946, psychophysicist Stanley Smith Stevens published a landmark taxonomy in Science categorizing measurement into four distinct hierarchical levels: nominal, ordinal, interval, and ratio. Each scale possesses specific mathematical properties that dictate which statistical operations and transformations are permissible.
| Measurement Scale | Defining Properties | Mathematical Relations Permitted | Central Tendency | Permissible Dispersion | Psychological Examples |
|---|---|---|---|---|---|
| Nominal | Categorical classification; mutual exclusivity and exhaustiveness; no quantitative magnitude or rank order. | Equivalence ( or ) | Mode | None (frequency counts, index of diversity) | Diagnostic categories (DSM-5 disorders), biological sex, experimental condition assignment (control vs. treatment). |
| Ordinal | Ranked order along a continuum; indicates relative standing ( or ), but intervals between adjacent ranks are unequal or unknown. | Greater than / less than (, ) | Median (also mode) | Range, Interquartile Range () | Likert scale ratings (e.g., to ), tournament placements (1st, 2nd, 3rd), socioeconomic status (low, middle, high). |
| Interval | Ordered ranks with equal quantitative distances between adjacent scale points; arbitrary zero point (zero does not indicate absence of the attribute). | Addition and subtraction (, ) | Mean, Median, Mode | Variance (), Standard Deviation () | Temperature in Celsius or Fahrenheit ( is not the absence of heat), calendar years, standard Wechsler deviation IQ scores. |
| Ratio | Equal intervals accompanied by a true, non-arbitrary absolute zero point representing complete absence of the measured property. | All arithmetic operations (, , , ); meaningful ratios | Mean, Median, Mode (also Geometric Mean) | Variance, Standard Deviation, Coefficient of Variation | Reaction time (milliseconds), number of errors on a spatial recall task, galvanic skin response (conductance in microsiemens), Kelvin temperature scale. |
The Critical Distinction: Interval vs. Ratio Scales
A perennial testing point on the GRE Subject Test in Psychology is the distinction between an arbitrary zero and an absolute zero:
- Arbitrary Zero (Interval): On an interval scale, zero is simply an agreed-upon convention. In Celsius, represents the freezing point of water, not the physical absence of thermal energy. Consequently, one cannot claim that is "twice as hot" as . Similarly, an individual with a Wechsler IQ of 140 does not possess "twice the intelligence" of someone with an IQ of 70, nor does an IQ score of 0 represent zero intellectual capacity.
- Absolute Zero (Ratio): On a ratio scale, zero signifies total absence. A participant who takes 400 milliseconds to respond to a visual stimulus took objectively twice as long as a participant responding in 200 milliseconds (). Ratios of measurement are mathematically valid only when the scale originates from an absolute zero.
2. Measures of Central Tendency
Measures of central tendency identify a single representative numerical value around which observations cluster.
The Mode
The mode is the most frequently occurring score in a distribution.
- Characteristics: A distribution may have a single mode (unimodal), two modes (bimodal), multiple modes (multimodal), or no mode if all scores occur with equal frequency.
- Unique Advantage: The mode is the only measure of central tendency that can be used with nominal data (e.g., identifying the most common psychiatric diagnosis in an inpatient clinic).
- Limitation: It ignores the numerical value of most observations in a quantitative dataset and can fluctuate markedly between samples.
The Median
The median () is the 50th percentile—the middle value that bisects a rank-ordered distribution such that exactly 50% of scores fall at or below it and 50% fall at or above it.
- Computation:
- For an odd sample size , the median is the observation at rank .
- For an even sample size , the median is the arithmetic mean of the two central observations at ranks and .
- Mathematical Property: The median minimizes the sum of absolute deviations across all observations:
- Robustness: The median is resistant to extreme scores (outliers). Whether the highest score in a sample of five reaction times () is or , the median remains unchanged at . Hence, the median is the preferred measure of central tendency for heavily skewed distributions (e.g., household income, housing prices, response latencies).
The Arithmetic Mean
The arithmetic mean (sample mean or population mean ) is the sum of all scores divided by the total number of observations:
- Foundational Algebraic Properties:
- Sum of Deviations Equals Zero: The sum of signed deviations of every individual score from the arithmetic mean is always identically zero:
- Least Squares Criterion: The mean minimizes the sum of squared deviations across all scores. No other constant produces a smaller value for :
- Sensitivity to Outliers: Because the mean incorporates the precise numerical value of every score, it is highly sensitive to extreme scores. A single outlier in the tail pulls the mean toward itself, rendering it misleading in skewed distributions.
3. Measures of Variability (Dispersion)
Central tendency captures the midpoint of a dataset, but variability describes how dispersed or tightly clustered the scores are around that center.
Range and Interquartile Range
- Range: The crude distance between the extreme scores:
- While intuitive, the range is highly unstable because it is determined entirely by two extreme scores and generally increases as sample size grows.
- Interquartile Range (): The range of the middle 50% of rank-ordered observations, defined as the distance between the 75th percentile (third quartile, ) and the 25th percentile (first quartile, ):
- The semi-interquartile range is defined as .
- In box-and-whisker plots (boxplots), the box spans the , the interior line indicates the median, and whiskers typically extend to beyond the quartiles. Observations outside this boundary are identified as outliers.
Variance and Standard Deviation
To quantify the dispersion of all individual data points around the mean, statisticians compute the Sum of Squared Deviations ():
Population Variance vs. Sample Variance (Bessel's Correction)
When calculating the variance of an entire population of size , the parameter is computed as:
However, when estimating the population variance from a sample of size , dividing by produces a biased estimator that systematically underestimates the true population variance. Because sample observations cluster more closely around the sample mean than around the true population mean , the sum of squared deviations from is smaller than the sum of squared deviations from .
To correct for this negative bias, German astronomer Friedrich Bessel introduced Bessel's correction, replacing with degrees of freedom () in the denominator of the sample variance ():
Dividing by renders an unbiased estimator, meaning that the expected value of the sample variance across repeated random samples equals the population parameter: .
Degrees of Freedom (df = N - 1) Concept:
If N = 4 scores must sum to a Mean of 10 (Total Sum = 40):
Score 1 = 8
Score 2 = 12
Score 3 = 11
Score 4 = MUST BE 9 <-- Exactly 1 score is mathematically constrained!
Only N - 1 = 3 scores are free to vary.
Standard Deviation
Because variance is expressed in squared units of measurement (e.g., or ), taking the positive square root returns the metric of dispersion to the original scale of measurement:
4. Distribution Shapes: Skewness and Kurtosis
Distributions in psychological research vary along two primary dimensions of shape: symmetry (skewness) and peakedness/tail weight (kurtosis).
Skewness (Asymmetry)
- Normal / Symmetric Distribution: Skewness is zero. The distribution can be folded in half along the median to produce mirror images. In a perfectly symmetric unimodal distribution:
- Positive Skew (Right-Skewed): The tail of the distribution extends asymptotically toward the positive (right) end of the horizontal axis. A cluster of low scores is balanced by a small number of extraordinarily high scores.
- Ordering of Central Tendencies: The mode remains at the peak of the cluster, the median is pulled slightly rightward, and the mean is pulled farthest into the long tail:
- Psychological Examples: Annual income, reaction times (participants cannot respond faster than physiological limits, but inattentive trials produce extreme positive latencies), number of depressive symptoms endorsed in non-clinical student samples.
- Negative Skew (Left-Skewed): The tail extends toward the negative (left) end of the horizontal axis. A cluster of high scores is accompanied by a small tail of very low scores.
- Ordering of Central Tendencies: The mean is pulled farthest toward the negative tail, followed by the median, while the mode sits at the high peak:
- Psychological Examples: Age at death in modern industrialized nations, scores on an exceptionally easy exam that produces a ceiling effect.
Positive Skew (Right-Skewed): Negative Skew (Left-Skewed):
Mode Mode
/\ /\
/ \ / \
/ \ / \
/ \ / \
/ Median\ / Median \
/ \ / \
/ Mean \ / Mean \
/ \________ ________/ \
------------------------> ------------------------>
Mode < Median < Mean Mean < Median < Mode
(Tail points Right) (Tail points Left)
Kurtosis (Peakedness and Tail Weight)
Kurtosis reflects the heaviness of a distribution's tails and the sharpness of its central peak relative to a standard normal curve:
- Mesokurtic: A normal distribution with zero excess kurtosis ().
- Leptokurtic (): Characterized by a sharp, elevated central peak and fat, heavy tails. More observations are concentrated near the mean and in the extreme tails than in a normal distribution, with fewer in the intermediate shoulders.
- Platykurtic (): Characterized by a broad, flat central peak and thin, light tails. Observations are dispersed more evenly across the range.
5. The Normal Distribution and Standardized Scores
The normal distribution (Gaussian distribution) is a continuous, symmetrical, mesokurtic bell-shaped distribution mathematically defined by two parameters: its mean () and standard deviation (). Its inflection points, where the curve transitions from convex to concave, occur at exactly .
The Empirical Rule (68–95–99.7 Rule)
In any normal distribution, fixed proportions of observations fall within integer standard deviation intervals from the mean:
- : Encompasses approximately of all scores ( between the mean and , and between the mean and ).
- : Encompasses approximately of all scores ( on either side of the mean; exactly encompasses ).
- : Encompasses approximately of all scores ( on either side of the mean).
- Beyond : Only of observations lie beyond three standard deviations ( in each extreme tail).
Standard Normal Distribution Areas
Mean
│
34.13% │ 34.13%
┌─────────────┼─────────────┐
13.59% │ │ │ 13.59%
┌────────────┤ │ ├────────────┐
2.14% │ │ │ │ │ 2.14%
───────┴────────────┴─────────────┴─────────────┴────────────┴───────
-3σ -2σ -1σ 0 +1σ +2σ +3σ
0.13% 0.13%
Standardized z-Scores
A -score represents the signed distance between an individual raw score () and the population mean () in units of standard deviation ():
- Properties of the -Score Distribution:
- The mean of a -score distribution is always .
- The standard deviation of a -score distribution is always .
- Converting raw scores to -scores does not normalize a non-normal distribution; it is a linear transformation that preserves the exact skewness, kurtosis, and relative distances of the raw distribution.
Percentile Equivalence in the Normal Curve
Because the areas under the standard normal curve are mathematically fixed, -scores map directly to percentile ranks:
- (the median)
Linear Transformations of Standardized Scores
To eliminate negative values and decimal fractions, psychometricians transform raw -scores into alternative standardized metrics using the general linear formula:
| Score Metric | New Mean () | New SD () | Transformation Formula | Typical Psychological Application |
|---|---|---|---|---|
| -Score | Foundational statistical standardization | |||
| -Score | Clinical personality inventories (MMPI-2, MCMI). A score of () denotes clinical significance. | |||
| Wechsler Deviation IQ | WAIS-5, WISC-V. Cutoff for intellectual disability is (; bottom ); gifted threshold is (). | |||
| Stanford-Binet (Early Form L-M) | Historical intelligence assessment | |||
| Graduate Record Exam (GRE General) | Traditional standardized aptitude testing | |||
| Stanine (Standard Nine) | Scaled integers | Educational tracking and military aptitude batteries |
Note
A classic GRE psychology calculation involves cross-metric conversion. For example, if an examinee scores a -score of on an MMPI subscale, their -score is . If mapped onto a Wechsler IQ scale with the same relative standing, their equivalent IQ would be (the 97.7th percentile).
A cognitive psychologist measures the time (in milliseconds) required for participants to identify whether an auditory probe was previously presented in a memory set. What scale of measurement does this dependent variable represent, and what is its primary mathematical property?
Nominal scale, because reaction times are classified into discrete experimental speed categories
Ordinal scale, because reaction times can only establish relative rank ordering between subjects
Ratio scale, because milliseconds have equal intervals and a true zero point meaning no time has elapsed
Interval scale, because differences between milliseconds are equal but reaction time has an arbitrary zero point
A clinical neuropsychologist administers a newly designed cognitive screening battery to a large normative cohort of healthy older adults. Because the screening battery was intentionally constructed to be accessible and non-taxing, most participants obtain near-perfect scores, though a small subset with subtle impairments score substantially lower. Which of the following relationships among central tendency metrics will characterize this distribution?
Mode < Median < Mean
Mean = Median = Mode
Median < Mean < Mode
Mean < Median < Mode
Why does the calculation of the sample variance formula incorporate degrees of freedom (N - 1) in the denominator rather than the total sample size (N)?
Scores cluster more tightly around the sample mean than the population mean, so N - 1 yields an unbiased estimate
Sample standard deviations are squared quantities that require a fractional reduction to match the original scale
Dividing by N - 1 is an arbitrary mathematical convention intended to make hand calculations simpler
The sample size N is reserved exclusively for non-parametric tests, whereas N - 1 is mandatory for all parametric statistics
An adolescent completes a standardized psychological assessment battery. On an executive functioning inventory that utilizes T-scores (Mean = 50, SD = 10), the adolescent achieves a score of 70. If this performance is transformed to a Wechsler Deviation IQ scale (Mean = 100, SD = 15), what would be the adolescent's corresponding score, and what percentile rank does it approximate in a normal distribution?
Score = 120; approximately 95th percentile
Score = 115; approximately 84th percentile
Score = 145; approximately 99.9th percentile
Score = 130; approximately 98th percentile
Sections you finish are checked off in the contents.