5.2 Mean, Median, Mode, Range, Interquartile Range, and Standard Deviation
Key Takeaways
The mean is the arithmetic average sensitive to extreme outliers; the median is the positional midpoint resistant to extreme values; the mode is the most frequently occurring value.
In skewed distributions, the mean is pulled in the direction of the long tail: right-skewed distributions exhibit Mean > Median, while left-skewed distributions exhibit Mean < Median.
The five-number summary (Min, Q1, Median, Q3, Max) divides ordered data into four quartiles of 25% each; the Interquartile Range (IQR = Q3 - Q1) measures the spread of the middle 50%.
The 1.5 × IQR rule identifies statistical outliers as any value strictly less than Q1 - 1.5 × IQR or strictly greater than Q3 + 1.5 × IQR.
Adding a constant to all data values shifts measures of center by that constant but leaves measures of spread completely unchanged; multiplying by a constant scales both center and spread.
Mean, Median, Mode, Range, Interquartile Range, and Standard Deviation
OpenExamPrep provides this quantitative reasoning review to help students master numerical summaries of central tendency and dispersion for the Texas Success Initiative Assessment 2.0 (TSIA2) Mathematics section. Being able to compute, compare, and critically evaluate statistical measures is crucial for analyzing data sets across academic disciplines and real-world workplace environments.
1. Measures of Central Tendency
Measures of central tendency identify a single numerical value that represents the center or typical value of a quantitative data distribution.
The Arithmetic Mean (Average)
The mean (denoted x̄ for a sample or μ for a population) is the sum of all observed values divided by the total number of observations n:
x̄ = (∑ x) ÷ n = (x₁ + x₂ + ... + xₙ) ÷ n
- Sensitivity (Non-Resistant): The mean incorporates every individual numerical value. Consequently, it is highly sensitive to extreme values (outliers) and distribution skewness.
- The Fundamental Sum Property: The formula can be rearranged to express the cumulative sum of all data values:
This relationship is essential for solving TSIA2 missing-score and target-average problems.∑ x = n × x̄
Worked Scenario: Target Examination Score
A student has completed 4 examinations in an introductory biology course with
scores of 74, 82, 86, and 78. What score must the student earn on the 5th
examination to achieve an overall course mean of 82?
Step 1: Determine the required cumulative point total for all 5 exams:
Required Sum = n × Target Mean = 5 × 82 = 410 points
Step 2: Calculate the current point sum from the first 4 exams:
Current Sum = 74 + 82 + 86 + 78 = 320 points
Step 3: Subtract the current sum from the required cumulative sum:
Fifth Exam Score = 410 - 320 = 90 points
Verification: (74 + 82 + 86 + 78 + 90) ÷ 5 = 410 ÷ 5 = 82.0.
The Median
The median (denoted M or Q2) is the physical middle value when observations are arranged in ascending numerical order:
- Odd Sample Size (n is odd): The median is the single middle observation located at position
(n + 1) ÷ 2.- Example: For ordered data
{4, 7, 9, 12, 18}(n = 5), the median is at position(5 + 1) ÷ 2 = 3→Median = 9.
- Example: For ordered data
- Even Sample Size (n is even): The median is the arithmetic mean of the two middle observations located at positions
n ÷ 2and(n ÷ 2) + 1.- Example: For ordered data
{6, 8, 11, 15, 19, 24}(n = 6), the middle positions are6 ÷ 2 = 3and3 + 1 = 4(values 11 and 15) →Median = (11 + 15) ÷ 2 = 26 ÷ 2 = 13.
- Example: For ordered data
- Resistance (Robustness): The median is resistant to extreme outliers because it depends entirely on the rank order of the observations, not on the numerical magnitude of the extreme endpoints.
The Mode
The mode is the observation that appears with the highest frequency in a data set:
- A distribution may be unimodal (one mode), bimodal (two modes), multimodal (three or more modes), or have no mode (if all values occur with equal frequency).
- The mode is the only measure of central tendency applicable to nominal categorical data (e.g., the modal vehicle color in a parking garage).
Weighted Mean and Combined Group Means
When subgroups have different sample sizes, or when individual scores carry different percentage weights (such as course grading policies), taking a simple unweighted average produces an incorrect result. The weighted mean must be used:
x̄_w = [ ∑ (wᵢ × xᵢ) ] ÷ [ ∑ wᵢ ]
Worked Scenario: Combined Group Mean
Lecture Section A has 30 students with a mean exam score of 76. Lecture Section B
has 20 students with a mean exam score of 86. What is the combined mean score
for all 50 students across both sections?
Step 1: Compute the cumulative points accumulated in Section A:
Sum_A = 30 students × 76 points = 2,280 points
Step 2: Compute the cumulative points accumulated in Section B:
Sum_B = 20 students × 86 points = 1,720 points
Step 3: Combine total points and divide by the total number of students:
Total Points = 2,280 + 1,720 = 4,000 points
Total Students = 30 + 20 = 50 students
Combined Mean = 4,000 ÷ 50 = 80.0 points
⚠️ Caution: Computing the simple average (76 + 86) ÷ 2 = 81.0 is INCORRECT
because Section A contains 10 more students than Section B, weighting the
combined average closer to 76.
2. Measures of Dispersion (Spread)
Measures of dispersion quantify the variability, scatter, or degree of spread among observations in a quantitative data distribution.
Range
The range is the simplest measure of dispersion, defined as the difference between the maximum and minimum observations:
Range = Maximum - Minimum
Although straightforward to calculate, the range relies exclusively on the two most extreme observations and conveys no information about the internal distribution of values.
The Five-Number Summary and Quartiles
The five-number summary partitions an ordered data set into four quarters, each containing approximately 25% of the data values:
- Minimum (Min): The smallest observation in the data set.
- First Quartile (Q1): The 25th percentile; the median of the lower half of data (excluding the overall median when n is odd).
- Median (Q2): The 50th percentile; the midpoint of the entire distribution.
- Third Quartile (Q3): The 75th percentile; the median of the upper half of data (excluding the overall median when n is odd).
- Maximum (Max): The largest observation in the data set.
Interquartile Range (IQR)
The Interquartile Range (IQR) measures the spread of the central 50% of the observations:
IQR = Q3 - Q1
Because it ignores the lowest 25% and highest 25% of the data, the IQR is resistant to outliers.
The 1.5 × IQR Outlier Rule
A standardized statistical criterion is used to detect potential outliers:
Lower Fence = Q1 - 1.5 × IQR
Upper Fence = Q3 + 1.5 × IQR
- Any observation strictly less than the Lower Fence is an outlier:
x < Q1 - 1.5 × IQR. - Any observation strictly greater than the Upper Fence is an outlier:
x > Q3 + 1.5 × IQR.
Worked Scenario: Outlier Identification Walkthrough
Consider the ordered exam score data set (n = 12):
{45, 68, 72, 75, 78, 80, 82, 85, 88, 90, 92, 100}
Step 1: Find the median (Q2):
n = 12 is even; average the 6th and 7th values: (80 + 82) ÷ 2 = 81.0
Step 2: Find Q1 (median of lower 6 values: {45, 68, 72, 75, 78, 80}):
Average the 3rd and 4th values: (72 + 75) ÷ 2 = 73.5
Step 3: Find Q3 (median of upper 6 values: {82, 85, 88, 90, 92, 100}):
Average the 3rd and 4th values: (88 + 90) ÷ 2 = 89.0
Step 4: Compute IQR:
IQR = Q3 - Q1 = 89.0 - 73.5 = 15.5
Step 5: Calculate 1.5 × IQR and establish fences:
1.5 × IQR = 1.5 × 15.5 = 23.25
Lower Fence = Q1 - 23.25 = 73.5 - 23.25 = 50.25
Upper Fence = Q3 + 23.25 = 89.0 + 23.25 = 112.25
Step 6: Evaluate data points against fences:
- Minimum value 45 < 50.25 → 45 is a confirmed lower outlier!
- Maximum value 100 < 112.25 → 100 is within the fence (not an outlier).
Box-and-Whisker Plots
A box plot graphically displays the five-number summary:
- The central rectangular box spans from Q1 to Q3, with the box width equal to the IQR.
- A vertical line inside the box marks the position of the Median.
- Whiskers extend outward from the box edges to the smallest and largest observations that fall within the fences.
- Outliers falling beyond the fences are plotted as isolated individual points (dots or asterisks).
Standard Deviation (Conceptual Interpretation)
The standard deviation (s for sample, σ for population) quantifies the typical distance that individual observations deviate from their arithmetic mean:
s = √[ ∑ (xᵢ - x̄)² ÷ (n - 1) ]
- Non-Negativity: Standard deviation is always non-negative:
s ≥ 0. It equals zero if and only if every single observation in the data set has the exact same value. - Spread Comparison: A data set whose values are clustered tightly around the mean has a small standard deviation; a data set whose values are widely dispersed has a large standard deviation.
- Non-Resistance: Like the mean, standard deviation is heavily influenced by outliers because distances from the mean are squared.
3. Distribution Shapes and Skewness: Mean vs. Median
The shape of a data distribution dictates the mathematical relationship between its mean and median:
| Distribution Profile | Visual Description | Central Tendency Relationship | Practical Examples |
|---|---|---|---|
| Symmetric (Bell-shaped) | Single central peak; balanced left and right tails | Mean ≈ Median ≈ Mode | Adult heights, standardized test scores |
| Right-Skewed (Positively Skewed) | Majority of data clustered at lower values; long tail stretches right | Mean > Median | Annual household incomes, home sale prices |
| Left-Skewed (Negatively Skewed) | Majority of data clustered at higher values; long tail stretches left | Mean < Median | Retirement ages, scores on an easy exam |
Understanding the Skewness Mechanism:
In a right-skewed distribution of home prices:
{ $200k, $220k, $230k, $250k, $270k, $290k, $1,500k }
- The median is $250k (the middle value).
- The mean is ($2,960k) ÷ 7 = $422.9k.
The single luxury home at $1,500,000 pulls the sum and the mean dramatically
into the right tail, while the median remains anchored at $250k.
Conclusion: For skewed distributions or data with outliers, the MEDIAN and IQR
are the preferred measures of center and spread.
4. Effect of Linear Transformations on Data
When every observation in a quantitative data set is modified by a linear transformation y = ax + b:
Adding or Subtracting a Constant (b)
- Measures of Center (Mean, Median, Mode): Shift by exactly b:
New Center = Old Center + b. - Measures of Spread (Range, IQR, Standard Deviation): Completely unchanged:
New Spread = Old Spread.- Why? Adding 5 points to every score moves the entire distribution 5 units to the right along the number line, but the relative distances between data points remain identical.
Multiplying or Dividing by a Constant (a)
- Measures of Center: Scaled by a:
New Center = a × Old Center. - Measures of Spread: Scaled by |a|:
New Spread = |a| × Old Spread.- Why? Converting currency or units (e.g., feet to inches, multiplying by 12) expands both the values and the distances between values by a factor of 12.
5. TSIA2 Exam Traps & Strategic Checkpoints
- Trap 1: Forgetting to Order Data First: Never identify the median or quartiles from an unordered list. Always arrange data in ascending numerical order first.
- Trap 2: Skewness Direction Confusion: Skewness is named after the direction of the tail, not the peak! A distribution with a peak on the left and a long tail stretching to the right is right-skewed (Mean > Median).
- Trap 3: Averaging Averages of Different Group Sizes: Never compute
(Mean₁ + Mean₂) ÷ 2unless the two groups have identical sample sizes. You must compute total points divided by total individuals. - Trap 4: Outlier Fence Anchor Points: In the 1.5 × IQR rule, subtract from Q1 (
Q1 - 1.5 × IQR) and add to Q3 (Q3 + 1.5 × IQR). Subtracting from the median or mean is an intentional test-maker distractor.
A college student has completed four statistics examinations with scores of 76, 82, 88, and 84. If the course grade is calculated as the arithmetic mean of five exams, what score must the student earn on the fifth examination to achieve an overall course mean of 85?
89
95
92
90
An ordered data set contains 11 values representing daily hours worked: 2, 3, 5, 6, 7, 8, 9, 10, 11, 13, 22. Using the standard 1.5 × IQR rule, which of the following statements correctly identifies the interquartile range and any outliers?
IQR = 8; there are no statistical outliers in the data set
IQR = 5; the value 2 is an outlier below the lower fence
IQR = 6; the value 22 is an outlier above the upper fence
IQR = 6; both 2 and 22 are outliers beyond the fences
A distribution of annual starting salaries for 1,200 entry-level software engineering graduates is strongly right-skewed (positively skewed) due to several graduates receiving exceptionally high equity bonuses. Which relationship between the mean and median salary is expected?
The mean and median will be exactly equal because salary data are quantitative continuous
The median will be significantly greater than the mean because most employees earn typical wages
The standard deviation will equal the interquartile range because of the skewness
The mean will be greater than the median because extreme high salaries pull the arithmetic average into the right tail
A professor realizes that an exam was excessively challenging and decides to add 8 points to every student's score. How does this adjustment affect the class mean and standard deviation?
The mean increases by 8 points, and the standard deviation remains unchanged
Both the mean and the standard deviation increase by 8 points
The mean remains unchanged, and the standard deviation increases by 8 points
The mean increases by 8 points, and the standard deviation increases by √8 points
Sections you finish are checked off in the contents.