12.2 DS Applied to Statistics, Averages, and Data Sets
Key Takeaways
- In Data Sufficiency, finding the arithmetic mean of a data set requires establishing the ratio of the sum of elements to the number of elements (Sum / n); knowing individual element values is never required if their aggregate sum is fixed.
- The median of an ordered set of n elements depends exclusively on middle positioning (the ((n+1)/2)-th element for odd n, or the mean of the two middle elements for even n); exterior values can vary infinitely without shifting the median.
- In evenly spaced sets (arithmetic progressions), the mean and median are mathematically identical and equal the average of the first and last terms ((First + Last) / 2).
- Standard deviation measures dispersion around the mean; adding or subtracting a constant to all terms leaves standard deviation completely unchanged, while multiplying all terms by constant k scales the standard deviation by |k|.
- A standard deviation of zero is uniquely equivalent to all elements in the data set being strictly identical, which is also equivalent to a range of zero.
12.2 DS Applied to Statistics, Averages, and Data Sets
Quick Summary: Statistical problems on GMAT Focus Data Sufficiency assess your grasp of aggregate structural properties and mathematical invariance rather than your ability to grind through descriptive statistics calculations. The arithmetic mean depends solely on the relationship $\text{Sum} = n \times \text{Mean}$. The median depends solely on the middle term(s) of an ordered set and is completely insulated from changes in extreme values. Standard deviation (SD) measures dispersion: shifting all elements by a constant $k$ preserves SD exactly, scaling all elements by $k$ scales SD by $|k|$, and an SD of zero is functionally equivalent to a range of zero (all elements are identical). For weighted averages, knowing the ratio of weights is fully sufficient to determine the combined average, even if the absolute sample sizes remain entirely unknown.
The Mean-Sum-Count Relationship in DS
The arithmetic mean of a data set containing $n$ elements is governed by the single fundamental identity:
Degrees of Freedom and Mean Sufficiency
In a data set of $n$ elements, there are $n$ degrees of freedom. However, if the question stem asks for the mean, you do not need $n$ separate equations to solve for each individual $x_i$. You only need to determine the aggregate sum $\sum x_i$ and the count $n$.
- Algebraic Cancellation: Test-makers frequently construct statements where unknown individual variables cancel out when summed. For example, if a prompt asks for the average of five numbers ${a, b, c, d, e}$ and Statement (1) provides $a + b = 20$ and $c + d + e = 35$, the sum is uniquely fixed at $20 + 35 = 55$, giving a mean of $55 / 5 = 11$, even though none of the five individual numbers can be determined!
- Linear Transformations: If a constant $k$ is added to every element in a data set, the new mean becomes $\text{Mean}{\text{new}} = \text{Mean}{\text{old}} + k$. If every element is multiplied by constant $c$, $\text{Mean}{\text{new}} = c \times \text{Mean}{\text{old}}$.
Median Sufficiency: Positional vs. Value Constraints
The median represents the numerical midpoint of an ordered distribution. Determining the median in Data Sufficiency requires understanding how ordered sets behave under odd versus even sample sizes:
The Positional Insulation Principle
A critical insight for GMAT Data Sufficiency is that the median is insulated against fluctuations in the values of non-median elements. Consider five numbers arranged in ascending order: $x_1 \le x_2 \le x_3 \le x_4 \le x_5$.
- The median is uniquely the middle term: $x_3$.
- To determine the median, you only need to know the value of $x_3$.
- The values of $x_1$ and $x_2$ can be any numbers less than or equal to $x_3$, and $x_4$ and $x_5$ can be any numbers greater than or equal to $x_3$. Their exact numerical values have zero impact on the median.
Symmetric and Evenly Spaced Sets
In any data set that is symmetric around its center—such as an arithmetic sequence or consecutive integers—the following identity holds:
On Data Sufficiency, if a statement establishes that a set consists of consecutive integers, consecutive multiples, or any evenly spaced sequence, any statement that provides the mean automatically provides the median, and vice versa!
Standard Deviation (SD) Sufficiency Without Calculation
Standard deviation measures the average dispersion of data values around their arithmetic mean:
The GMAT DS Golden Rule of Standard Deviation:
You will almost never be required to calculate the exact numerical standard deviation using this formula on the GMAT. Instead, DS questions test four structural invariance properties:
Property 1: Translation Invariance (Adding/Subtracting a Constant)
Adding or subtracting a constant $c$ to every term in a set shifts the entire distribution along the number line without altering its internal spacing. The distance between each point and the new mean remains identical:
Sufficiency Impact: If Statement (1) asserts that Set B is formed by adding 15 to each member of Set A, and Set A has a known standard deviation of 4.2, then Set B's standard deviation is immediately fixed at 4.2. Statement (1) is sufficient.
Property 2: Scale Factor Multiplications
Multiplying every term in a data set by a constant $k$ scales the distance of each point from the mean by $|k|$:
Sufficiency Impact: Multiplying by $-3$ multiplies the standard deviation by $|-3| = 3$. If the initial standard deviation is known, the new standard deviation is uniquely determined.
Property 3: Zero Standard Deviation
A standard deviation of zero indicates that there is zero dispersion around the mean:
If a statement establishes that $\sigma = 0$, all elements in the set are identical. Conversely, if a statement confirms that the range of a set is 0, $\sigma$ is uniquely fixed at 0.
Property 4: Equidistant Clustering
If a data set consists of values clustered symmetrically around a center (e.g., ${\mu - d, \mu - d, \mu, \mu + d, \mu + d}$), the standard deviation depends entirely on $d$ and sample size $n$, completely independent of the value of the mean $\mu$!
Weighted Averages Sufficiency: The Ratio of Weights
When combining two groups with different averages ($A_1$ and $A_2$) and sample sizes ($w_1$ and $w_2$), the combined weighted average is:
Dividing the numerator and denominator by $w_2$ exposes the fundamental DS shortcut:
The Proportionality Law of Weighted Averages in DS
To determine the combined weighted average of two subgroups with known subgroup averages $A_1$ and $A_2$:
- You do NOT need the absolute number of elements in either group ($w_1$ or $w_2$)!
- You ONLY need the RATIO of the weights: $\frac{w_1}{w_2}$ (or percentage distribution $p_1$ and $p_2$).
| Stated Information in DS | Can Combined Average Be Determined? | Rationale |
|---|---|---|
| Group 1 Average + Group 2 Average + Total Headcount ($w_1 + w_2$) | NO | Weight ratio $\frac{w_1}{w_2}$ is unconstrained; average could be weighted 99% to Group 1 or 99% to Group 2 |
| Group 1 Average + Group 2 Average + Weight Ratio (e.g., $w_1 = 3 w_2$) | YES | Proportions fix the lever arm directly: $A_{\text{combined}} = \frac{3 A_1 + A_2}{4}$ |
| Group 1 Average + Group 2 Average + Percentage Share ($w_1$ is 40% of total) | YES | $A_{\text{combined}} = 0.40 A_1 + 0.60 A_2$ |
Complete Worked DS Scenarios
Worked Example 1: Median Sufficiency Under Inequality Ordering
Question Stem:
Seven employees in a department earned different performance bonus amounts last year: $b_1 < b_2 < b_3 < b_4 < b_5 < b_6 < b_7$. What was the median bonus amount earned by the seven employees?Statement (1): The fourth-highest bonus amount earned was $12,500.
Statement (2): The sum of the lowest three bonus amounts was $24,000, and the sum of the highest three bonus amounts was $48,000.
Analytical Deconstruction
Because there are $n = 7$ distinct bonus amounts arranged in strict ascending order, the median is the $\frac{7+1}{2} = 4\text{th}$ element, which is $b_4$.
-
Evaluating Statement (1) Alone:
Counting from either the top or the bottom of a 7-element ordered set, the 4th element is identical:- 1st highest: $b_7$
- 2nd highest: $b_6$
- 3rd highest: $b_5$
- 4th highest: $b_4$ Statement (1) explicitly fixes $b_4 = $12,500$. Because $b_4$ is the exact median of 7 sorted values, the median is uniquely determined as $12,500. Statement (1) alone is SUFFICIENT.
-
Evaluating Statement (2) Alone:
Statement (2) provides $b_1 + b_2 + b_3 = 24,000$ and $b_5 + b_6 + b_7 = 48,000$. This gives the sums of the outer flanks, but leaves the middle term $b_4$ completely unconstrained. For example, $b_4$ could be $10,000$ or $15,000$ without violating the given conditions (e.g., $b_3 < b_4 < b_5$). Statement (2) alone is NOT sufficient. Conclusion: Statement (1) alone is sufficient, but statement (2) alone is not sufficient.
Worked Example 2: Standard Deviation Shift Invariance
Question Stem:
Data set $X$ consists of numbers ${x_1, x_2, x_3, x_4, x_5}$, and Data set $Y$ consists of numbers ${x_1 + 8, x_2 + 8, x_3 + 8, x_4 + 8, x_5 + 8}$. Is the standard deviation of Data set $Y$ greater than 5?Statement (1): The standard deviation of Data set $X$ is 6.2.
Statement (2): The arithmetic mean of Data set $Y$ is 28.
Analytical Deconstruction
Data set $Y$ is formed by adding the constant $c = 8$ to every element in Data set $X$. By the translation invariance property of dispersion, adding a constant shifts the mean by 8 but leaves every deviation from the mean unchanged: $(x_i + 8) - (\mu_X + 8) = x_i - \mu_X$. Therefore:
The question "Is $\sigma_Y > 5$?" is mathematically equivalent to "Is $\sigma_X > 5$?".
-
Evaluating Statement (1) Alone:
Statement (1) gives $\sigma_X = 6.2$. Because $\sigma_Y = \sigma_X$, we have $\sigma_Y = 6.2$. Since $6.2 > 5$, the answer to the question is a definitive, unconditional YES. Statement (1) alone is SUFFICIENT. -
Evaluating Statement (2) Alone:
Statement (2) gives the mean of Data set $Y$: $\mu_Y = 28$. The mean provides the central location of the data set but reveals nothing about its spread or standard deviation. Data set $Y$ could be ${28, 28, 28, 28, 28}$ with $\sigma_Y = 0$ (answer NO), or ${0, 0, 28, 56, 56}$ with $\sigma_Y > 5$ (answer YES). Statement (2) alone is NOT sufficient. Conclusion: Statement (1) alone is sufficient, but statement (2) alone is not sufficient.
High-Frequency Traps in Statistics DS
- The Full Calculation Trap: Believing that finding the mean, median, or standard deviation requires determining every individual value in the data set. GMAT questions are deliberately crafted so that individual values cannot be solved, but the aggregate statistical parameter is uniquely fixed.
- Conflating Range with Standard Deviation: Range is $\text{Max} - \text{Min}$, which measures extreme boundary distance. Standard deviation measures the average squared distance of all points from the center. Knowing that two sets have the identical range does not mean they have the same standard deviation.
- Even vs. Odd Median Oversight: For even sample sizes ($n=6, 8, 10$), test-takers often assume the median must be an element in the set. For even $n$, the median is the arithmetic mean of the two middle elements and might not appear in the data set at all.
- Absolute Weight Fallacy in Weighted Averages: Assuming that calculating a weighted average requires absolute sample counts rather than recognizing that relative weight percentages or ratios are completely sufficient.
A data set consists of 5 integers: {x₁, x₂, x₃, x₄, x₅}. What is the standard deviation of the data set? Statement (1): The range of the data set is 0. Statement (2): Every integer in the data set is equal to 14.
A set of five distinct integers is arranged in increasing order: k₁ < k₂ < k₃ < k₄ < k₅. What is the median of the set? Statement (1): k₃ = 27. Statement (2): The average (arithmetic mean) of the five integers is 27.
At a technology consulting firm, the staff consists entirely of junior analysts and senior consultants. The average annual salary of junior analysts is $70,000, and the average annual salary of senior consultants is $120,000. What is the average annual salary of all staff members at the firm? Statement (1): The firm employs 45 junior analysts. Statement (2): Junior analysts represent exactly 60 percent of all staff members at the firm.