14.2 Data Handling and Probability
Key Takeaways
- MAT data handling is school statistics plus probability rules, denser than AQL Quantitative Literacy chart-reading: you compute measures and apply probability, not only read a campus table
- The 2015 MAT booklet lists histograms, line graphs, pie charts, ogives, and box-and-whisker plots; current Test Content says interpret representations and measures of data
- Mean follows every value; median and mode resist outliers; a constant data set has standard deviation 0; adding a constant leaves SD unchanged
- An ogive is cumulative frequency against the upper class boundary — the median is the first class whose running total passes n/2
- Use a Venn diagram for overlapping sets and a tree for sequential draws; P(A ∪ B) = P(A) + P(B) − P(A ∩ B); independence is a product rule, not ‘no overlap’
14.2 Data Handling and Probability
Quick Answer: MAT data handling is school statistics plus probability rules, not QL chart-reading. You interpret histograms, line graphs, pie charts, ogives, and box-and-whisker plots; compute and compare mean, median, mode, and standard deviation; trace tree and Venn diagrams; and judge how outliers move the measures. Calculators are not permitted.
MAT stats versus QL data literacy
Quantitative Literacy on the AQL asks you to read a campus table, a fee pie, or a water-use line graph and extract a total or a percentage in context. The MAT paper, under the 2015 booklet heading 4. DATA HANDLING and PROBABILITY and the current Test Content line interpret various representations and measures of data, goes further:
- Measurement: mean, median, mode, range, quartiles, interquartile range, standard deviation — and what happens to each when a value is added, removed, or rescaled.
- Representation: histograms, line graphs, pie charts, ogives (cumulative-frequency curves), box-and-whisker plots.
- Probability: listing, product along a tree, union and intersection in a Venn diagram, complementary events, independence versus mutual exclusion.
The 2015 achievement levels spell the density. Basic: identify a central-tendency measure and read a simple graph or table. Intermediate: use variability, tree and Venn diagrams, probability rules, and standard deviation. Proficient: several representations at once, outliers pulling the mean, and multi-step probability. That is why this section is denser than QL data literacy: you still read the display, then you compute a measure or a probability from it without a calculator.
Measures without a calculator
Mean, median, mode
For a list, the mean is the sum divided by n. The median is the middle value once the list is ordered (the average of the two middle values if n is even). The mode is the most frequent value (a list may be unimodal, bimodal, or have no repeating value).
Worked list: 4, 6, 6, 7, 9, 10, 18.
- Sum = 60, n = 7, mean = 60/7 = 8 4/7 (leave the mixed number; do not chase a decimal expansion).
- The list is already ordered; median = 7.
- Mode = 6.
The 18 sits far to the right. It lifts the mean above the median. That is the signature of positive skew: mean > median. A low outlier (negative skew) pulls the mean below the median. MAT often asks which measure is least affected by the outlier: the median (and the mode), not the mean.
If every value is increased by 3, the mean and median each rise by 3 and the standard deviation is unchanged. If every value is multiplied by 2, mean and median double and the standard deviation is multiplied by 2 (absolute value of the scale factor).
Five MAT practice scores have mean 12, so the total is 60. Remove an outlier of 20: the remaining total is 40, and the mean of the four remaining scores is 10. That one-line total-minus-outlier move is the MAT version of “impact of an outlier,” and it needs no calculator.
Standard deviation
Standard deviation measures spread about the mean. A tight cluster has a small SD; two clumps far from the mean have a large SD. You will rarely be asked to grind a messy square root by hand. You will be asked which of four tiny data sets has the largest SD, what happens when an outlier is inserted, and that a constant data set such as 9, 9, 9, 9 has SD 0.
Mini calculation that stays exact: data 1, 5. Mean = 3. Deviations −2 and +2. Squared deviations 4 and 4. Using the population form (divide by n), variance = 8/2 = 4, SD = 2. If a stem used the sample form (divide by n − 1), variance = 8/1 = 8, SD = 2√2. MAT numbers are chosen so the intended division is obvious, or the question is qualitative. Do not invent a calculator-needed square root.
Outliers and the box-and-whisker
A five-number summary is minimum, Q1, median, Q3, maximum. The IQR = Q3 − Q1. A common fence in school box-plot work is:
- lower fence: Q1 − 1.5 × IQR
- upper fence: Q3 + 1.5 × IQR
A point beyond a fence is an outlier; the whisker stops at the last non-outlier.
Worked: ordered marks 2, 4, 5, 6, 7, 8, 9, 20, n = 8.
Lower half 2, 4, 5, 6 → Q1 = (4+5)/2 = 4.5. Upper half 7, 8, 9, 20 → Q3 = (8+9)/2 = 8.5. Median = (6+7)/2 = 6.5. IQR = 4. Upper fence = 8.5 + 1.5×4 = 8.5 + 6 = 14.5. The 20 is an outlier. The right whisker ends at 9, not at 20.
Removing 20 drops the mean sharply and the SD. The median of the remaining seven values is 6 — a much smaller shift than the mean’s. That is the Proficient-level “impact of outliers on measures of central tendency and variability” skill in the 2015 table.
Representations
| Display | What it is | MAT trap |
|---|---|---|
| Histogram | Adjacent bars; area tracks frequency on continuous class intervals | Treating it like a bar chart of named categories, or ignoring unequal class width |
| Line graph | Ordered x (often time) joined by segments | Reading a steep segment as a large percentage change when the base is small |
| Pie chart | Parts of one whole; 360° = 100% | Comparing pies with different totals as if slices were raw counts |
| Ogive | Cumulative frequency against the upper class boundary | Reading a height as a class frequency instead of a running total |
| Box-and-whisker | Five-number summary plus possible outlier dots | Assuming the box is symmetric, or using the maximum as Q3 |
Ogive — median class
40 candidates. Cumulative frequencies at upper bounds 10, 20, 30, 40, 50 are 5, 14, 26, 34, 40. Class frequencies are therefore 5, 9, 12, 8, 6.
The median is the 20th ordered mark (n/2). The cumulative frequency first exceeds 20 after 14 at the bound 20, in the class whose upper bound is 30. So the median lies in 20 ≤ x < 30.
Q1 is the 10th value (class 10–20). Q3 is the 30th value (class 30–40). Sketch the ogive as a rising polygonal line and read across from 20 on the cumulative-frequency axis. The arithmetic is comparing 20 with 5, 14, 26 — no calculator.
Histogram versus pie
A histogram with unequal class width is not read by bar height alone if the paper used frequency density. Many school histograms keep equal width, so height is proportional to frequency; check the axis label. A pie of “preferred residence” with slices 90°, 120°, 150° is 25%, 33⅓%, 41⅔% of that survey only. A line graph of weekly practice scores is for ordered change, not for composition of a whole.
Probability: trees and Venn diagrams
Venn (two sets)
In a tutorial group of 30, 18 take extra algebra work, 16 take extra geometry, and 9 take both.
- Only algebra: 18 − 9 = 9
- Only geometry: 16 − 9 = 7
- Neither: 30 − (9 + 9 + 7) = 5
- Algebra or geometry: 18 + 16 − 9 = 25
P(both | takes algebra) = 9/18 = 1/2. P(exactly one) = 16/30 = 8/15. P(geometry but not algebra) = 7/30.
The addition rule is P(A ∪ B) = P(A) + P(B) − P(A ∩ B). If A and B are mutually exclusive, the intersection is empty and you drop the last term. Independent events satisfy P(A ∩ B) = P(A)P(B); that is a product rule, not a Venn picture of “no overlap.” Overlap and independence are different ideas — MAT exploits that confusion.
Tree (sequential, no replacement)
A box holds 4 blue cards and 2 red cards. Two cards are drawn without replacement.
P(both blue) = (4/6)×(3/5) = (2/3)×(3/5) = 2/5.
P(different colours) = P(blue then red) + P(red then blue) = (4/6)(2/5) + (2/6)(4/5) = 8/30 + 8/30 = 16/30 = 8/15.
P(at least one red) = 1 − P(both blue) = 1 − 2/5 = 3/5. The complement is usually faster than listing three favourable paths.
With replacement, second-branch denominators would stay 6. Read the stem. A trap value 4/9 is (2/3)×(2/3), which pretends the first blue card was replaced.
Independence check
A fair coin and a fair die: P(Heads and 6) = (1/2)×(1/6) = 1/12. The events do not interfere. Drawing cards without replacement does interfere: P(second blue | first blue) = 3/5, not 4/6.
Method on the day
Read the title, units, and whether a graph is cumulative. For a measure, order the data before you hunt the median. For an outlier question, compute IQR fences or at least compare mean versus median. For probability, decide tree (order) versus Venn (sets) before arithmetic. All of it is pencil work on the MAT paper.
Forty candidates have cumulative frequencies 5, 14, 26, 34, 40 at the upper class bounds 10, 20, 30, 40, 50. In which class does the median lie?
A box holds 4 blue cards and 2 red cards. Two cards are drawn at random without replacement. What is the probability that both cards are blue?
In a group of 30 students, 18 take extra algebra, 16 take extra geometry, and 9 take both. What is the probability that a student chosen at random takes geometry but not algebra?