8.1 Data Representation and Analysis

Key Takeaways

  • NBT QL data items ask you to decode a table or chart and then interpret it in a higher-education context, not only to read an axis label
  • Mixed tables combine counts, rates, and categories; the largest occupied-bed count is not automatically the highest occupancy rate
  • Pie charts show parts of one whole; grouped bars compare side by side; stacked bars show composition; broken lines show change over time; scatter plots show association
  • The mean follows outliers, the median is the middle value, and the mode is the most frequent value; choose the measure that matches the story
  • Tree branches must add back to their parent total; a leftover that does not match the title means you are reading a subset or an incomplete split
Last updated: September 2026

Decode the display, then interpret it

The Quantitative Literacy (QL) component of South Africa's National Benchmark Tests sits inside the Academic and Quantitative Literacy (AQL) paper. It is not a second Mathematics (MAT) paper. The National Benchmark Tests Project, through the Centre for Educational Assessments, describes data representation and analysis as the ability to derive information from contextualised tables, charts and diagrams and then to interpret what that information means. Independent OpenExamPrep study material for NBT QL trains that as a two-step habit: decode the display, then say what the numbers imply for a real campus decision.

QL items use tables with several rows and columns and mixed data types. A single table may mix whole-number counts, percentages, categories, and time labels. The trap is to treat every column as if it answered the same question.

Consider a fictional first-year residence snapshot at Marula University, a public university in Limpopo, at the start of the academic year.

ResidenceBedsOccupied week 1Occupied week 6NSFAS-fundedSelf-fundedOther fundingMean wait (days)
Magoleng24022823617454811
Baobab180162179121411718
Mopani32028830124048139
Ivory9071882249174

Week 1 occupancy rates are 228/240 = 95% at Magoleng, 162/180 = 90% at Baobab, 288/320 = 90% at Mopani, and 71/90 ≈ 79% at Ivory. Mopani housed the most students (288), but Magoleng filled the largest share of its beds. Ivory looks "empty" if you only count vacant beds (19), yet Mopani had 32 vacant beds — a larger count of empty rooms and still a higher fill rate, because Mopani is much bigger. On QL, "highest occupancy" is incomplete until you know whether the setter wants a count or a rate.

Check that mixed columns are consistent. NSFAS-funded + self-funded + other funding at Magoleng is 174 + 54 + 8 = 236, which matches week 6 occupancy, not week 1 and not the bed total. If a question asks for the NSFAS share of Magoleng residents in week 6, the denominator is 236 occupied beds, giving 174/236 ≈ 73.7%, not 174/240. Using the design capacity as the base quietly changes the story.

Pie charts show parts of one whole. A pie of Magoleng's week 6 funding mix is appropriate because NSFAS, self-funded, and other are mutually exclusive categories that add to the occupied total. A pie is a poor choice for comparing Magoleng with Ivory, or for showing change from week 1 to week 6: those are not slices of a single whole. If someone draws two pies, one per week, you still cannot read a trend off a slice angle without converting back to counts or rates. QL often asks whether a pie is even a fair display, not only which slice is largest.

Simple bar charts compare separate categories with a common unit. Grouped (compound) bar charts place two or more bars side by side for each category — for example NSFAS versus self-funded counts at each residence — so you can compare funding mix across residences without forcing the bars to add to a total you care about. Stacked bars pile segments that do add to a total: NSFAS, self-funded, and other stacked to week 6 occupancy. Read a stacked bar from the baseline for the bottom segment only. The middle segment's length is its own count; its starting height is not its value. Candidates who read a stacked mid-segment as if it started at zero invent a number the chart never drew.

A broken-line graph joins successive time points. Marula's campus clinic logged influenza-like visits in the first eight teaching weeks as 18, 22, 31, 40, 38, 27, 21, 19. The line rises to week 4, then falls. "Broken" here means the graph is a polyline through those points, not a smooth curve you should interpolate as if visits were continuous between Wednesdays. You may estimate a midpoint only if the question asks for an approximate reading; you may not invent a week-3.5 value and treat it as measured. The interpretation is seasonal: visits peaked after registration week, then declined — not "the clinic is steadily busier."

Scatter plots show paired measurements for the same students. A first-year chemistry practical at Marula plotted tutorial sessions attended (horizontal) against practical mark out of 100 (vertical). Most points rise from left to right: students who attended more tutorials tended to score higher. That is an association, not a proof that tutorials cause the marks. One student attended 11 of 12 tutorials and scored 38. That point is an outlier relative to the cloud. It does not erase the overall positive pattern, and it does not prove that tutorials harm performance. On QL, name the outlier, then still describe the main trend.

Tree diagrams split a group into successive branches. Marula's 400 first-years split by faculty (Health Sciences 120, Education 80, Commerce 200), then each faculty splits into campus-shuttle users versus walk-or-taxi. Every child branch must add back to its parent: 48 shuttle + 72 walk = 120 Health students. If a tree's last row does not sum to the title total, the diagram is incomplete or you are reading a subset. Trees on QL are often classification devices, not MAT-style probability engines.

Mean, median, and mode are tools for a story, not decorations. Seven practical marks were 48, 51, 53, 55, 56, 58 and 94. The mean is 415 ÷ 7 ≈ 59.3. The median, the middle value once ordered, is 55. There is no mode; every mark appears once. The 94 is an outlier. Remove it and the mean becomes 321 ÷ 6 = 53.5 (a drop of about 5.8 marks) while the median becomes (53 + 55) / 2 = 54 (a drop of 1). The mean followed the extreme score; the median barely moved. A second example: ten weekly library no-show penalties of 0, 0, 1, 1, 1, 1, 2, 3, 3, 12. The mode is 1 (most frequent), the median is 1, and the mean is 24/10 = 2.4. Reporting only the mean would make no-shows sound routine; the mode and median say the typical week is a single penalty, with one extraordinary week of 12.

National reports treat Quantitative Literacy as a demanding domain: the 2025-intake National Benchmark Tests report listed a full-cohort QL mean near 49 percent. That figure is overall QL, not a published per-subdomain cut-score. Treat data items as interpretation under time pressure, not giveaways. Work without a calculator: cancel factors, estimate, and compare rather than chasing extra decimals. Read the title, units, and footnotes before you touch an option. Then interpret: what would a residence manager, a clinic nurse, or a first-year actually do with this number?

Loading diagram...
QL habit: decode a display, then interpret it
Marula University week 1 occupancy rate by residence
Test Your Knowledge

Using the Marula University residence table, week 1 occupied beds were Magoleng 228 of 240, Baobab 162 of 180, Mopani 288 of 320, and Ivory 71 of 90. Which residence had the highest occupancy rate in week 1?

A
B
C
D
Test Your Knowledge

A chemistry practical set has marks 48, 51, 53, 55, 56, 58 and 94. What happens if the 94 is treated as an outlier and removed?

A
B
C
D
Test Your Knowledge

On a scatter plot of tutorial sessions attended versus practical mark, most points rise from left to right, but one student attended 11 tutorials and scored 38. The most defensible reading is:

A
B
C
D