11.1 Data Representations: Frequency Tables, Histograms, Stem-and-Leaf, Dot Plots & Box Plots
Key Takeaways
Categorical data (nominal, ordinal) is represented using frequency tables and bar charts with distinct spaces between bars, whereas numerical data (discrete, continuous) requires continuous scales such as histograms, dot plots, and stem-and-leaf plots.
Histograms group continuous quantitative observations into equal-width, non-overlapping intervals (bins) where touching adjacent bars reflect continuous intervals across the domain.
Stem-and-leaf plots preserve individual raw values while displaying overall distribution shape, requiring an explicit key to define place values and split stems when observations cluster tightly.
Dot plots display individual discrete observations as points along a continuous number line, making clusters, gaps, peaks, and potential outliers immediately discernible for smaller data sets.
Box plots summarize numerical data via the five-number summary (Minimum, Q1, Median, Q3, Maximum), divide distributions into four equal-count quartiles (25% each), and identify outliers via Tukey's 1.5 × IQR rule.
11.1 Data Representations: Frequency Tables, Histograms, Stem-and-Leaf, Dot Plots & Box Plots
Statistical literacy in middle-grades mathematics begins with the ability to collect, organize, and transform raw observations into structured visual representations. Data displays do not merely present numbers; they reveal the underlying shape, center, variability, and anomalies of empirical distributions. A central responsibility of mathematics educators in grades 4–8 is developing students' conceptual understanding of when and why specific graphical displays are chosen, how to construct them with mathematical fidelity, and how to extract valid inferences from their structural properties.
Categorical vs. Numerical Data Types
Before selecting or constructing any graphical display, an analyst must first identify the fundamental nature of the variable under investigation. Confusing categorical attributes with numerical measurements is one of the most persistent conceptual hurdles in early statistical learning.
1. Categorical (Qualitative) Data
Categorical variables represent characteristics, labels, or qualities that place individual entities into distinct groups or categories. Categorical data is divided into two distinct levels of measurement:
- Nominal Data: Unordered, mutually exclusive categories where numerical values serve purely as labels with no inherent hierarchy or quantitative meaning. Examples include favorite sports, bus route numbers, hair colors, or student ID numbers. Arithmetic operations (such as calculating a mean) on nominal labels are completely meaningless.
- Ordinal Data: Qualitative categories that possess a natural, sequential order or rank, yet the differences (intervals) between ranks are neither uniform nor mathematically quantifiable. Examples include Likert satisfaction scales (e.g., Strongly Disagree, Disagree, Neutral, Agree, Strongly Agree), competition finishes (1st, 2nd, 3rd place), or t-shirt sizes (Small, Medium, Large, Extra Large). While we know that 1st place finished ahead of 2nd place, the time interval separating them is not fixed or identical to the interval separating 2nd and 3rd.
2. Numerical (Quantitative) Data
Numerical variables represent measurable or countable quantities for which standard arithmetic operations (addition, subtraction, multiplication, averaging) are mathematically valid. Quantitative data is categorized as follows:
- Discrete Data: Numerical values resulting from a counting process, existing only as isolated, distinct points along a number line. Values typically comprise non-negative integers. Examples include the number of books checked out of a library, the number of siblings a student has, or the number of questions answered correctly on a test. One cannot have 2.37 siblings.
- Continuous Data: Numerical values resulting from a continuous measurement process along an uninterrupted continuum. Between any two continuous values, an infinite number of intermediate values theoretically exist. The precision of continuous data is restricted solely by the sensitivity of the measuring instrument. Examples include student sprint times in seconds, mass in grams, temperature in degrees Celsius, or heights in centimeters.
Frequency Distributions and Relative Frequency Tables
A frequency distribution is a structured tabular summary that organizes raw data into mutually exclusive classes or categories, displaying the count of observations (absolute frequency, ) falling within each class. To convert raw observations into a frequency table:
- Identify the range of the data: .
- For numerical data, determine appropriate, non-overlapping intervals (bins) of equal width.
- Tally the number of observations within each interval.
- Calculate the relative frequency, which expresses each class count as a proportion or percentage of the total sample size ():
- Optionally compute the cumulative frequency and cumulative relative frequency by maintaining a running sum of frequencies from the lowest class up through the current class.
Step-by-Step Frequency Table Construction
Consider the raw scores of 20 middle school students on a 50-point mathematics challenge:
With , minimum value 24, and maximum value 50, we establish uniform bin widths of 5 units, choosing left-inclusive intervals :
| Score Interval (Bin) | Tally | Absolute Frequency () | Relative Frequency () | Percentage (%) | Cumulative Frequency | Cumulative Relative Frequency |
|---|---|---|---|---|---|---|
| I | 1 | 5% | 1 | 0.05 (5%) | ||
| I | 1 | 5% | 2 | 0.10 (10%) | ||
| II | 2 | 10% | 4 | 0.20 (20%) | ||
| IIIII | 5 | 25% | 9 | 0.45 (45%) | ||
| IIIII | 5 | 25% | 14 | 0.70 (70%) | ||
| IIIII I | 6 | 30% | 20 | 1.00 (100%) | ||
| Total | 20 | 1.00 | 100% | — | — |
Sum Check Rule: The sum of absolute frequencies must equal the sample size (), and the sum of relative frequencies must equal exactly ().
Bar Charts vs. Histograms: The Critical Distinction
A universal pedagogical misconception among middle school students is treating histograms and bar charts as interchangeable. Educators must reinforce the mathematical and structural distinctions between them:
| Structural Feature | Bar Chart (Bar Graph) | Histogram |
|---|---|---|
| Data Type | Categorical (Nominal or Ordinal) | Numerical (Continuous or Grouped Discrete) |
| Horizontal Axis (-axis) | Distinct, non-numeric category names or qualitative labels | Continuous numerical scale partitioned into equal-width bins |
| Bar Spacing | Distinct spaces/gaps between adjacent bars to emphasize separate categories | No spaces/gaps between adjacent bars; bars touch to signify continuous intervals |
| Meaning of Gaps | Gaps represent visual separation between distinct qualitative categories | A gap indicates an interval with an absolute frequency of zero () |
| Bar Ordering | Nominal bars can be rearranged in any order (e.g., alphabetically or Pareto order) | Numerical bin order is strictly fixed along the real number line |
| Bar Width | Width is arbitrary and decorative; only height conveys meaning | Width represents the interval span (bin width); area is proportional to frequency |
Boundary Rules in Histograms
When continuous observations land exactly on a bin boundary, a standardized boundary convention must be applied across the entire data set. The standard convention is left-inclusive and right-exclusive (), meaning the lower bound is included in the bin, but the upper bound is deferred to the next adjacent bin. For example, in the interval , a score of exactly belongs in that bin, but a score of belongs in the bin. The only exception occurs in the final bin, which is closed on both ends to capture the maximum value ().
Stem-and-Leaf Plots: Preserving Individual Data Values
A stem-and-leaf plot (stemplot) is an exploratory display that organizes quantitative data while retaining every individual raw data value. It partitions each numerical observation into two components:
- Stem: The leading digit or digits representing higher place values.
- Leaf: The single trailing digit representing the lowest place value.
Construction Rules
- Stems are written vertically in a column in ascending order from top to bottom, followed by a vertical dividing line.
- Leaves are written horizontally to the right of their corresponding stem in strictly ascending order from left to right.
- Every leaf must consist of exactly one single-digit integer ( through ). If a raw value contains more than one trailing digit, the data must be rounded or truncated before plotting.
- Duplicate values must each have their own individual leaf recorded (e.g., two scores of 38 produce two leaves of 8).
- The Essential Key/Legend: A stem-and-leaf plot is mathematically uninterpretable without an explicit key. For example, the stem-leaf pair could represent , , , or . The key establishes the place value context:
Split Stems for Clustered Data
When data values cluster densely within a narrow numerical range, a standard stemplot yields very few stems with overwhelmingly long rows of leaves, obscuring the underlying distribution shape. To resolve this, educators introduce split stems:
- Each stem digit is repeated twice: the first stem holds leaves , and the second stem holds leaves .
- Alternatively, stems can be split five times (leaves , , , , ).
Comparative Back-to-Back Stem-and-Leaf Plots
A back-to-back stem-and-leaf plot shares a single central column of stems to compare two related data sets simultaneously (such as Class A vs. Class B):
- The leaves for the group on the right (Class B) are recorded in conventional ascending order from left to right away from the center.
- The leaves for the group on the left (Class A) are recorded in ascending order from right to left away from the center.
Notice that for Class A, on stem 3, the leaves reading outward from the stem to the left are , representing the scores .
Dot Plots (Line Plots): Granular Displays for Small Data Sets
A dot plot (historically called a line plot in early elementary curricula) displays individual discrete observations as dots or points placed vertically above a continuous horizontal number line.
Characteristics and Pedagogical Value
- Granular Precision: Every single observation in the data set is represented by a physical dot, making sample size () immediately countable by summing the dots.
- Direct Shape Identification: The vertical stacking of dots produces an immediate visual histogram-like profile, revealing peaks (modes), clusters (concentrations of data), gaps (intervals with no observations), and apparent extreme values (outliers).
- Scale Fidelity: The horizontal axis is an authentic, continuous number line with uniform intervals. Numbers with zero frequency must still be labeled to preserve geometric proportion.
- Limitation: Dot plots become visually cluttered, unwieldy, and difficult to interpret when sample sizes exceed to observations or when data spans thousands of distinct continuous values.
Box Plots (Box-and-Whisker Plots) & The Five-Number Summary
A box plot is an exploratory graphical display that summarizes a quantitative distribution based on positional ranks rather than raw counts. It visualizes the Five-Number Summary:
- Minimum (): The lowest value in the data set. (In a modified box plot, the lower whisker stops at the lowest value that is not an outlier.)
- First Quartile (): The 25th percentile; 25% of the data values lie at or below .
- Second Quartile / Median (): The 50th percentile; the physical middle value dividing the distribution into two equal halves.
- Third Quartile (): The 75th percentile; 75% of the data values lie at or below (and 25% lie above it).
- Maximum (): The highest value in the data set. (In a modified box plot, the upper whisker stops at the highest value that is not an outlier.)
Calculating Quartiles: The Standard Middle School Method
In grades 4–8 mathematics (the Moore and McCabe convention), quartiles are calculated through a straightforward algorithmic sequence:
- Arrange the observations in ascending numerical order.
- Locate the median ():
- If is odd, the median is the single middle observation at position .
- If is even, the median is the arithmetic mean of the two central observations at positions and .
- Partition the ordered data set into a lower half and an upper half:
- If is odd, strictly exclude the median from both halves.
- If is even, split the data set cleanly at the midpoint into two equal halves.
- Calculate as the median of the lower half of the data.
- Calculate as the median of the upper half of the data.
The Interquartile Range ()
The Interquartile Range quantifies the spread of the central 50% of the distribution:
Because relies entirely on the middle half of the ranked data, it is a resistant statistic that remains completely unaffected by extreme values or outliers in the tails.
Identifying Outliers Using Tukey's Rule
John Tukey established a mathematically rigorous criterion for identifying potential outliers using fences constructed around the central box:
- Lower Inner Fence:
- Upper Inner Fence:
Outlier Decision Rule:
- Any observation strictly less than the lower fence () is classified as a lower outlier.
- Any observation strictly greater than the upper fence () is classified as an upper outlier.
Modified Box Plots
In a standard box plot, whiskers extend all the way to the absolute minimum and maximum values. In a modified box plot (the preferred standard in modern statistics education):
- Whiskers extend outward from the box only to the most extreme data points that fall inside the fences (the adjacent values).
- Any data point falling outside the fences is plotted individually as an isolated dot, asterisk (), or open circle.
The Four Quartile Regions
A box plot partitions any numerical distribution into four distinct spatial zones:
- Lower Whisker: From Minimum to (contains approximately 25% of the data).
- Lower Box: From to Median () (contains approximately 25% of the data).
- Upper Box: From Median () to (contains approximately 25% of the data).
- Upper Whisker: From to Maximum (contains approximately 25% of the data).
Critical Understanding: A longer whisker or box section does not contain more data points; it signifies that the 25% of observations in that region are more widely dispersed across the number line (lower density).
Circle Graphs (Pie Charts)
The framework lists pie charts among the displays teachers must organize and interpret. A circle graph shows how a whole divides into categories. Each sector's central angle is proportional to its relative frequency:
In a class survey of 40 students, 14 who chose soccer get a sector of . Circle graphs work only when the categories are mutually exclusive and together make up 100% of the whole. They are a poor choice for comparing groups of different sizes or for showing change over time.
Summary of Graphical Displays: Strengths and Limitations
| Display Type | Data Type | Primary Strengths | Limitations | Optimal Research Question |
|---|---|---|---|---|
| Bar Chart | Categorical | Clear visual comparison of discrete category frequencies; easy to interpret. | Cannot show continuous distribution shape or numerical spread. | "Which category has the highest or lowest frequency?" |
| Histogram | Continuous Numerical | Summarizes large data sets into intervals; clearly reveals distribution shape, center, and skewness. | Obscures individual raw data values; appearance varies based on chosen bin width. | "What is the overall shape and spread of scores across uniform intervals?" |
| Stem-and-Leaf | Quantitative (Discrete/Continuous) | Retains exact individual numerical values; displays distribution shape like a histogram. | Unusable for massive data sets (); requires clear key to avoid ambiguity. | "What are the exact scores of all participants while comparing distribution shapes?" |
| Dot Plot | Discrete Numerical | Retains all individual data points; highlights clusters, gaps, repeated values, and modes. | Cumbersome for large sample sizes () or highly continuous decimals. | "Where do individual values cluster, and where are the gaps in a small class sample?" |
| Box Plot | Continuous Numerical | Excellent for comparing multiple distributions side-by-side; clearly identifies outliers via five-number summary. | Masks distribution details such as bimodal peaks, clusters, and individual data points. | "How do the medians, spreads (IQR), and outliers of two or more groups compare?" |
Worked Step-by-Step Mathematical Examples
Worked Example 1: Calculating Five-Number Summary, Outlier Fences, and Modified Box Plot Bounds
Problem: A physical education teacher records the 100-meter sprint times (in seconds) for a group of 15 middle school students:
Determine the five-number summary, calculate the interquartile range (), identify any mathematical outliers using Tukey's rule, and specify the endpoints of the whiskers for a modified box plot.
Solution:
-
Step 1: Order data and confirm sample size: The data set is already sorted in ascending order with observations.
-
Step 2: Find the median (): Since is odd, the median is the single middle observation at position :
- Step 3: Determine the lower and upper halves: Excluding the median (), the lower half contains the 7 values below position 8:
The upper half contains the 7 values above position 8:
-
Step 4: Find and : For the lower half ( values), the median is the 4th value: . For the upper half ( values), the median is the 4th value: .
-
Step 5: Compute the Interquartile Range ():
- Step 6: Compute Tukey's Outlier Fences:
-
Step 7: Identify Outliers:
-
Minimum value is . Since , there are no lower outliers.
-
Maximum value is . Since , the observation is a confirmed upper outlier.
-
Step 8: Specify Modified Box Plot Whiskers:
-
Lower whisker extends from () down to the absolute minimum: .
-
Upper whisker extends from () up to the largest non-outlying observation: .
-
The value is plotted as an isolated point at on the number line.
-
Five-Number Summary reported: ; in the modified box plot the upper whisker stops at and is plotted as an outlier.
Worked Example 2: Interpreting a Comparative Back-to-Back Stem-and-Leaf Plot
Problem: A mathematics teacher analyzes the test scores from two sections using the back-to-back stem-and-leaf plot below ():
Determine the median score and range for both sections, and provide a comparative analysis of student achievement.
Solution:
- Step 1: Reconstruct raw scores for Section A in ascending order: Remember that Section A leaves read outward to the left (from stem outward):
- Stem 5:
- Stem 6:
- Stem 7:
- Stem 8:
- Stem 9: Total count . The 8th value is the median: .
- Step 2: Reconstruct raw scores for Section B in ascending order: Section B leaves read standard left to right away from the stem:
- Stem 5:
- Stem 6:
- Stem 7:
- Stem 8:
- Stem 9: Total count . The 8th value is the median: .
- Step 3: Comparative Analysis: Section B demonstrates slightly higher overall performance, with a median score () exceeding Section A's median (). Furthermore, Section B has 7 students scoring in the 80s and 90s combined (), whereas Section A has 6 students scoring in the 80s and 90s (). Both classes exhibit similar overall dispersion, with ranges of 40 and 45 points respectively.
Diagnostic Misconceptions & Pedagogical Strategies
- The "Length Equals Count" Box Plot Fallacy: Students frequently assume that a longer box section or whisker represents "more data points." Pedagogical Strategy: Use a dot plot positioned directly beneath a box plot on the same number line. Have students count the number of dots inside each quartile segment. Explicitly reinforce that each segment contains exactly 25% of the data; a longer segment indicates lower density (values are spaced further apart), while a compressed segment indicates high density (values cluster tightly).
- Conflating Histograms with Bar Graphs: Students routinely draw histograms with spaces between bars or attempt to construct histograms for nominal categories. Pedagogical Strategy: Require students to label the horizontal axis of histograms at the boundary edges of the bins rather than centering labels beneath bars. Emphasize that because the horizontal axis is a continuous number line, bars must touch to represent contiguous numerical intervals.
- Omitting the Key in Stem-and-Leaf Displays: Students often create stem-and-leaf plots without a key, rendering the display mathematically ambiguous. Pedagogical Strategy: Present students with an unkeyed display showing and challenge them to guess whether it represents 12 years old, meters, pounds, or 12 cents. This diagnostic exercise demonstrates the absolute necessity of place value keys.
An educator presents students with the ordered data set representing the number of books read during a summer reading challenge: 4, 7, 8, 9, 12, 14, 15, 17, 19, 20, 22, 24, 25, 28, 48. Using the standard middle-school quartile convention (excluding the median) and Tukey's 1.5 × IQR rule, what is the interquartile range (IQR), and which value, if any, is classified as an outlier?
IQR = 13; there are no outliers because 48 is within two standard deviations of the sample mean.
IQR = 15; the value 48 is an outlier because it exceeds the upper fence of 46.5.
IQR = 15; both 4 and 48 are outliers because they lie beyond the lower fence of 5 and upper fence of 45.
IQR = 24; the value 48 is an outlier because the maximum observation is more than twice the third quartile.
A teacher displays a back-to-back stem-and-leaf plot comparing the test scores of two class sections. Several students conclude that Class A must have scored lower overall than Class B because the leaves on the left side (Class A) read backward from right to left. What structural characteristic of stem-and-leaf plots should the teacher clarify to address this confusion?
In a back-to-back stem-and-leaf plot, the shared stem occupies the central column, and leaves for the group on the left must be read outward from the stem toward the left, preserving ascending numerical order away from the center.
Stem-and-leaf plots should only be read from left to right regardless of layout; reading backward indicates that Class A's scores are negative quantities.
Leaves in a stem-and-leaf plot represent frequencies rather than actual data values, so directionality has no relationship to individual student scores.
Class A's values cannot be determined without first converting the stem-and-leaf display into a categorical bar chart with equal intervals.
Which of the following visual representations is most appropriate for displaying a data set of 120 continuous measurements of student sprint times (in seconds) to analyze the frequency distribution across uniform time intervals, and what distinguishes this graphical format from a standard bar chart?
A bar chart; adjacent bars touch to show that the student categories form a single continuous athletic group.
A dot plot; individual points are stacked at exact tenths of a second to prevent overlap across large data sets.
A histogram; adjacent bars touch without gaps because the horizontal axis represents contiguous numerical intervals rather than discrete categorical labels.
A box plot; individual bars represent the five-number summary intervals with spaces between the quartiles.
Sections you finish are checked off in the contents.