11.1 Statistics: Data Displays, Measures & Investigations
Key Takeaways
Bar graphs and pictographs display categorical data, while dot plots, histograms, stem-and-leaf plots, and box plots display numerical data.
The median resists outliers better than the mean, so it often better represents skewed data.
Range and interquartile range describe how spread out the data are.
In a normal distribution, about 68% of values fall within one standard deviation of the mean and about 95% within two.
A random sample represents a population, while a biased sample can produce misleading conclusions.
Overview & Exam Relevance
Competency 005 (Probability and Statistics) of the TExES Core Subjects EC-6 (391) Mathematics subject exam (902) covers how teachers guide students through statistical investigation: collecting, organizing, displaying, and interpreting data. This section focuses on statistics; probability and the mathematical processes of Competency 006 follow in the next sections. In modern Texas elementary curricula, statistics and probability are not isolated arithmetic procedures; they provide the empirical tools through which students collect, organize, represent, interpret, and draw evidence-based conclusions from data.
Candidates must select appropriate graphical displays for categorical and numerical data, compute and compare measures of center and spread for different distribution shapes, recognize when the normal distribution applies, and design simple statistical investigations that answer real questions.
Graphical Representations of Data (Competency 005)
Selecting an effective data display requires distinguishing between categorical (qualitative) data (non-numeric labels, such as eye color or vehicle type) and numerical (quantitative) data (measurable or countable quantities, such as height, test scores, or temperature).
DATA REPRESENTATION DECISION FRAMEWORK
│
├── Categorical Data (Labels / Groups)
│ ├── Pictograph (Scaled key/legend; early elementary)
│ ├── Bar Graph (Discrete separated bars; single & double comparisons)
│ └── Circle Graph / Pie Chart (Part-to-whole percentage proportions)
│
└── Numerical Data (Continuous or Countable Measurements)
├── Dot Plot / Line Plot (Number line with dots; small datasets n < 30)
├── Stem-and-Leaf Plot (Maintains raw values; stem = tens, leaf = ones)
├── Histogram (Continuous bins/intervals; touching bars; no gaps)
├── Box Plot (5-Number Summary: Min, Q1, Median, Q3, Max; shows IQR)
└── Scatter Plot (Bivariate pairs; analyzes correlation/association)
Comprehensive Comparison of Graphical Displays
| Graphical Display | Data Type | Structural Features | Optimal Pedagogical Purpose | Common Student Misconception |
|---|---|---|---|---|
| Pictograph | Categorical | Icons represent fixed quantities; requires a clear key/legend | Introducing data representation in Grades K–3; visual counting | Ignoring the key (e.g., counting 4 icons as 4 items when each icon represents 5 items). |
| Bar Graph | Categorical | Bars separated by distinct spaces; vertical or horizontal orientation | Comparing frequencies across distinct, non-overlapping categories | Confusing bar graphs with histograms; failing to read fractional bar heights accurately. |
| Dot Plot (Line Plot) | Numerical | Individual dots or X's stacked above a continuous number line | Visualizing distributions, clusters, peaks (mode), gaps, and outliers for small sets () | Placing dots on numbers that have zero frequency rather than leaving the number blank. |
| Stem-and-Leaf Plot | Numerical | Stems display leading place values; leaves display trailing digits; requires a key | Organizing data sequentially while preserving every individual raw data value | Omitting the key; failing to list repeated leaf values (e.g., writing 5 once for scores of 45 and 45). |
| Histogram | Continuous Numerical | Bars touch without spaces (spaces only occur if an interval has zero frequency); equal interval bins | Displaying frequency distributions for large numerical datasets | Mistaking histograms for bar graphs; assuming bars touch arbitrarily rather than representing continuous bins. |
| Circle Graph (Pie Chart) | Categorical | Slices represent proportional percentages of a whole ( or ) | Displaying relative part-to-whole proportions | Believing larger circle graphs necessarily represent larger total quantities than smaller ones. |
| Box Plot (Box-and-Whisker) | Numerical | Displays the Five-Number Summary (Min, , Median, , Max); box width represents IQR | Visualizing median, spread, quartiles (each representing of data), and skewness | Assuming the wider whisker or box section contains more data points (each quartile contains exactly ). |
| Scatter Plot | Bivariate Numerical | Coordinate points plotted on Cartesian axes | Analyzing association/correlation (positive, negative, none; linear vs non-linear) | Conflating correlation with direct causation. |
Measures of Central Tendency & Variability
Statistical summary values describe the center and dispersion of numerical datasets.
Measures of Central Tendency
- Mean (Arithmetic Average): The sum of all observations divided by the total number of observations (). Conceptually modeled as the "fair share" (leveling stacks of cubes) or the physical "balance point" of a data distribution on a fulcrum.
- Median (Physical Center): The middle observation when data are arranged in numerical order. If the sample size is odd, it is the exact middle value; if is even, it is the arithmetic average of the two middle values.
- Mode (Most Frequent): The data value that occurs with the highest frequency. A dataset may be unimodal (one mode), bimodal (two modes), multimodal, or have no mode. The mode is the only measure of central tendency applicable to nominal categorical data (e.g., modal favorite ice cream flavor).
Measures of Variability (Spread)
- Range: The simplest measure of dispersion: .
- Interquartile Range (IQR): The difference between the third quartile (, 75th percentile) and the first quartile (, 25th percentile): The IQR measures the spread of the middle of the data and is completely unaffected by extreme scores.
- Formal Outlier Threshold: A data value is mathematically classified as an outlier if:
Impact of Outliers and Distribution Skewness
A vital concept tested on the TExES 391 is how data distributions dictate the choice of statistical measures:
- Symmetric Distribution (Bell-Shaped): . The Mean is preferred because it incorporates all mathematical values.
- Right-Skewed Distribution (Positive Skew): Stretches toward high extreme values. Extreme high outliers pull the mean to the right: .
- Left-Skewed Distribution (Negative Skew): Stretches toward low extreme values. Extreme low outliers pull the mean to the left: .
Core Statistical Rule: Because the mean and range are heavily distorted by extreme outliers, the Median is the superior measure of central tendency and the IQR is the superior measure of variability whenever a dataset is skewed or contains outliers.
Designing Statistical Investigations
The framework expects teachers to design, conduct, analyze, and interpret statistical investigations of real-world problems. A common four-step cycle is:
- Formulate a statistical question, one that anticipates variability ("How long do fifth graders spend reading each night?" rather than "How long did I read last night?").
- Collect data with a plan: a survey, a measurement, or an experiment. A random sample gives every member of the population an equal chance to be chosen. A biased sample, such as asking only the soccer team about favorite sports, can mislead.
- Analyze data with displays and summary statistics.
- Interpret results in context, including what the data cannot show.
The Normal Distribution
Many measurements, such as the heights of students of the same age, form a normal distribution: a symmetric, bell-shaped curve where the mean, median, and mode are equal. In a normal distribution:
- about 68% of values fall within one standard deviation of the mean;
- about 95% fall within two standard deviations; and
- about 99.7% fall within three standard deviations.
If fifth graders' heights are normally distributed with mean 55 inches and standard deviation 3 inches, about 95% of fifth graders are between 49 and 61 inches tall. This is an inference about the population from the shape of the graph.
Reading Graphs Critically
Students should check titles, labels, scales, and sources. A vertical axis that does not start at zero can exaggerate small differences, and a pictograph's key changes what each icon means. Drawing conclusions "from any data graph" requires reading these features first.
A sixth-grade teacher compiles the weekly independent reading logs for seven students: 40 minutes, 45 minutes, 50 minutes, 50 minutes, 55 minutes, 60 minutes, and 340 minutes. The teacher asks the class to identify which measure of central tendency best represents the typical reading time of the group. Which measure should be selected, and why?
The mean, because it calculates the mathematical fair share by incorporating all numerical values in the dataset.
The mode, because 50 minutes appears with the highest frequency and represents the peak of the data distribution.
The median, because the extreme outlier of 340 minutes substantially inflates the mean, whereas the median remains resistant to extreme values.
The range, because it evaluates the full dispersion between the lowest and highest values in the distribution.
A student wants to know the favorite after-school activity of all fourth graders at her school, so she surveys the 12 members of the robotics club. What is the main problem with her plan?
The sample is too large to analyze.
The sample is biased because robotics club members are not representative of all fourth graders.
Surveys can never be used to answer statistical questions.
She should use a line graph instead of a survey.
Sections you finish are checked off in the contents.