11.1 Statistics: Data Displays, Measures & Investigations

Key Takeaways

  • Bar graphs and pictographs display categorical data, while dot plots, histograms, stem-and-leaf plots, and box plots display numerical data.

  • The median resists outliers better than the mean, so it often better represents skewed data.

  • Range and interquartile range describe how spread out the data are.

  • In a normal distribution, about 68% of values fall within one standard deviation of the mean and about 95% within two.

  • A random sample represents a population, while a biased sample can produce misleading conclusions.

Last updated: October 2026

Overview & Exam Relevance

Competency 005 (Probability and Statistics) of the TExES Core Subjects EC-6 (391) Mathematics subject exam (902) covers how teachers guide students through statistical investigation: collecting, organizing, displaying, and interpreting data. This section focuses on statistics; probability and the mathematical processes of Competency 006 follow in the next sections. In modern Texas elementary curricula, statistics and probability are not isolated arithmetic procedures; they provide the empirical tools through which students collect, organize, represent, interpret, and draw evidence-based conclusions from data.

Candidates must select appropriate graphical displays for categorical and numerical data, compute and compare measures of center and spread for different distribution shapes, recognize when the normal distribution applies, and design simple statistical investigations that answer real questions.


Graphical Representations of Data (Competency 005)

Selecting an effective data display requires distinguishing between categorical (qualitative) data (non-numeric labels, such as eye color or vehicle type) and numerical (quantitative) data (measurable or countable quantities, such as height, test scores, or temperature).

DATA REPRESENTATION DECISION FRAMEWORK
│
├── Categorical Data (Labels / Groups)
│   ├── Pictograph (Scaled key/legend; early elementary)
│   ├── Bar Graph (Discrete separated bars; single & double comparisons)
│   └── Circle Graph / Pie Chart (Part-to-whole percentage proportions)
│
└── Numerical Data (Continuous or Countable Measurements)
    ├── Dot Plot / Line Plot (Number line with dots; small datasets n < 30)
    ├── Stem-and-Leaf Plot (Maintains raw values; stem = tens, leaf = ones)
    ├── Histogram (Continuous bins/intervals; touching bars; no gaps)
    ├── Box Plot (5-Number Summary: Min, Q1, Median, Q3, Max; shows IQR)
    └── Scatter Plot (Bivariate pairs; analyzes correlation/association)

Comprehensive Comparison of Graphical Displays

Graphical DisplayData TypeStructural FeaturesOptimal Pedagogical PurposeCommon Student Misconception
PictographCategoricalIcons represent fixed quantities; requires a clear key/legendIntroducing data representation in Grades K–3; visual countingIgnoring the key (e.g., counting 4 icons as 4 items when each icon represents 5 items).
Bar GraphCategoricalBars separated by distinct spaces; vertical or horizontal orientationComparing frequencies across distinct, non-overlapping categoriesConfusing bar graphs with histograms; failing to read fractional bar heights accurately.
Dot Plot (Line Plot)NumericalIndividual dots or X's stacked above a continuous number lineVisualizing distributions, clusters, peaks (mode), gaps, and outliers for small sets (n<30n < 30)Placing dots on numbers that have zero frequency rather than leaving the number blank.
Stem-and-Leaf PlotNumericalStems display leading place values; leaves display trailing digits; requires a keyOrganizing data sequentially while preserving every individual raw data valueOmitting the key; failing to list repeated leaf values (e.g., writing 5 once for scores of 45 and 45).
HistogramContinuous NumericalBars touch without spaces (spaces only occur if an interval has zero frequency); equal interval binsDisplaying frequency distributions for large numerical datasetsMistaking histograms for bar graphs; assuming bars touch arbitrarily rather than representing continuous bins.
Circle Graph (Pie Chart)CategoricalSlices represent proportional percentages of a whole (100%100\% or 360∘360^\circ)Displaying relative part-to-whole proportionsBelieving larger circle graphs necessarily represent larger total quantities than smaller ones.
Box Plot (Box-and-Whisker)NumericalDisplays the Five-Number Summary (Min, Q1Q_1, Median, Q3Q_3, Max); box width represents IQRVisualizing median, spread, quartiles (each representing 25%25\% of data), and skewnessAssuming the wider whisker or box section contains more data points (each quartile contains exactly 25%25\%).
Scatter PlotBivariate NumericalCoordinate points plotted on Cartesian axes (x,y)(x, y)Analyzing association/correlation (positive, negative, none; linear vs non-linear)Conflating correlation with direct causation.

Measures of Central Tendency & Variability

Statistical summary values describe the center and dispersion of numerical datasets.

Measures of Central Tendency

  1. Mean (Arithmetic Average): The sum of all observations divided by the total number of observations (xˉ=∑xn\bar{x} = \frac{\sum x}{n}). Conceptually modeled as the "fair share" (leveling stacks of cubes) or the physical "balance point" of a data distribution on a fulcrum.
  2. Median (Physical Center): The middle observation when data are arranged in numerical order. If the sample size nn is odd, it is the exact middle value; if nn is even, it is the arithmetic average of the two middle values.
  3. Mode (Most Frequent): The data value that occurs with the highest frequency. A dataset may be unimodal (one mode), bimodal (two modes), multimodal, or have no mode. The mode is the only measure of central tendency applicable to nominal categorical data (e.g., modal favorite ice cream flavor).

Measures of Variability (Spread)

  1. Range: The simplest measure of dispersion: Range=Maximum−Minimum\text{Range} = \text{Maximum} - \text{Minimum}.
  2. Interquartile Range (IQR): The difference between the third quartile (Q3Q_3, 75th percentile) and the first quartile (Q1Q_1, 25th percentile): IQR=Q3−Q1\text{IQR} = Q_3 - Q_1 The IQR measures the spread of the middle 50%50\% of the data and is completely unaffected by extreme scores.
  3. Formal Outlier Threshold: A data value is mathematically classified as an outlier if: Value<Q1−1.5(IQR)orValue>Q3+1.5(IQR)\text{Value} < Q_1 - 1.5(\text{IQR}) \quad \text{or} \quad \text{Value} > Q_3 + 1.5(\text{IQR})

Impact of Outliers and Distribution Skewness

A vital concept tested on the TExES 391 is how data distributions dictate the choice of statistical measures:

  • Symmetric Distribution (Bell-Shaped): Mean≈Median≈Mode\text{Mean} \approx \text{Median} \approx \text{Mode}. The Mean is preferred because it incorporates all mathematical values.
  • Right-Skewed Distribution (Positive Skew): Stretches toward high extreme values. Extreme high outliers pull the mean to the right: Mode<Median<Mean\text{Mode} < \text{Median} < \text{Mean}.
  • Left-Skewed Distribution (Negative Skew): Stretches toward low extreme values. Extreme low outliers pull the mean to the left: Mean<Median<Mode\text{Mean} < \text{Median} < \text{Mode}.

Core Statistical Rule: Because the mean and range are heavily distorted by extreme outliers, the Median is the superior measure of central tendency and the IQR is the superior measure of variability whenever a dataset is skewed or contains outliers.


Designing Statistical Investigations

The framework expects teachers to design, conduct, analyze, and interpret statistical investigations of real-world problems. A common four-step cycle is:

  1. Formulate a statistical question, one that anticipates variability ("How long do fifth graders spend reading each night?" rather than "How long did I read last night?").
  2. Collect data with a plan: a survey, a measurement, or an experiment. A random sample gives every member of the population an equal chance to be chosen. A biased sample, such as asking only the soccer team about favorite sports, can mislead.
  3. Analyze data with displays and summary statistics.
  4. Interpret results in context, including what the data cannot show.

The Normal Distribution

Many measurements, such as the heights of students of the same age, form a normal distribution: a symmetric, bell-shaped curve where the mean, median, and mode are equal. In a normal distribution:

  • about 68% of values fall within one standard deviation of the mean;
  • about 95% fall within two standard deviations; and
  • about 99.7% fall within three standard deviations.

If fifth graders' heights are normally distributed with mean 55 inches and standard deviation 3 inches, about 95% of fifth graders are between 49 and 61 inches tall. This is an inference about the population from the shape of the graph.


Reading Graphs Critically

Students should check titles, labels, scales, and sources. A vertical axis that does not start at zero can exaggerate small differences, and a pictograph's key changes what each icon means. Drawing conclusions "from any data graph" requires reading these features first.

Test Your Knowledge

A sixth-grade teacher compiles the weekly independent reading logs for seven students: 40 minutes, 45 minutes, 50 minutes, 50 minutes, 55 minutes, 60 minutes, and 340 minutes. The teacher asks the class to identify which measure of central tendency best represents the typical reading time of the group. Which measure should be selected, and why?

A

The mean, because it calculates the mathematical fair share by incorporating all numerical values in the dataset.

B

The mode, because 50 minutes appears with the highest frequency and represents the peak of the data distribution.

C

The median, because the extreme outlier of 340 minutes substantially inflates the mean, whereas the median remains resistant to extreme values.

D

The range, because it evaluates the full dispersion between the lowest and highest values in the distribution.

Test Your Knowledge

A student wants to know the favorite after-school activity of all fourth graders at her school, so she surveys the 12 members of the robotics club. What is the main problem with her plan?

A

The sample is too large to analyze.

B

The sample is biased because robotics club members are not representative of all fourth graders.

C

Surveys can never be used to answer statistical questions.

D

She should use a line graph instead of a survey.

Sections you finish are checked off in the contents.