8.1 Data Collection, Organization & Sampling
Key Takeaways
- Primary data is collected directly by the researcher for a specific purpose, whereas secondary data is obtained from existing sources like Statistics Canada or EQAO.
- A census collects data from every member of a population, while a sample collects data from a representative subset to infer population characteristics.
- Probability sampling methods include simple random, systematic, and stratified sampling, which minimize selection bias compared to non-probability convenience sampling.
- Sampling bias occurs when a sample does not represent the population, non-response bias occurs when selected individuals fail to participate, and measurement bias arises from flawed tools or leading questions.
- Categorical data describes qualities or labels (nominal or ordinal), whereas numerical data represents quantities that are either discrete (countable whole numbers) or continuous (measurable real numbers).
8.1 Data Collection, Organization & Sampling
Quick Summary: Statistical literacy begins with sound data collection practices. This section examines the distinction between primary and secondary data sources, compares full-population censuses against representative samples, details probability sampling protocols (simple random, systematic, stratified) and non-probability convenience sampling, identifies key sources of sampling and measurement bias, and classifies statistical variables into categorical (nominal/ordinal) and numerical (discrete/continuous) types.
Primary vs. Secondary Data Sources
Data collection is the foundation of empirical inquiry in education and mathematics. Educators and researchers must distinguish between primary data and secondary data based on origin and ownership.
1. Primary Data
Primary data consists of original data collected directly by the researcher or educator first-hand for a specific, intended inquiry.
- Collection Methods: Direct classroom observations, diagnostic student testing, customized surveys, structured interviews, controlled educational experiments.
- Advantages: Tailored specifically to the research question; high degree of methodological control; known data quality, accuracy, and operational definitions.
- Disadvantages: Time-consuming to collect; labor-intensive; often expensive; limited by sample size resources.
2. Secondary Data
Secondary data consists of pre-existing data originally gathered, compiled, and published by another individual, agency, or institution for a different purpose.
- Collection Methods: Statistics Canada census databases, Education Quality and Accountability Office (EQAO) provincial assessment repositories, Ministry of Education public datasets, published academic journals, historical attendance records.
- Advantages: Cost-effective; instantly accessible; provides large-scale, population-level datasets that would be impossible for an individual educator to gather independently.
- Disadvantages: May not align perfectly with the specific research context; potential lack of control over original collection errors or measurement definitions; risk of outdated information.
| Feature | Primary Data | Secondary Data |
|---|---|---|
| Source | Direct first-hand collection | Pre-existing published sources |
| Collector | The active researcher/educator | Third-party agency (e.g., Statistics Canada, EQAO) |
| Specificity | High (custom-designed for topic) | Moderate to Low (requires adaptation) |
| Cost & Time | High resource investment | Low cost, immediately available |
| Control | Full methodological control | Dependent on original author's rigor |
Populations, Censuses, and Samples
To draw valid conclusions from data, researchers define the scope of their study using statistical definitions of populations and subsets.
graph TD
subgraph PopulationScope["Statistical Target Domain"]
POP["Target Population<br/>(Complete Set of All N Individuals)"]
POP --> CENSUS["Census<br/>(Collects data from EVERY individual | N = n)"]
POP --> SAMPLE["Sample<br/>(Collects data from subset n < N)"]
SAMPLE --> PROB["Probability Sampling<br/>(Random, Stratified, Systematic)"]
SAMPLE --> NONPROB["Non-Probability Sampling<br/>(Convenience, Voluntary Response)"]
end
Definitions:
- Population ($N$): The entire group of individuals, objects, or items under investigation that share a defined characteristic (e.g., all Grade 9 students in Ontario, or all elementary schools in a school district).
- Census: An exhaustive research investigation that collects data from every single member of the target population ($n = N$).
- Pros: Eliminates sampling error; provides exact population parameters.
- Cons: Often prohibitively expensive, logistically complex, and time-prohibitive for large populations.
- Sample ($n$): A representative subset of individuals selected from the target population ($n < N$) designed to reflect the population's characteristics.
- Pros: Feasible, efficient, and cost-effective.
- Cons: Subject to sampling error; requires careful sampling design to avoid bias.
Sampling Methods & Designs
Sampling methods fall into two primary categories: probability sampling (where every member has a known, non-zero chance of selection) and non-probability sampling.
1. Simple Random Sampling (SRS)
Every individual in the target population has an equal and independent chance of being selected, and every possible sample of size $n$ has an equal likelihood of selection.
- Mechanism: Assign each population member a unique identifier (1 to $N$), then use a random number generator or lottery system to select $n$ units.
- Example: Drawing student ID numbers from a computer algorithm to select 30 students from a school roster of 600.
2. Systematic Sampling
Individuals are selected at regular, fixed numerical intervals from an ordered population list.
- Mechanism: Calculate the sampling interval $k = \frac{N}{n}$ (where $N$ is population size and $n$ is sample size). Select a random starting point between 1 and $k$, then select every $k^{\text{th}}$ individual thereafter.
- Example: To select $n = 50$ students from a school list of $N = 1,000$, $k = \frac{1000}{50} = 20$. Randomly pick a start between 1 and 20 (e.g., 7th student), then pick every 20th student (27th, 47th, 67th...).
- Risk: If the population list contains periodic patterns (cyclical trends), systematic sampling can introduce systematic bias.
3. Stratified Random Sampling
The population is partitioned into mutually exclusive subgroups called strata based on shared characteristics (e.g., grade level, gender, geographic zone). A simple random sample is then drawn from each stratum in proportion to its representation in the overall population.
- Proportional Allocation Formula: where $N_i$ is the population size of stratum $i$, $N$ is total population, $n$ is total desired sample size, and $n_i$ is the sample size for stratum $i$.
- Advantage: Ensures key sub-populations are represented accurately without being under-sampled.
4. Convenience Sampling (Non-Probability)
Selecting individuals who are easily accessible, available, or willing to participate.
- Example: Surveying only the students sitting in the cafeteria during first period lunch.
- Critical Limitation: High risk of severe sampling bias; results cannot be generalized to the broader population.
| Sampling Method | Selection Strategy | Key Advantage | Key Risk / Disadvantage |
|---|---|---|---|
| Simple Random | Equal probability via random numbers | Unbiased, mathematically simple | Requires complete numbered list of population |
| Systematic | Select every $k^{\text{th}}$ item ($k = N/n$) | Easy to implement in field | Periodicity in list introduces bias |
| Stratified | Random sample within relative strata | Guarantees representation of subgroups | Requires prior knowledge of population strata |
| Convenience | Select available/accessible units | Fast, cheap, convenient | Severe selection bias; non-generalizable |
Sources of Statistical Bias
Bias refers to any systematic error in sampling, data collection, or measurement that causes sample statistics to consistently overestimate or underestimate a population parameter.
1. Sampling Bias (Selection Bias)
Occurs when the sampling methodology systematically excludes or under-represents certain segments of the population.
- Example: Conducting an online survey about internet access in rural Ontario schools naturally excludes families without reliable internet connection.
- Voluntary Response Bias: Occurs when sample participants self-select (e.g., call-in polls or open web surveys), attracting individuals with strong negative or positive extreme opinions.
2. Non-Response Bias
Occurs when individuals selected for a sample fail or refuse to respond to a survey, and non-responders differ systematically from responders in their characteristics or opinions.
- Example: A mail-back questionnaire sent to parents regarding school budget cuts where only highly dissatisfied parents return the form.
3. Measurement / Response Bias
Occurs when the measuring instrument or survey question design influences respondents to give inaccurate, untrue, or biased answers.
- Causes: Leading question phrasing (e.g., "Do you agree that our hard-working teachers deserve higher funding?"), sensitive topics causing social desirability bias, or poorly calibrated physical measuring tools.
Classification of Data Variables
Data collected in research is classified into categorical or numerical structures, determining the appropriate statistical calculations and graphs.
graph LR
DATA["Data Variables"] --> CAT["Categorical (Qualitative)"]
DATA --> NUM["Numerical (Quantitative)"]
CAT --> NOM["Nominal<br/>(No natural order: Eye colour, Subject)"]
CAT --> ORD["Ordinal<br/>(Natural order: Grade 9-12, Likert scale)"]
NUM --> DISC["Discrete<br/>(Countable whole numbers: Absences)"]
NUM --> CONT["Continuous<br/>(Measurable real numbers: Height, Time)"]
1. Categorical (Qualitative) Data
Data that describes attributes, qualities, or labels.
- Nominal Data: Categories with no intrinsic mathematical order or ranking (e.g., favorite subject: Math, Science, Art; transportation mode: Bus, Walk, Car).
- Ordinal Data: Categories that possess a natural meaningful order or ranking, but numerical differences between ranks cannot be precisely measured (e.g., grade level: Grade 9, 10, 11, 12; survey agreement: Disagree, Neutral, Agree).
2. Numerical (Quantitative) Data
Data expressed as numerical quantities where arithmetic operations are meaningful.
- Discrete Numerical Data: Data values that are countable, separate numbers (typically non-negative integers) with gaps between possible values (e.g., number of students absent: 0, 1, 2, 3; number of books checked out).
- Continuous Numerical Data: Data values that can take any real number value along a continuous scale within a given interval, limited only by measuring instrument precision (e.g., student height: 168.4 cm; running time: 14.25 seconds; temperature: 21.5°C).
Comprehensive Worked Sampling & Bias Scenario
Scenario Context: A high school principal in an Ontario secondary school wishes to survey student perspectives on a proposed new cell phone policy. The school has a total population of $N = 1,200$ students, broken down by grade level as follows:
- Grade 9: $360$ students
- Grade 10: $300$ students
- Grade 11: $300$ students
- Grade 12: $240$ students
The principal decides to collect a representative sample of $n = 200$ students using stratified random sampling proportional to population size.
Tasks:
- Calculate the exact number of students to sample from each grade stratum ($n_9, n_{10}, n_{11}, n_{12}$).
- Evaluate a competing proposal where the principal puts an open-access survey QR code on the school main office counter and analyzes the first 200 responses. Identify the sampling technique and sources of bias.
- Classify four survey variables collected: (a) Student Grade Level, (b) Daily Phone Usage (minutes), (c) Number of Phone Notification Interruptions per Class, and (d) Primary Cell Phone Operating System (iOS/Android).
Step-by-Step Solution:
-
Step 1: Calculate Stratified Proportional Sample Sizes
- Total population $N = 1,200$. Target sample $n = 200$.
- Overall sampling ratio $\frac{n}{N} = \frac{200}{1,200} = \frac{1}{6} \approx 0.1667$.
- Grade 9 Sample ($n_9$):
- Grade 10 Sample ($n_{10}$):
- Grade 11 Sample ($n_{11}$):
- Grade 12 Sample ($n_{12}$):
- Check Sum: $60 + 50 + 50 + 40 = 200$ students.
-
Step 2: Evaluate Competing QR-Code Proposal
- Sampling Technique: Convenience sampling combined with voluntary response sampling.
- Sources of Bias:
- Sampling / Selection Bias: Excludes students who do not visit the main office; over-represents students sent to the office or student council members.
- Voluntary Response Bias: Only students with strong emotional opinions (e.g., highly agitated against phone restrictions) will take the time to scan and respond.
-
Step 3: Classify Survey Data Variables
- (a) Grade Level (Gr 9, 10, 11, 12): Categorical Ordinal data (categories with a natural grade order).
- (b) Daily Phone Usage (minutes): Numerical Continuous data (time can be measured continuously in fractions of minutes).
- (c) Number of Notification Interruptions: Numerical Discrete data (countable whole numbers: 0, 1, 2, 3...).
- (d) Operating System (iOS, Android, Other): Categorical Nominal data (unordered categories).
A high school principal wants to sample 120 students from a school of 800 students (200 in Grade 9, 240 in Grade 10, 180 in Grade 11, and 180 in Grade 12). If she uses stratified random sampling proportional to grade population, how many Grade 10 students should be selected?
A researcher conducts a survey on student reading habits by posting a flyer on a library bulletin board asking students to scan a QR code and fill out an online questionnaire. Which primary source of statistical bias does this survey design introduce?
Which of the following variables collected in a school health study represents continuous numerical data?