7.3 Sampling Distributions, Probability and Inferential Statistics

Key Takeaways

  • TOS competencies A.2.1 and A.2.2 require candidates to apply sampling distributions using sex-disaggregated data and to relate probability theory to inferential statistics.
  • Probability sampling — simple random, systematic, stratified, cluster and multistage — permits generalisation; non-probability sampling does not.
  • The sampling distribution of the mean is the distribution of sample means from repeated samples, and its spread is the standard error.
  • Larger samples reduce the standard error, narrowing confidence intervals and increasing precision.
  • A 95 percent confidence interval means that 95 percent of intervals constructed this way would contain the true population value.
Last updated: September 2026

7.3 Sampling Distributions, Probability and Inferential Statistics

Blueprint anchor. TOS sub-topic A.2 Statistical Models carries 4 items, of which A.2.1 requires applying sampling distribution using sex-disaggregated data and A.2.2 requires relating probability theory and inferential statistics.


1. Why Sample at All

A municipality with 18,000 households cannot be fully enumerated for every planning question. Sampling substitutes a manageable subset for the whole, and the price of that convenience is sampling error — the difference between the sample statistic and the true population parameter. Inferential statistics is the discipline of quantifying that price.


2. Probability Sampling

Every unit has a known, non-zero chance of selection. Only probability samples support statistical generalisation.

DesignProcedureWhen to use in welfare workCaution
Simple randomEvery unit equally likely; draw by lot or random numbersSmall, complete and accessible listRequires a full sampling frame
SystematicSelect every k-th unit after a random startLong client registers, queue-based listsFails if the list has a hidden periodicity
StratifiedDivide into strata, then sample within eachWhen a subgroup must be represented — sex, sector, barangayRequires stratum membership to be known in advance
ClusterSample whole groups, then study all units insideGeographically dispersed populations; sitios as clustersLess precise than simple random for the same size
MultistageSuccessive sampling of clusters then unitsProvincial and national household surveysComplexity in estimating error

Proportionate stratified sampling allocates the sample to strata in proportion to their share of the population; disproportionate allocation oversamples a small stratum so that it can be analysed separately. When the TOS asks for sampling using sex-disaggregated data, it is asking for stratification by sex so that findings can be reported and compared for women and men rather than merged into an average that hides both.


3. Non-Probability Sampling

DesignProcedureLegitimate useLimitation
PurposiveDeliberately select information-rich casesQualitative studies; key informantsNo statistical generalisation
QuotaFill preset category counts by availabilityRapid appraisal under time pressureSelection bias inside quotas
Convenience / accidentalWhoever is availablePilot testing an instrumentWeakest external validity
SnowballRespondents refer othersHidden and stigmatised populationsNetwork-bounded sample

[!IMPORTANT] Standing rule. A satisfaction survey completed by whoever happened to visit the office cannot be reported as "clients are 87 percent satisfied". It reports only on those who came and chose to answer. Attaching population-level language to a convenience sample is the single most common statistical misuse in agency reporting, and it is examinable.


4. Probability Theory in One Page

  • Probability ranges from 0 (impossible) to 1 (certain), and is often expressed as a percentage.
  • Independent events do not affect each other's probability; mutually exclusive events cannot both occur.
  • The addition rule governs "or" questions; the multiplication rule governs "and" questions for independent events.
  • The law of large numbers holds that as sample size grows, the sample mean converges on the population mean.

Probability supplies the bridge to inference: because we know how sample statistics behave in the long run, we can attach a quantified confidence to a single sample's result.


5. The Sampling Distribution and the Standard Error

Imagine drawing many samples of the same size from one population and computing each sample's mean. The distribution of those means is the sampling distribution of the mean. Three properties matter:

  1. Its centre equals the population mean.
  2. Its spread — the standard error of the mean — is smaller than the spread of individual observations.
  3. By the central limit theorem, it approaches a normal shape as sample size increases, even when the underlying variable is skewed.

The practical consequence: the standard error shrinks as sample size grows, so larger samples give more precise estimates. It shrinks with the square root of sample size, which is why quadrupling a sample only halves the error — a point that disciplines unrealistic survey plans.


6. Confidence Intervals

A confidence interval is a range of plausible values for the population parameter, with a stated confidence level.

Estimate: 62 percent of assisted households are female-headed, 95 percent confidence interval 57 to 67 percent.

Correct interpretation: if this sampling and estimation procedure were repeated many times, about 95 percent of the intervals produced would contain the true population proportion.

Incorrect interpretations to reject in an item:

  • "There is a 95 percent chance the true value is 62 percent."
  • "95 percent of households are between 57 and 67 percent female-headed."
  • "The result is 95 percent accurate."

A wider interval means less precision. Narrowing it requires a larger sample, lower variability, or accepting a lower confidence level.


7. Worked Practice Application

A municipal social welfare office must estimate the prevalence of unmet health-service needs among 4,200 senior citizens across 24 barangays, and must be able to report separately for women and men and for upland and lowland barangays.

Design decision.

  • A convenience sample taken at the Office of the Senior Citizens Affairs counter would over-represent mobile, already-connected seniors — precisely the group least likely to have unmet needs. Rejected.
  • Stratified sampling by sex and by upland/lowland location is required, because both are reporting domains.
  • Because upland seniors are only about 15 percent of the population, proportionate allocation would yield too few for separate analysis, so disproportionate allocation oversamples the upland stratum, with weighting applied when a municipality-wide figure is reported.
  • Within strata, systematic selection from the senior citizens registry is practical, provided the registry is not ordered in a way that repeats every k-th entry.

Reporting discipline. The report gives the estimate, the confidence interval, the sample size for each domain, and an explicit note that the registry excludes unregistered seniors — the most likely source of undercoverage bias. Naming the limitation is part of the professional product, not an admission of failure.

Test Your Knowledge

A municipal office surveys whoever visits the counter during one week and reports that "87 percent of senior citizens in the municipality are satisfied with services." What is the principal defect?

A
B
C
D
Test Your Knowledge

A study reports that 62 percent of assisted households are female-headed, with a 95 percent confidence interval of 57 to 67 percent. Which interpretation is correct?

A
B
C
D
Test Your Knowledge

Upland seniors form only 15 percent of a municipality's senior population, yet the office must report findings separately for them. Which sampling decision follows?

A
B
C
D