7.4 Basic Statistics: Frequency Distributions, Central Tendency, Risk Ratios & Odds Ratios

Key Takeaways

  • A frequency distribution lists each category with its count, percent (relative frequency) and often a cumulative percent; the percents must sum to 100.

  • The mean is the arithmetic average, the median is the middle value of ordered data, and the mode is the most frequent value; the median is preferred for skewed data such as days from diagnosis to treatment.

  • Relative risk (risk ratio) = risk in the exposed ÷ risk in the unexposed, calculated from cohort data; a value of 1 means no association.

  • Odds ratio = (a × d) ÷ (b × c) from a 2×2 table, the measure used in case-control studies; for rare diseases it approximates the relative risk.

  • A 95% confidence interval for a ratio that includes 1.0 means the association is not statistically significant at the 0.05 level.

Last updated: September 2026

Why registrars need basic statistics

Registrars answer data requests, prepare cancer committee and annual reports, and support quality and research studies. You do not need advanced biostatistics for the exam, but you must be able to summarize registry data correctly and interpret the risk measures that appear in the studies your data support.

Types of data

TypeDefinitionRegistry examples
Nominal (categorical)Named categories with no orderPrimary site, sex, race, Class of Case
OrdinalOrdered categories without equal spacingStage group (I–IV), grade (1–3), performance status
Discrete numericCountsNumber of positive lymph nodes
Continuous numericMeasurements on a scaleAge, tumor size in mm, days from diagnosis to treatment

The data type decides the summary. Categorical data are summarized with counts and percents, and numeric data with measures of central tendency and spread.

Frequency distributions

A frequency distribution lists each value or category with the number of cases (frequency), the relative frequency (percent of the total), and often the cumulative percent.

Example: stage at diagnosis for 400 analytic breast cases

AJCC stage groupFrequencyPercentCumulative percent
06015.0%15.0%
I18045.0%60.0%
II10025.0%85.0%
III369.0%94.0%
IV123.0%97.0%
Unknown123.0%100.0%
Total400100.0%—

Rules of thumb:

  • Always report the denominator. "45% stage I" means nothing without "n = 400."
  • Show unknowns. Hiding them inflates the other percentages and conceals a data quality problem.
  • Group continuous data into mutually exclusive, exhaustive intervals, such as age groups 0–14, 15–39, 40–64 and 65+, with no gaps or overlaps.
  • Choose the right chart. Use bar charts for categories, histograms for grouped continuous data, line graphs for trends over time, and pie charts only for a few parts of one whole.

Measures of central tendency

MeasureDefinitionBest use
MeanSum of the values ÷ number of valuesRoughly symmetric numeric data
MedianThe middle value when the data are ordered (the average of the two middle values when n is even)Skewed data or data with outliers
ModeThe most frequent valueCategorical data; finding the most common category

Worked example: days from diagnosis to first treatment for 9 patients

Ordered values: 12, 14, 15, 18, 21, 22, 22, 30, 160.

  • Mean = (12 + 14 + 15 + 18 + 21 + 22 + 22 + 30 + 160) ÷ 9 = 314 ÷ 9 ≈ 34.9 days
  • Median = the 5th value = 21 days
  • Mode = 22 days, the only value that appears twice

The one 160-day outlier pulls the mean up by about 14 days. The median better represents a typical patient, which is why time-to-treatment reports usually show the median.

Measures of spread

  • Range: maximum − minimum (160 − 12 = 148 days).
  • Interquartile range (IQR): the 75th percentile − the 25th percentile. It covers the middle 50% of the data and resists outliers.
  • Standard deviation: the typical distance of values from the mean. It is most useful with roughly bell-shaped data.

Ratios, proportions and rates

  • Ratio: one quantity divided by another, where the two may be unrelated (male-to-female cases, 1.2:1).
  • Proportion: a part divided by the whole, where the numerator is included in the denominator (the percent of cases diagnosed at stage I).
  • Rate: a proportion over a period of time in a population at risk (incidence per 100,000 per year; see Section 7.3).

Risk and the risk ratio (relative risk)

In a cohort study, exposed and unexposed people are followed forward in time. Arrange the results in a 2×2 table:

DiseaseNo diseaseTotal
Exposedaba + b
Unexposedcdc + d
  • Risk (cumulative incidence) in the exposed = a ÷ (a + b); in the unexposed = c ÷ (c + d).
  • Relative risk (RR), also called the risk ratio = [a ÷ (a + b)] ÷ [c ÷ (c + d)].
  • Risk difference (attributable risk) = the risk in the exposed − the risk in the unexposed.

Worked example: 1,000 smokers and 2,000 non-smokers are followed for 10 years.

Lung cancerNo lung cancerTotal
Smokers309701,000
Non-smokers61,9942,000
  • The risk in smokers is 30/1,000 = 0.030, and in non-smokers it is 6/2,000 = 0.003.
  • RR = 0.030 ÷ 0.003 = 10. Smokers had 10 times the 10-year risk.
  • The risk difference is 0.030 − 0.003 = 0.027, or 27 excess cases per 1,000 smokers.

The odds ratio

In a case-control study, people are selected because they have the disease (cases) or do not (controls), and past exposure is compared. The design fixes how many people have the disease, so risk cannot be calculated. The study uses the odds ratio (OR) instead:

  • OR = (a × d) ÷ (b × c), the cross-product of the 2×2 table.

Worked example: 200 mesothelioma cases and 400 controls are asked about asbestos exposure.

CasesControls
Exposed120 (a)100 (b)
Unexposed80 (c)300 (d)

OR = (120 × 300) ÷ (100 × 80) = 36,000 ÷ 8,000 = 4.5. The odds of past asbestos exposure were 4.5 times as high among cases as among controls. When a disease is rare, the OR approximates the RR.

Interpreting ratios

Value of RR or ORInterpretation
1.0No association
Greater than 1.0Exposure associated with higher risk (a possible risk factor)
Less than 1.0Exposure associated with lower risk (a possible protective factor)
  • 95% confidence interval: if the interval for an RR or OR includes 1.0 (for example, 0.8–2.3), the association is not statistically significant at the 0.05 level. If it excludes 1.0 (for example, 1.4–3.9), it is.
  • Association is not causation. Bias, confounding (such as age or smoking) and chance must be considered before concluding that an exposure causes cancer.
Loading diagram...
Choosing a summary measure
Test Your Knowledge

In a case-control study, 90 of 150 cases and 60 of 300 controls report a particular exposure. What is the odds ratio?

A

6.0

B

3.0

C

1.5

D

0.17

Test Your Knowledge

A registrar reports the number of days from diagnosis to surgery for 51 patients. Most waited 20–35 days, but three waited more than 200 days because of delayed referrals. Which summary best represents the typical wait?

A

The mean, because it uses every value

B

The median, because it is not pulled toward the extreme values

C

The mode, because it is always the most accurate average

D

The range, because it shows the longest wait

Test Your Knowledge

In a cohort study, 40 of 2,000 exposed workers and 10 of 2,000 unexposed workers develop bladder cancer. What is the relative risk, and what does it mean?

A

4.0; exposed workers had four times the risk of the unexposed workers

B

0.25; exposure was protective

C

30; there were 30 more cases in the exposed group

D

4.0; the result proves the exposure causes bladder cancer

Sections you finish are checked off in the contents.