Nursing Research & Biostatistics: Research Process, Designs, Sampling, Ethics & Statistical Measures

Key Takeaways

  • A true experiment has three features: manipulation of the independent variable, a control group, and randomisation.

  • Simple random, systematic, stratified and cluster sampling are probability methods; convenience, purposive, quota and snowball sampling are non-probability methods.

  • A Type I error rejects a true null hypothesis (a false positive), and a Type II error fails to reject a false null hypothesis (a false negative).

  • In a normal distribution about 68% of values lie within 1 SD of the mean, about 95% within 2 SD and about 99.7% within 3 SD.

  • The chi-square test examines association between categorical variables, the t-test compares two means and ANOVA compares three or more means.

Last updated: October 2026

Nursing research is the systematic inquiry that builds knowledge for nursing practice. Biostatistics applies statistics to health data. Recruitment questions test vocabulary (variable, hypothesis, sampling), the order of the research process, and simple calculations of mean, median, mode and standard deviation.


1. Types of Research

ClassificationTypes
By purposeBasic (pure)—builds theory; applied—solves practical problems
By approachQuantitative—numbers, measurement, statistics; qualitative—words, meanings, experiences
Quantitative designsExperimental (true experiment, e.g., a randomised controlled trial); quasi-experimental (lacks randomisation or a control group); non-experimental (descriptive, correlational, survey, ex post facto)
Qualitative designsPhenomenology (lived experience), grounded theory (generates theory), ethnography (culture), case study, historical research
By timeCross-sectional (one point in time); longitudinal (over time—prospective or retrospective)

A true experiment has three features: manipulation of the independent variable, a control group, and randomisation.


2. Steps of the Research Process

  1. Identify and state the research problem.
  2. Review the literature.
  3. Select a conceptual or theoretical framework.
  4. State the objectives and hypotheses.
  5. Define the variables operationally.
  6. Choose the research design.
  7. Define the population and select the sample.
  8. Develop the tool; establish reliability and validity.
  9. Conduct a pilot study (a small trial run to test feasibility and the tool).
  10. Collect data.
  11. Analyse and interpret the data.
  12. Communicate findings (report, publication) and use them in practice.

3. Variables and Hypotheses

  • Independent variable: the presumed cause, manipulated by the researcher (e.g., a teaching programme).
  • Dependent variable: the outcome measured (e.g., knowledge score).
  • Extraneous (confounding) variable: an uncontrolled factor that may affect the outcome (e.g., prior education).
  • Research (alternative) hypothesis (H₁): predicts a relationship or difference. It may be directional ("will increase") or non-directional ("will differ").
  • Null hypothesis (H₀): states that there is no relationship or difference; statistical tests try to reject it.
H₀ is actually trueH₀ is actually false
Reject H₀Type I error (α)—false positiveCorrect decision (power = 1 − β)
Do not reject H₀Correct decisionType II error (β)—false negative

By convention, a result is statistically significant when p < 0.05, meaning a result this extreme would occur less than 5% of the time if H₀ were true.


4. Population and Sampling

  • Population: the whole group of interest. Sample: the subset studied. Sampling frame: the list from which the sample is drawn.
Probability sampling (each member has a known chance)Non-probability sampling
Simple random—lottery or random-number tableConvenience—whoever is available
Systematic—every kth unit after a random startPurposive (judgemental)—chosen for specific characteristics
Stratified—divide into strata (e.g., by age), then sample randomly from eachQuota—fill set numbers in each category without randomisation
Cluster—randomly select whole groups (villages, wards)Snowball—participants refer others (useful for hidden groups)
Multistage—combine methods in stages

Systematic sampling example: from a list of 600 patients, a sample of 60 is needed, so the sampling interval is 600 ÷ 60 = 10; pick a random start between 1 and 10 and then every 10th patient.


5. Data-Collection Tools and Their Quality

  • Questionnaire (self-administered), interview schedule, observation checklist, rating scales such as the 5-point Likert scale (strongly agree → strongly disagree), and physiological measures (BP, blood glucose).
  • Reliability (consistency): test-retest, inter-rater, split-half, and internal consistency (Cronbach's alpha, commonly 0.7 or above is acceptable).
  • Validity (accuracy): content validity (expert review), criterion validity, construct validity.

6. Research Ethics

  • Nuremberg Code (1947): voluntary consent is essential.
  • Declaration of Helsinki (1964, revised since): ethical principles for medical research.
  • Belmont Report (1979): respect for persons, beneficence and justice.
  • In India, the ICMR National Ethical Guidelines for Biomedical and Health Research Involving Human Participants (2017) require ethics committee approval, informed consent, confidentiality, voluntary participation with the right to withdraw, and protection of vulnerable groups.

7. Biostatistics

Levels of measurement (NOIR)

LevelCharacteristicsExample
NominalCategories with no orderBlood group, sex, religion
OrdinalOrdered categories, unequal intervalsPain mild/moderate/severe; socio-economic class
IntervalEqual intervals, no true zeroTemperature in °C
RatioEqual intervals and a true zeroWeight, height, pulse rate

Measures of central tendency

  • Mean = sum of values ÷ number of values (affected by extreme values).
  • Median = middle value when data are arranged in order (preferred for skewed data).
  • Mode = most frequent value.
  • Empirical relation in moderately skewed data: Mode ≈ 3 Median − 2 Mean.

Worked example: pulse rates of 7 patients: 72, 76, 80, 80, 84, 88, 108.

  • Mean = 588 ÷ 7 = 84
  • Median (4th value) = 80
  • Mode = 80
  • Range = 108 − 72 = 36

The mean exceeds the median because the single high value (108) pulls it up—a positively skewed distribution.

Measures of dispersion

  • Range = highest − lowest value.
  • Variance = average of the squared deviations from the mean; standard deviation (SD) = √variance.
  • Coefficient of variation = (SD ÷ mean) × 100, used to compare variability of different measures.

The normal distribution

A symmetrical, bell-shaped curve in which mean = median = mode. About 68% of values lie within ±1 SD, 95% within ±2 SD (exactly 95% within ±1.96 SD) and 99.7% within ±3 SD.

Correlation and tests of significance

  • Correlation coefficient (r) ranges from −1 to +1; 0 means no linear relationship. Pearson's r for interval/ratio data; Spearman's rank correlation for ordinal data. Correlation does not prove causation.
  • t-test: compares the means of two groups (paired t-test for before-after measurements in the same group).
  • ANOVA: compares means of three or more groups.
  • Chi-square test: tests association between categorical variables (e.g., smoking status and TB).

Diagrams and graphs

GraphUse
Bar diagramDiscrete or categorical data; bars separated by gaps
HistogramContinuous grouped data; adjacent bars with no gaps
Frequency polygonJoins the mid-points of histogram bars
Pie chartProportions of a whole
Line graphTrends over time
Scatter diagramRelationship (correlation) between two variables
OgiveCumulative frequency; used to read the median
Test Your Knowledge

A researcher divides nursing students by year of study and then randomly selects an equal proportion from each year. Which sampling technique is this?

A

Stratified random sampling

B

Multistage cluster sampling

C

Non-random quota sampling

D

Systematic random sampling

Test Your Knowledge

The haemoglobin values (g/dL) of five women are 8, 9, 10, 10 and 13. What are the mean and median?

A

Mean 10, median 10

B

Mean 10.4, median 10

C

Mean 9, median 10

D

Mean 10, median 9

Test Your Knowledge

A study concludes that a new dressing heals wounds faster when, in reality, it has no effect. What kind of error has occurred?

A

Measurement reliability error

B

Type II error

C

Type I error

D

Sampling frame error

Sections you finish are checked off in the contents.