3.2 Steps of Research & Experimental Methodologies

Key Takeaways

  • The scientific research process moves sequentially through problem identification, literature review, hypothesis development, research design, sampling, data collection, analysis, interpretation, and formal reporting.
  • In experimental methodology, the independent variable is actively manipulated to observe its causal effect upon the measured dependent variable while controlling for extraneous and confounding variables.
  • Internal validity assesses whether observed experimental outcomes are solely attributable to the manipulated treatment, whereas external validity reflects the generalizability of findings across populations and settings.
  • Probability sampling techniques (simple random, systematic, stratified, cluster) provide known, non-zero selection probabilities, whereas non-probability methods (purposive, quota, snowball, convenience) rely on subjective selection.
  • A Type I error (alpha) occurs when a true null hypothesis is incorrectly rejected (false positive), while a Type II error (beta) occurs when a false null hypothesis fails to be rejected (false negative).
Last updated: August 2026

Steps of Research & Experimental Methodologies

Conducting rigorous academic research requires mastery of a structured methodological pipeline, sound experimental controls, appropriate sampling frameworks, and formal statistical hypothesis testing procedures.


1. The Step-by-Step Research Pipeline

The scientific research process follows a canonical sequential pipeline, where each phase establishes the prerequisite foundation for subsequent steps:

┌───────────────────────────────────┐
│ 1. Formulate the Research Problem │ (Define scope, background & core question)
└─────────────────┬─────────────────┘
                  ▼
┌───────────────────────────────────┐
│ 2. Extensive Literature Review    │ (Identify knowledge gaps & conceptual models)
└─────────────────┬─────────────────┘
                  ▼
┌───────────────────────────────────┐
│ 3. Formulate Testable Hypotheses  │ (Null H0 vs Alternative H1; Directional/Non-directional)
└─────────────────┬─────────────────┘
                  ▼
┌───────────────────────────────────┐
│ 4. Develop the Research Design    │ (Blueprint: Experimental, Descriptive, or Correlational)
└─────────────────┬─────────────────┘
                  ▼
┌───────────────────────────────────┐
│ 5. Select Sampling Design         │ (Probability vs Non-Probability; Determine Sample Size)
└─────────────────┬─────────────────┘
                  ▼
┌───────────────────────────────────┐
│ 6. Collect Empirical Data         │ (Questionnaires, Structured Schedules, Observational Tools)
└─────────────────┬─────────────────┘
                  ▼
┌───────────────────────────────────┐
│ 7. Analyze & Process Data         │ (Descriptive stats, Inferential testing: t-test, ANOVA, Chi-Square)
└─────────────────┬─────────────────┘
                  ▼
┌───────────────────────────────────┐
│ 8. Test Hypotheses & Interpret    │ (Retain/Reject H0, Relate to theoretical literature)
└─────────────────┬─────────────────┘
                  ▼
┌───────────────────────────────────┐
│ 9. Prepare the Research Report    │ (Format thesis chapters, document limitations & implications)
└───────────────────────────────────┘

2. Experimental Research & Variable Taxonomy

Experimental research is the gold standard for establishing cause-and-effect relationships through active manipulation, control, and random assignment.

Classification of Variables

  • Independent Variable (IV / Treatment / Predictor): The antecedent condition or variable actively manipulated by the researcher to observe its effect (e.g., Computer-Assisted Instruction vs. Traditional Lecture).
  • Dependent Variable (DV / Criterion / Outcome): The consequential variable that is measured and observed to change in response to IV manipulation (e.g., Student Academic Achievement Scores).
  • Extraneous Variables: Uncontrolled variables not central to the study that could inadvertently influence the dependent variable (e.g., room temperature, baseline IQ, socioeconomic background).
  • Confounding Variables: An extraneous variable that covaries systematically with the independent variable, making it impossible to disentangle which variable caused the observed change in the dependent variable.
  • Moderating Variable: A secondary independent variable that modifies the direction or strength of the relationship between the IV and DV (e.g., student gender moderating the effect of instructional method on performance).
  • Mediating (Intervening) Variable: An internal explanatory mechanism or process through which the independent variable influences the dependent variable (e.g., instructional method increases student motivation, which in turn elevates achievement).

Control Mechanisms in Experiments

  1. Random Assignment (Randomization): Equates experimental and control groups across measured and unmeasured extraneous characteristics before treatment.
  2. Matching: Pairing subjects with identical attributes (e.g., matched IQ scores) across treatment arms.
  3. Statistical Control: Utilizing covariate adjustments such as Analysis of Covariance (ANCOVA) to partial out extraneous variance mathematically.

Threats to Experimental Validity

Type of ValidityCore DefinitionMajor Threats to Validity
Internal ValidityThe degree to which changes in the DV can be confidently attributed solely to the IV without alternative explanations.History (unplanned external events), Maturation (biological/psychological aging over time), Testing (sensitization from repeated pre-tests), Instrumentation (changes in scoring calibration), Statistical Regression (extreme scores regressing toward the mean), Selection Bias (non-equivalent comparison groups), Experimental Mortality / Attrition (differential drop-out rates).
External ValidityThe degree to which experimental findings can be generalized across diverse populations, settings, and times.Hawthorne Effect (participants altering behavior because they know they are observed), Pretest Sensitization (pretest altering responsiveness to treatment), Multiple-Treatment Interference (carryover effects from sequential treatments).

Quasi-Experimental Designs

When true randomization is ethically or practically unfeasible (such as using intact school classrooms), researchers employ Quasi-Experimental Designs (e.g., Non-Equivalent Control Group Pretest-Posttest Design or Interrupted Time-Series Design), controlling for confounds through statistical modeling.


3. Sampling Techniques: Probability vs. Non-Probability

Sampling is the process of selecting a representative subset of units from a defined target population to make valid statistical generalizations.

                                   SAMPLING TECHNIQUES
                                            │
            ┌───────────────────────────────┴───────────────────────────────┐
            ▼                                                               ▼
   PROBABILITY SAMPLING                                           NON-PROBABILITY SAMPLING
 (Every unit has known, non-zero P)                              (Subjective, Non-random selection)
            │                                                               │
 ┌──────────┼──────────┬──────────┐                      ┌──────────┬───────┼──────────┐
 ▼          ▼          ▼          ▼                      ▼          ▼       ▼          ▼
Simple   Systematic Stratified  Cluster               Purposive   Quota  Snowball  Convenience
Random   (k = N/n)  (Homog.   (Heterog.              (Judgment) (Non-    (Chain-   (Accidental)
                     strata)   clusters)                         random  referral)
                                                                 strata)

Probability Sampling Methods

  1. Simple Random Sampling: Every member of the population has an equal, independent probability of selection (lottery method or random number generators).
  2. Systematic Sampling: Selecting units at uniform intervals from an ordered sampling frame using the sampling fraction $k = N / n$, where $N$ is total population size and $n$ is desired sample size, starting from a randomly determined starting point between $1$ and $k$.
  3. Stratified Random Sampling: The heterogeneous population is divided into mutually exclusive, internally homogeneous subgroups (strata) based on key characteristics (e.g., rural vs. urban, income levels), after which simple random samples are drawn from each stratum (proportionate or disproportionate).
  4. Cluster Sampling: The population is naturally divided into internally heterogeneous groups (clusters, e.g., entire schools or geographical districts). Complete clusters are randomly sampled, and all individuals within chosen clusters are surveyed (often executed as Multi-stage Cluster Sampling).

Non-Probability Sampling Methods

  1. Purposive / Judgmental Sampling: Researchers hand-pick specific information-rich cases based on expert judgment and defined study criteria.
  2. Quota Sampling: Sets demographic quotas (e.g., 50% male, 50% female) mirroring population proportions, but units are selected conveniently rather than randomly.
  3. Snowball (Chain-Referral) Sampling: Initial participants recruit subsequent eligible peers from their social network, ideal for hidden, marginalized, or hard-to-reach populations (e.g., rare disease patients, underground artists).
  4. Convenience / Accidental Sampling: Selecting readily available and accessible participants (e.g., surveying students in the university cafeteria).

4. Hypothesis Testing & Decision Matrix (Type I vs. Type II Errors)

In statistical hypothesis testing, researchers evaluate empirical evidence against a default baseline assertion—the Null Hypothesis ($H_0$), which posits no true difference or relationship—compared to the Alternative Hypothesis ($H_1$).

The Four-Quadrant Decision Matrix

True State of Nature $\rightarrow$Null Hypothesis ($H_0$) is TRUENull Hypothesis ($H_0$) is FALSE
Researcher Rejects $H_0$Type I Error ($\alpha$)<br/>(False Positive / Producer's Risk)<br/>Probability = $\alpha$ (Level of Significance)Correct Decision<br/>(True Positive)<br/>Probability = $1 - \beta$ (Statistical Power)
Researcher Retains / Fails to Reject $H_0$Correct Decision<br/>(True Negative)<br/>Probability = $1 - \alpha$ (Confidence Level)Type II Error ($\beta$)<br/>(False Negative / Consumer's Risk)<br/>Probability = $\beta$
  • Level of Significance ($\alpha$): The probability of committing a Type I error, conventionally set a priori at $0.05$ (5%) or $0.01$ (1%).
  • Statistical Power ($1 - \beta$): The probability of correctly rejecting a false null hypothesis. Power is enhanced by increasing sample size ($n$), increasing effect size, or lowering measurement error.
Loading diagram...
Hypothesis Testing Error Framework
Test Your Knowledge

A statistical analyst conducts a hypothesis test at the 0.05 significance level and concludes that a new teaching method produces superior test scores, leading to the rejection of the null hypothesis. In reality, the teaching method has no effect, and the observed difference was due to random sampling fluctuation. What type of error was committed?

A
B
C
D
Test Your Knowledge

What is the structural difference between Stratified Random Sampling and Cluster Sampling?

A
B
C
D
Test Your Knowledge

In a longitudinal experimental evaluation of an educational intervention spanning two academic years, several participants from the experimental group move out of the district, resulting in uneven attrition between groups. Which threat to internal validity does this represent?

A
B
C
D
Test Your Knowledge

A sociologist wishes to study the lived experiences of undocumented migrant laborers who are difficult to locate through public directories. The researcher interviews three accessible individuals and asks them to refer other acquaintances within their network. What sampling method is being used?

A
B
C
D