Psychological Testing, Intelligence Theories, and Standardization

Key Takeaways

  • Intelligence testing evolved from Binet-Simon (mental age) to Terman's Stanford-Binet (Ratio IQ = [MA/CA] * 100) and Wechsler's WAIS/WISC (Deviation IQ based on normal distribution).
  • Theories of intelligence range from Spearman's general intelligence ('g') and Thurstone's primary abilities to Gardner's Multiple Intelligences and Sternberg's Triarchic Theory.
  • Psychometric evaluation requires Standardization, high Reliability (consistency), and strong Validity (accuracy of measure).
  • IQ scores follow a normal bell curve with a Mean of 100 and Standard Deviation of 15 (68% scoring between 85 and 115).
  • Intelligence is shaped by nature-nurture interactions, with heritability estimated at 50–80%, alongside the Flynn Effect demonstrating generational gains in IQ scores.
Last updated: July 2026

Psychological Testing, Intelligence Theories, and Standardization

Psychometrics is the subfield of psychology devoted to the quantitative measurement of mental abilities, traits, and educational achievements. Central to psychometrics is the assessment of intelligence—the global capacity to think rationally, act purposefully, profit from experience, and deal effectively with the environment.


History of Intelligence Testing

Modern intelligence testing originated in France in the early 20th century:

  • Alfred Binet and Théodore Simon (1905): Commissioned by the French government to identify children needing special education, they created the Binet-Simon Scale. They introduced the concept of Mental Age (MA)—the chronological age that typically corresponds to a given level of performance.
  • Lewis Terman (1916): A Stanford professor who adapted Binet's test for American school children, creating the Stanford-Binet Intelligence Scale. Terman adopted German psychologist William Stern's formula for the Ratio Intelligence Quotient (IQ):

Ratio IQ=(Mental Age (MA)Chronological Age (CA))×100\text{Ratio IQ} = \left( \frac{\text{Mental Age (MA)}}{\text{Chronological Age (CA)}} \right) \times 100

Example: A 10-year-old child (CA = 10) who scores at the level of an average 12-year-old (MA = 12) has a Ratio IQ of $(12 / 10) \times 100 = 120$. Ratio IQ works well for children but fails for adults whose mental development plateaus while chronological age continues to rise.

  • David Wechsler (1939): Designed the Wechsler Adult Intelligence Scale (WAIS-IV) and Wechsler Intelligence Scale for Children (WISC-V). Wechsler abandoned Ratio IQ in favor of Deviation IQ, which measures an individual's performance relative to a standardized representative sample of age-matched peers using a normal distribution curve.

Theories of Intelligence

Psychologists have long debated whether intelligence is a single general capacity or a collection of distinct mental abilities:

  • Charles Spearman (Two-Factor Theory): Used factor analysis (a statistical technique that correlates test items) to argue that intelligence consists of a general mental capacity called the 'g' factor (general intelligence) that underlies all cognitive tasks, alongside specific abilities called 's' factors.
  • L.L. Thurstone: Rejected Spearman's single 'g' factor, proposing that intelligence comprises 7 Primary Mental Abilities (verbal comprehension, word fluency, number facility, spatial visualization, associative memory, perceptual speed, and reasoning).
  • Howard Gardner (Theory of Multiple Intelligences): Proposed that intelligence is not a single scalar metric but exists as 8 or 9 distinct modalities: Linguistic, Logical-Mathematical, Spatial, Musical, Bodily-Kinesthetic, Interpersonal, Intrapersonal, Naturalistic, and (optionally) Existential. Gardner argued that brain damage can selectively impair one modality while leaving others intact.
  • Robert Sternberg (Triarchic Theory of Intelligence): Proposed three sub-theories of intelligence:
    1. Analytical Intelligence: Academic problem-solving and test-taking abilities.
    2. Creative Intelligence: Adapting novel solutions to new situations.
    3. Practical Intelligence: "Street smarts" and applying knowledge to real-world context.
  • Daniel Goleman: Popularized Emotional Intelligence (EQ), emphasizing self-awareness, impulse control, empathy, and social relationship management as crucial predictors of life success.

Psychometric Requirements: Standardization, Reliability, and Validity

For any psychological test to be scientifically meaningful, it must meet three rigorous criteria:

1. Standardization

Standardization involves establishing uniform administration and scoring procedures, as well as administering the test to a large, representative pre-test sample (normative group) to establish benchmarks.

On modern IQ tests, raw scores are mapped to a Normal Distribution (Bell Curve):

  • Mean = 100, Standard Deviation (SD) = 15.
  • 68.26% of the population falls within 1 SD of the mean (scores between 85 and 115).
  • 95.44% falls within 2 SDs (scores between 70 and 130).
  • 99.74% falls within 3 SDs (scores between 55 and 145).
  • An IQ below 70 combined with adaptive behavior deficits indicates Intellectual Disability; an IQ above 130 indicates Giftedness.

2. Reliability (Consistency)

Reliability refers to the consistency and stability of test results across repeated testing:

  • Test-Retest Reliability: Administering the same test twice to the same group and calculating correlation.
  • Alternate-Form Reliability: Administering two equivalent versions of a test to the same group.
  • Split-Half Reliability: Dividing a single test into two halves (e.g., odd vs. even questions) and checking consistency between scores.

3. Validity (Accuracy)

Validity refers to whether a test actually measures what it claims to measure:

  • Content Validity: The degree to which test items represent the entire domain of material being assessed.
  • Criterion-Related / Predictive Validity: How accurately test scores predict performance on a future external criterion (e.g., SAT scores predicting first-year college GPA).
  • Construct Validity: The extent to which a test measures an abstract theoretical construct (e.g., intelligence or anxiety).

Note: A test can be highly reliable without being valid, but a test cannot be valid without first being reliable.


Nature vs. Nurture and the Flynn Effect

  • Heritability: The proportion of variation among individuals in a population that is attributable to genetic differences. Intelligence heritability estimates range between 50% and 80%, supported by twin studies showing identical twins raised apart have higher IQ correlations than fraternal twins raised together.
  • Environmental Factors: Malnutrition, poverty, sensory deprivation, and lack of schooling decrease IQ performance, while enriched environments boost cognitive development.
  • Flynn Effect: Named after James Flynn, this phenomenon documents a widespread, worldwide generational rise in average IQ scores over the 20th century (approximately 3 points per decade), forcing test developers to re-standardize test norms periodically.

Theories of Intelligence Comparison

TheoristModel NameCore Structure / ModalitiesKey Feature
Charles SpearmanTwo-Factor TheoryGeneral factor ('g') + specific factors ('s')Derived via factor analysis of test items
L.L. ThurstonePrimary Mental Abilities7 independent mental factorsRejects single 'g' factor
Howard GardnerMultiple Intelligences8 to 9 distinct intelligence modalitiesBrain damage can isolate individual modalities
Robert SternbergTriarchic TheoryAnalytical, Creative, and PracticalIntegrates academic skills with real-world application
Daniel GolemanEmotional Intelligence (EQ)Empathy, self-regulation, self-awarenessPredicts interpersonal success and life outcomes
Test Your Knowledge

Using William Stern's Ratio IQ formula, what is the calculated IQ of an 8-year-old child who demonstrates a mental age of 10?

A
B
C
D
Test Your Knowledge

On a standardized intelligence test with a Mean of 100 and a Standard Deviation of 15, approximately what percentage of the population scores between 85 and 115?

A
B
C
D
Test Your Knowledge

If a newly developed college admissions test accurately predicts students' first-year college grade point averages (GPA), the test demonstrates high:

A
B
C
D