14.1 Standardized Tests: Achievement, Aptitude, Ability & Norm vs. Criterion

Key Takeaways

  • A standardized test uses the same content, administration, scoring, and interpretation for every taker.

  • Achievement tests measure what students have learned, aptitude tests predict success in a specific future area, and ability tests measure general reasoning.

  • Individually administered ability tests such as the WISC-V report standard scores with a mean of 100 and a standard deviation of 15.

  • Norm-referenced tests compare a student with a norm group; criterion-referenced tests compare performance with fixed standards.

  • To compare a class with students nationally, use a norm-referenced test; to check mastery of standards, use a criterion-referenced test.

Last updated: September 2026

The Psychometric Landscape in Secondary Education

Secondary educators in grades 7 through 12 regularly interact with a broad spectrum of assessment data. On any given week, a high school teacher may review classroom quiz results, district interim benchmark assessments, state accountability reports, Advanced Placement (AP) exam breakdowns, and standardized college-readiness scores such as the PSAT, SAT, or ACT. Misinterpreting these quantitative metrics leads to erroneous instructional decisions, flawed academic tracking, and distorted parent-teacher communications.

To make sound instructional adjustments and excel on the Praxis PLT: Grades 7-12 (5624) exam, teachers must possess a rigorous understanding of educational psychometrics. This begins with distinguishing between norm-referenced and criterion-referenced measurement frameworks, mastering raw versus derived score conversions, interpreting normal distribution curves, avoiding the widespread pitfalls of grade-equivalent scores, and communicating testing data accurately and empathetically to adolescents and their families.


Standardized Tests: Achievement, Aptitude, and Ability

ETS asks candidates to understand the types and purposes of standardized tests (achievement, aptitude, and ability), to recognize the data each provides, and to distinguish norm-referenced from criterion-referenced scoring. A test is standardized when every taker faces the same content, administration conditions, scoring rules, and score interpretation. Standardized tests can be norm-referenced, criterion-referenced, or both.

TypeWhat it measuresExamplesData it providesTypical uses
AchievementWhat a student has learned in a content areaState assessments, end-of-course exams, AP exams, Iowa Assessments, NAEP (reported for groups, not individuals)Scale scores, performance levels, percentile ranks, subscores by standardAccountability, program evaluation, placement, instructional planning
AptitudePotential to learn or perform in a specific future area; designed to predictThe ASVAB for military occupations, differential aptitude batteries, language or music aptitude testsScores and percentile ranks tied to predicted successCareer guidance, placement, selection
Ability (cognitive ability or intelligence)General reasoning ability: verbal, quantitative, and nonverbalIndividually given tests such as the WISC-V or Stanford-Binet (given by school psychologists); group tests such as the Cognitive Abilities Test (CogAT)Standard scores (often mean 100, standard deviation 15), index scores, percentile ranksSpecial education evaluations, gifted identification

The categories overlap, because every test measures some learned skills. College admission tests such as the SAT were once called aptitude tests but today describe themselves as measures of college and career readiness. On the PLT, match the purpose: past learning in a course points to achievement, prediction of success in a specific future domain points to aptitude, and general cognitive capacity points to ability.

To compare a class's performance with students nationally, a teacher needs a norm-referenced test; to find out whether students mastered specific standards, a criterion-referenced test.

Norm-Referenced vs. Criterion-Referenced Assessment Paradigms

All educational assessments derive meaning by comparing a student's performance against a reference frame. Psychometrics categorizes tests into two fundamental paradigms based on this reference frame:

Norm-Referenced Tests (NRT): Comparative Relative Standing

A Norm-Referenced Test (NRT) is designed to determine an individual student's relative standing in comparison to a clearly defined normative sample of peers who took the identical assessment under standardized conditions.

  • Core Purpose: To rank test-takers along a continuum of achievement, distinguish between high and low performers, and identify extreme percentiles (e.g., qualifying students for gifted and talented programs or special education interventions).
  • Underlying Statistical Distribution: NRT items are intentionally selected to produce a wide spread of scores that conforms to a symmetrical normal distribution (bell curve). Psychometricians intentionally discard test items that almost all students answer correctly or that almost all students miss, because items with near-universal pass rates fail to discriminate among test-takers.
  • Typical Secondary Examples: The SAT, ACT, PSAT, Iowa Assessments, TerraNova, and Wechsler Intelligence Scale for Children (WISC).

Criterion-Referenced Tests (CRT): Absolute Standards-Based Mastery

A Criterion-Referenced Test (CRT) measures an individual student's performance against predetermined, explicit learning objectives, curricular standards, or performance benchmarks, completely independent of how other students perform.

  • Core Purpose: To evaluate whether a student has mastered specific content knowledge, critical skills, or performance standards (e.g., Can the student solve quadratic equations by factoring? Can the student identify faulty logic in an editorial?).
  • Underlying Statistical Distribution: CRTs do not aim to create a bell curve. Theoretically, if instruction is extraordinarily effective, 100% of students could achieve an "Advanced" or "Proficient" rating. Cut scores (performance standards) are established by panels of educators and psychometricians using systematic standard-setting methodologies (such as the Angoff method).
  • Typical Secondary Examples: State standards-based accountability exams, Advanced Placement (AP) subject examinations (scored 1 to 5 based on cut scores), career and technical certification tests, driver's licensing exams, and teacher-made end-of-unit chapter exams.

Psychometric Comparison: NRT vs. CRT

Assessment DimensionNorm-Referenced Tests (NRT)Criterion-Referenced Tests (CRT)
Primary ObjectiveRank students; determine relative competitive percentile standing among peers.Measure individual mastery of explicit, predefined curriculum standards.
Reference FrameworkThe normative peer group (national, state, or demographic standardization sample).Predetermined criterion, learning objective, cut score, or performance rubric.
Item Selection CriteriaItems with moderate difficulty and high discrimination to spread scores across a bell curve; items answered correctly by all are removed.Items directly aligned to content standards, regardless of whether 10% or 95% of students master them.
Score Reporting UnitsPercentile ranks, stanines, z-scores, T-scores, normal curve equivalents (NCEs).Raw scores, percentage correct, performance levels (Below Basic, Proficient, Advanced), pass/fail.
Instructional UtilityLimited for day-to-day lesson adjustments; useful for long-term placement, screening, and program eligibility.High for day-to-day instructional adjustments; pinpoints exact standard deficits requiring immediate reteaching.
Classroom ScenarioA counselor evaluates 8th-grade PSAT percentile ranks to recommend honors track placement.A chemistry teacher administers a stoichiometry test to determine if students can balance chemical equations before moving to thermodynamics.

Test Your Knowledge

A secondary curriculum team is reviewing assessment data from two different testing instruments: an end-of-course state biology exam used to verify student competency for graduation, and a national aptitude exam used to identify candidates for a regional STEM honors program. Which statement correctly identifies the psychometric design and primary purpose of these two assessments?

A

The state biology exam is norm-referenced to rank students along a bell curve, while the national aptitude exam is criterion-referenced to verify mastery of state science standards

B

The state biology exam is criterion-referenced to evaluate mastery against predefined state standards, while the national aptitude exam is norm-referenced to compare performance against a normative peer group

C

Both assessments are norm-referenced tests designed to ensure equal distribution across stanines 1 through 9

D

Both assessments are criterion-referenced tests utilizing identical standard cut scores to determine pass/fail status

Test Your Knowledge

As part of a special education evaluation, a school psychologist gives an individually administered test that yields a full-scale score with a mean of 100 and a standard deviation of 15. Which type of standardized test is this?

A

An achievement test of course content

B

An aptitude test for a specific occupation

C

A criterion-referenced end-of-course exam

D

An ability (cognitive or intelligence) test

Test Your Knowledge

Which test is designed primarily to predict success in a specific future area rather than to measure what a student has already learned in a course?

A

A state end-of-course biology exam

B

The Armed Services Vocational Aptitude Battery (ASVAB)

C

A teacher-made unit test on the Civil War

D

A state mathematics assessment aligned to grade-level standards

Sections you finish are checked off in the contents.