2.3 Accuracy, Precision & Measurement Error

Key Takeaways

  • Accuracy is how close a measurement is to the true value; precision is how close repeated measurements are to each other - you can have one without the other
  • Darts clustered tightly but off the bullseye are precise but not accurate; darts scattered around the bullseye are accurate on average but not precise
  • Common error sources include the limits of the instrument, human reaction time on a stopwatch, and environmental factors like wind or temperature
  • Repeating a measurement and taking the average makes a result more reliable because random errors partly cancel out
  • An anomalous result - one that sits far from the pattern - should be checked, repeated, and left out of the average if it is clearly a mistake
Last updated: August 2026

Accuracy and Precision Are Not the Same Thing

These two words are used loosely in everyday speech, but in science they mean different things - and telling them apart is an official Paper I topic, where the framework asks students to differentiate between accuracy and precision in experiments.

  • Accuracy describes how close a measurement is to the true value. If a block of cheese really has a mass of 250 g and your balance reads 249 g, that is an accurate measurement.
  • Precision describes how close repeated measurements are to each other. If you weigh the cheese three times and get 262 g, 263 g and 262 g, those readings are very precise - they agree beautifully - but they are not accurate, because they are all about 12 g too heavy (perhaps the balance was not zeroed).

The classic picture is a dartboard (target) analogy. Imagine four players each throwing three darts:

  1. Darts tight together on the bullseye - accurate and precise. This is the goal.
  2. Darts tight together but off in one corner - precise but not accurate. Something is consistently wrong, like an unzeroed balance. This kind of consistent, one-directional error is called a systematic error.
  3. Darts scattered widely but centred roughly around the bullseye - accurate on average but not precise. Each throw wanders randomly; this is random error.
  4. Darts scattered in a far corner - neither accurate nor precise.

A useful memory hook: precision is about agreement with each other, accuracy is about agreement with the truth. A broken watch that always reads 3:15 is very precise - and permanently inaccurate.

Where Measurement Error Comes From

No measurement is perfect. Every reading carries some measurement error, the gap between the reading and the true value. Knowing the main sources helps you answer "why might these results differ?" questions:

  • Instrument limits. A ruler marked in millimetres cannot honestly report hundredths of a millimetre. Every instrument has a smallest division, and any reading is uncertain by about half of that division. Choosing a finer instrument (a 10 mL cylinder instead of a beaker) reduces this error.
  • Human reaction time. Timing a sprint with a handheld stopwatch involves pressing start and stop by hand, and human reaction time is roughly 0.1-0.3 seconds. Two people timing the same race will rarely agree to the tenth of a second. Automatic timing gates remove this error - which is why athletics uses them.
  • Environmental factors. A metal ruler expands slightly on a hot day; a breeze moves a pendulum; a draughty windowsill cools one beaker faster than another. Conditions that change between trials quietly shift the readings.
  • Reading mistakes. Parallax error, misread scales and skipped zeroing all count as avoidable human errors rather than unavoidable limits.

A mistake (recording 68 instead of 86) is different from an error (an unavoidable limit of the method). Mistakes can be eliminated by care; errors can only be reduced.

Repeating Measurements and Averaging

Because random error pushes readings sometimes high and sometimes low, taking the same measurement several times and calculating the average (mean) gives a result more likely to sit near the true value - the random wobbles partly cancel each other out. Scientists say this makes the result more reliable: if someone else repeated your experiment the same way, they would probably get a similar answer.

To average, add the readings and divide by how many there are. Timing a marble rolling down a ramp five times might give 2.4 s, 2.6 s, 2.5 s, 2.5 s and 2.5 s: the sum is 12.5 s, divided by 5 gives an average of 2.5 s. Three trials is usually considered a sensible minimum in school science; five is better when time allows.

Anomalous Results

In any set of repeats, one reading sometimes sits oddly far from the others - an anomalous result (also called an outlier). Suppose four trials give 2.4 s, 2.5 s, 4.9 s and 2.5 s. The 4.9 s reading almost certainly came from a fumbled stopwatch press or the marble catching on the ramp. The correct response has three steps:

  1. Notice it - look for values that break the pattern.
  2. Investigate and repeat - if possible, do that trial again properly.
  3. Leave it out of the average - including 4.9 s would drag the mean up to about 3.1 s, which represents none of the real trials well. With it excluded, the average is a much more honest 2.47 s.

What you must never do is quietly delete a result simply because you dislike it, or keep one because it helps your favourite answer. The rule is: an outlier may be excluded only when you can point to a concrete reason it went wrong.

Significant Figures - a Gentle Introduction

Senior papers occasionally touch on significant figures, the digits in a measurement that genuinely carry information. The gentle version is this: your answer should not pretend to be more precise than your instruments. If your stopwatch reads to tenths of a second and your average works out to 2.466666... s, reporting "2.466666 s" is false precision - the extra digits are invented detail. Rounding to 2.5 s (or 2.47 s at most) matches what the equipment can really tell you. Zeros at the start of a number (0.35 g) are placeholders, not significant; zeros trapped between other digits (3.05 g) or after a decimal point (2.50 s) do count. You will not need heavy significant-figure arithmetic for ICAS, but recognising an over-precise answer - and avoiding it in your own science reports - is exactly the kind of judgement the Reasoning and Problem Solving questions reward.

The big picture of this section: measuring well is not about getting lucky once. It is about choosing the right tool, reading it carefully, repeating the measurement, averaging honestly, and treating strange results with suspicion rather than wishful thinking. Those habits are what ICAS Observing and Measuring questions - and real scientists - are looking for.

Test Your Knowledge

A student weighs the same apple four times on a kitchen scale and gets 212 g, 213 g, 212 g and 213 g. The apple's true mass is 198 g. Which description of these readings is correct?

A
B
C
D
Test Your Knowledge

Two students time the same 50 m sprint with handheld stopwatches and get noticeably different times. What is the most likely source of the error, and the best way to reduce it?

A
B
C
D
Test Your Knowledge

A group measures the time for a sugar cube to dissolve in five trials: 58 s, 61 s, 59 s, 95 s and 60 s. What should they do to find the best value?

A
B
C
D