Introduction
Ask a class why their measurements didn’t all match and someone will say “human error.” It’s the standard answer, and it’s almost always wrong or at least, it’s a tiny slice of the real answer dressed up as the whole thing.
Take a straightforward case. Thirty students measure the same steel bar with the same caliper and get thirty slightly different numbers. Nobody was careless. Nobody misread the scale. The variation came from somewhere, and “human error” only accounts for a fraction of it.
The caliper’s own calibration was off by an amount nobody could see. The workshop was four degrees warmer than the reference temperature, so the bar was genuinely longer than it would have been at 20 °C. The bar itself wasn’t perfectly round. Some students squeezed the jaws harder than others.
Every one of those is a source of measurement uncertainty. Recognising them is the actual skill the arithmetic that follows is straightforward, but you can’t calculate an uncertainty for a source you never noticed.
This is a map of where measurement doubt comes from, including the sources that get overlooked most often, and how to work out which ones are worth your attention.

What Counts as a Source (and What Doesn’t)
A source of uncertainty is anything that would make a repeat of your measurement come out differently, or that shifts every measurement away from the true value by an unknown amount.
That definition does a lot of work, so it’s worth pulling apart.
The first half covers things that vary the operator’s judgement, small temperature swings, where exactly you position the part. Repeat the measurement and you’d get a slightly different answer.
The second half covers things that don’t vary but are still wrong in a way you can’t quantify. If your caliper reads 0.02 mm high on every single measurement, repeating it a thousand times won’t reveal that. It’s still a source of uncertainty, and it’s the more dangerous kind.
What isn’t a source
Mistakes. Reading 24.3 when the scale says 34.3, spilling half a sample, using the wrong formula. These are blunders. You don’t widen your uncertainty to cover them you spot them, discard the result, and measure again. Uncertainty describes the doubt that remains when everything was done correctly.
Known, corrected errors. If your calibration certificate says the instrument reads 0.02 mm high and you subtract 0.02 mm from every reading, that error is gone. What remains is the uncertainty in that correction, which is smaller. Correcting a known error converts a big problem into a small one; it doesn’t eliminate it entirely.
That distinction between what you can correct and what you can only account for runs through everything that follows.
The Five Families of Uncertainty Sources
Metrologists usually sort sources into five groups, borrowed from the fishbone diagrams used in quality control. The grouping isn’t sacred, but it’s a good checklist because it stops you fixating on the instrument and forgetting everything else.
Machine the instrument
The obvious family, and rarely the only important one.
Its calibration uncertainty comes from the certificate and is often the single biggest contributor. Resolution limits how finely you can read it. Drift accumulates between calibrations, so an instrument certified eleven months ago is not the instrument that was certified. Then there’s hysteresis (readings differ depending on whether you approached from above or below), zero error, non-linearity across the range, and the instrument’s own repeatability.
Man the operator
Not “human error” in the vague sense, but specific, describable effects.
Parallax when reading an analogue scale from an angle. Reaction time in any hand-timed measurement, typically around 0.2 seconds and partly systematic since most people press late. Measuring force how hard you close a caliper’s jaws changes the reading, and it changes it more on soft materials. Judgement calls like deciding when a titration has reached its endpoint. And variation between operators, which is often larger than variation within one operator.
Material the thing being measured
Consistently underrated, and in testing laboratories frequently the largest source of all.
Surface roughness means the “diameter” depends on whether you happened to land on a peak or a valley. Form error a shaft that isn’t perfectly round gives different answers at different angles. Elastic deformation under measuring force, obvious with rubber, real with aluminium. Inhomogeneity, where your sample isn’t representative of the batch. And for chemical or biological work, moisture content, stability and degradation over time.
Method how you’re doing it
Alignment matters more than students expect. Tilt an instrument by angle θ and you introduce a cosine error, which grows quickly. There’s also Abbe error, where the measuring scale isn’t in line with the thing being measured the reason a caliper is less accurate than a micrometer even when both are perfectly calibrated.
Beyond that: where you chose to measure on a part that isn’t uniform, sampling decisions, approximations in your calculation model, and in chemistry, recovery and interferences.
Medium – the environment
For dimensional work, temperature dominates almost everything else. Steel expands about 11.5 µm per metre per degree; aluminium roughly twice that. Measure a 500 mm aluminium bar five degrees above 20 °C and it’s genuinely 0.06 mm longer than its nominal size that’s not an error, it’s physics, and it’s why metrology labs are temperature-controlled.
Add humidity (matters for polymers, paper and electronics), vibration, air buoyancy for precision weighing, air pressure, lighting, and electromagnetic interference for electrical measurements.
The Sources Almost Everyone Misses
The five families give you a checklist. These are the specific items that slip through it, drawn from the errors that show up most in student lab reports.
The zero. Not checking that a caliper reads zero when closed, or that a balance is tared. A zero offset is systematic, invisible in your data, and completely uncorrected by repeating the measurement.
Thermal soak. A part taken straight off a machine tool, or held in your hand for two minutes, is not at room temperature. Precision inspection requires letting parts sit until they’ve equalised sometimes hours for large components.
The reference standard itself. Your instrument was calibrated against a master, and that master has its own uncertainty. It’s already inside the number on your certificate, but students often treat a calibrated instrument as though it were exact.
Measuring force on soft materials. A caliper measuring a rubber grommet compresses it. The reading is real; the dimension isn’t.
Sample inhomogeneity. In materials testing, food analysis and environmental work, the question “is this sub-sample representative?” often dominates the entire budget and it’s the source most likely to be left out of a student’s list entirely.
Drift since calibration. A certificate is a snapshot. Eleven months later, the instrument has moved.
Operator anchoring. If students can see previous results before recording their own, the data clusters artificially tight. The spread you measure is smaller than the spread that actually exists, which makes your Type A uncertainty look better than it is.
Time. Some quantities genuinely change while you’re measuring them a cooling sample, a settling suspension, a warming room. Your reading is a snapshot of something moving.
Working Out Which Sources Actually Matter
Listing sources is only half the job. If you tried to quantify every possible influence, you’d never finish and it wouldn’t change your answer much anyway.
The reason is that uncertainty components combine as squares, which is brutally unforgiving to small contributors.
Suppose your largest source is 0.010 mm and you’re wondering whether to bother with one worth 0.002 mm a fifth the size:
√(0.010² + 0.002²) = 0.0102 mm
Including it changed the combined uncertainty by 2%. That’s noise. As a rough screening rule, anything smaller than about a fifth of your largest source can be noted and set aside. At a third the size it starts to be worth including, and anything comparable to the largest source definitely is.
What this looks like in practice
Here’s a caliper measuring a 150 mm aluminium bar on a workshop bench, with five sources quantified:
| Source | Family | u (mm) | Share of total |
|---|---|---|---|
| Calibration certificate | Machine | 0.015 | 62% |
| Temperature (± 5 °C) | Medium | 0.010 | 27% |
| Repeatability (10 readings) | Man / Machine | 0.0047 | 6% |
| Resolution | Machine | 0.0029 | 2% |
| Measuring force | Man | 0.0029 | 2% |
| Combined | 0.019 | 100% |
Two things stand out. Calibration and temperature together account for 89% of the total and neither of them is visible while you’re measuring. Meanwhile resolution, the source students name first, contributes 2%.
If you wanted to improve this measurement, buying a caliper with finer resolution would achieve essentially nothing. Getting a better calibration, or moving the work into a temperature-controlled room, would transform it.
That’s the practical payoff of thinking about sources properly. It tells you where to spend effort and money, and it’s usually not where your intuition points.
Random or Systematic? A Different Way to Sort the Same List
The five families tell you where a source comes from. There’s a second classification that tells you what to do about it, and it cuts across all five.
Random sources vary unpredictably from reading to reading repeatability, operator judgement, small environmental fluctuations. They scatter your results around a centre. Averaging helps: take n readings and the uncertainty in the mean falls by √n.
Systematic sources shift every reading in the same direction by the same amount zero error, a calibration offset, a consistent temperature deviation, a cosine error from an instrument you keep tilting the same way. They move your whole cluster off-target.
Here’s the part that matters: repeating a measurement does absolutely nothing for a systematic source. You can take five hundred readings, watch your standard deviation shrink beautifully, and still be 0.02 mm off because the instrument reads high. Your uncertainty statement will look impressively small and your answer will still be wrong.
Systematic sources have to be handled differently corrected using a calibration certificate where possible, or estimated and included in the budget where not. This is also why a precise-looking result can be badly inaccurate, and why calibration exists as a discipline at all.
Most real sources sit somewhere between the two. Reaction time is a good example: partly random scatter, partly a consistent tendency to press late.
Frequently Asked Questions
What are the main sources of measurement uncertainty?
They’re usually grouped into five families: the instrument (calibration, resolution, drift, hysteresis), the operator (parallax, measuring force, judgement, reaction time), the material being measured (surface roughness, form error, deformation, inhomogeneity), the method (alignment, sampling, model approximations), and the environment (temperature, humidity, vibration, buoyancy). For dimensional measurement, calibration and temperature typically dominate.
Is human error a source of measurement uncertainty?
Partly, but the phrase does more harm than good. Specific operator effects parallax, measuring force, reaction time, endpoint judgement are genuine sources you can quantify. Outright mistakes like misreading a scale are not uncertainty at all; those results should be discarded and the measurement repeated.
Why is temperature such a big source of uncertainty?
Because materials physically change size with temperature, and the effect scales with length. Steel expands roughly 11.5 µm per metre per degree Celsius, aluminium about twice that.
On a 500 mm aluminium part measured 5 °C above the 20 °C reference, that’s a real change of about 0.06 mm larger than most other uncertainty sources combined, which is why precision measurement rooms are temperature-controlled.
How do I know which sources to include in an uncertainty budget?
Include anything comparable in size to your largest source. Because components combine as squares, a source one-fifth the size of the biggest one changes the combined result by around 2% small enough to note and set aside. A source one-third the size contributes about 5% and is usually worth including.
What’s the difference between a random and a systematic source?
Random sources scatter readings unpredictably around a central value and are reduced by averaging more measurements. Systematic sources shift every reading in the same direction and are not reduced by averaging at all — they must be corrected using calibration data or estimated and included in the budget.
Does taking more readings reduce measurement uncertainty?
Only the random part of it, and with diminishing returns since the improvement goes as √n. Going from 10 readings to 40 halves that component. It has no effect whatsoever on systematic sources such as calibration offsets, resolution limits or temperature deviations.
Which source of uncertainty is usually the largest?
It depends heavily on the measurement. In dimensional metrology it’s typically calibration uncertainty or temperature. In chemical and materials testing it’s often sample inhomogeneity or the reference material. In hand-timed physics experiments it’s usually operator reaction time. Building a budget is how you find out rather than guess.
Conclusion
The instinct when a measurement misbehaves is to blame the instrument or the person holding it. Usually both are innocent.
Doubt enters a measurement from five directions the machine, the operator, the material, the method and the environment and the ones doing the real damage tend to be invisible while you work. A calibration offset you can’t see. A room four degrees too warm. A part that isn’t quite round. None of these announce themselves, and none of them are fixed by measuring more carefully.
What makes this worth learning properly is that identifying sources changes what you do. Once you’ve seen that calibration and temperature account for nearly 90% of an uncertainty budget, buying a finer instrument stops looking like an improvement and starts looking like a waste of money. The list of sources isn’t an academic exercise; it’s a diagnosis.
And if you take one thing from the random-versus-systematic split: repeating a measurement fixes only half your problems. The other half needs a calibration certificate.


