Posted on Leave a comment

repeatability vs reproducibility: What’s the Difference?

repeatability vs reproducibility

A quality engineer runs a check on a shop floor. The same inspector measures the same bracket ten times with the same gauge, and the readings sit within 0.01 mm of each other. Excellent, she notes. The gauge is tight.

The next week, three inspectors on three shifts measure that same bracket with that same gauge. The spread is now 0.04 mm four times worse. Nothing about the gauge changed. Nothing about the bracket changed.

What changed was the conditions, and that’s the whole distinction. Repeatability is how consistent a measurement is when you hold everything steady. Reproducibility is how consistent it is when you deliberately let things vary different people, different equipment, different days, sometimes different laboratories entirely.

Both are measures of precision. Both matter. But they answer different questions, and confusing them leads to a specific and expensive mistake: trusting a measurement system because it looked consistent under conditions that never occur in real production.


The Difference in One Sentence

Repeatability is precision under the most similar conditions possible. Reproducibility is precision under deliberately different ones.

The item being measured stays the same in both cases. That’s the fixed point. Everything else is what separates the two.

What variesRepeatabilityReproducibility
The item measuredSameSame
OperatorSame personDifferent people
InstrumentSame gaugeDifferent gauges (usually)
Laboratory / locationSameDifferent
Time between readingsShort same sessionDays, weeks, or longer
EnvironmentSameMay differ
MethodSameSame

Notice that method stays constant in both. If you change the procedure itself, you’re no longer measuring reproducibility you’re comparing two different methods, which is a separate exercise.

The practical consequence: reproducibility is always at least as large as repeatability, and normally larger. It has to be, because reproducibility conditions include all the variation present under repeatability conditions plus the between-operator and between-lab variation on top. If you ever calculate a reproducibility smaller than your repeatability, something has gone wrong with the study.

Why the distinction earns its keep

A single technician measuring one part repeatedly is not how measurement works in a real factory or a real laboratory. Parts get inspected across three shifts, by different people, on different gauges, over months.

Repeatability tells you the best your measurement system can possibly do. Reproducibility tells you what it actually does. Optimising for the first while ignoring the second is how a company ends up with an inspection process that passes its calibration checks and still ships bad parts.


The Third Tier Nobody Mentions

Most articles stop at two. The international standard for precision, ISO 5725, actually defines three, and the middle one is the tier most real laboratories operate in day to day.

Repeatability conditions – same operator, same equipment, same lab, short interval. The narrowest possible.

Intermediate precision conditions – same laboratory, but something deliberately varies: a different day, a different analyst, a different instrument, a different batch of reagent. You’re still in one building with one method.

Reproducibility conditions – different laboratories entirely. Different everything, same method.

Intermediate precision is what you’re actually measuring when you compare Monday’s results with Thursday’s, or when a second analyst in the same lab runs the same test. Calling that “reproducibility” is technically wrong and will get flagged in an audit, because true reproducibility requires a multi-laboratory study the kind run through proficiency testing schemes and collaborative trials.

If you write method validation reports or work toward accreditation, this is a distinction worth getting right. For a student lab report, knowing the tier exists is usually enough.


How Manufacturing Measures Both: Gauge R&R

In manufacturing quality, repeatability and reproducibility are measured together in a study called Gauge R&R, part of Measurement System Analysis (MSA). The R&R is literally the two words in this article’s title.

A word of warning first, because it catches people out. Gauge R&R uses a narrower definition of reproducibility than metrology does. In a standard R&R study, all operators use the same gauge, so reproducibility means purely operator-to-operator variation. The ISO 5725 sense different labs, different equipment is broader. Same word, two scopes, depending on which room you’re standing in.

The terms

A typical study has three operators measure ten parts three times each, and the variation gets split up:

  • EV (Equipment Variation) = repeatability the gauge’s own scatter
  • AV (Appraiser Variation) = reproducibility the operator-to-operator scatter
  • GRR = the two combined
  • PV (Part Variation) = the genuine spread between the parts themselves
  • TV (Total Variation) = everything together

They combine in quadrature, the same root-sum-of-squares used everywhere else in uncertainty work:

GRR = √(EV² + AV²)
TV = √(GRR² + PV²)
%GRR = (GRR ÷ TV) × 100

A worked study

Say a study on a machined shaft diameter returns:

EV (repeatability) = 0.012 mm
AV (reproducibility) = 0.009 mm
PV (part variation) = 0.085 mm

Combining the measurement system’s own variation:

GRR = √(0.012² + 0.009²) = √0.000225 = 0.015 mm

Against total variation:

TV = √(0.015² + 0.085²) = √0.00745 = 0.0863 mm
%GRR = (0.015 ÷ 0.0863) × 100 = 17.4%

There’s one more number worth calculating, the number of distinct categories how many genuinely different groups your gauge can separate the parts into:

NDC = 1.41 × (PV ÷ GRR) = 1.41 × (0.085 ÷ 0.015) = 8
 Bar chart breaking down gauge R&R into equipment variation, appraiser variation and part variation

What the Numbers Actually Tell You to Fix

A percentage on its own is just a grade. The useful part is the diagnosis underneath it.

Is the system good enough?

The widely used AIAG thresholds:

%GRRVerdict
Under 10%Acceptable
10–30%Conditionally acceptable — depends on cost, criticality, application
Over 30%Not acceptable — the measurement system needs work

Alongside that, NDC should be at least 5. Below that, the gauge can’t reliably distinguish one part from another, and using it for process control is close to guesswork.

The study above came out at 17.4% with an NDC of 8. Conditionally acceptable usable, worth improving, and not something to sign off without a conversation about what the part does.

Which one is bigger?

This is the question that turns a number into an action.

If EV is larger than AV, the gauge is the problem. Look at its resolution, its condition, the fixturing, whether it’s the right instrument for the tolerance you’re chasing. Training operators harder won’t help they’re already more consistent than the tool.

If AV is larger than EV, the people are the problem or more fairly, the procedure is. Different operators are doing something differently, which usually means the work instruction is ambiguous, the part is awkward to locate, or nobody agreed on where exactly to measure. Buying a better gauge won’t fix this at all.

In the example above, EV (0.012) exceeds AV (0.009), so the gauge deserves attention before the training budget does.

That single comparison is the most practically valuable output of an R&R study, and it’s the reason the two concepts are worth separating rather than lumping together as “precision”.

The limits used in laboratories

Testing labs express the same idea differently, using ISO 5725’s repeatability and reproducibility limits:

r = 2.8 × sr (repeatability limit)
R = 2.8 × sR (reproducibility limit)

Where sr and sR are the respective standard deviations. The 2.8 comes from 1.96 × √2 the spread you’d expect between two results at 95% confidence.

The interpretation is refreshingly concrete. If a method has sr = 0.05 g, then r = 0.14 g, meaning two results produced by the same analyst on the same day should differ by less than 0.14 g about 95% of the time. A bigger gap than that is a signal to investigate rather than average.


Where People Get These Two Wrong

Treating repeatability as accuracy. Both repeatability and reproducibility are measures of precision how tightly results cluster. Neither says anything about whether the cluster is in the right place. A gauge with a 0.5 mm zero offset can be beautifully repeatable and consistently wrong by half a millimetre. Precision is about scatter; accuracy needs calibration against a traceable standard.

Using “reproducibility” for a same-lab study. If nothing left the building, that’s intermediate precision. True reproducibility means different laboratories.

Assuming the Gauge R&R meaning is universal. In an R&R study, reproducibility means operator-to-operator variation with one shared gauge. In ISO 5725 it means between-laboratory variation. Both are correct within their own framework just be clear which you’re using before quoting a number to someone.

Confusing it with the reproducibility crisis in research. When scientists talk about a study being reproducible, they mean an independent team can run the experiment and reach the same conclusion. That’s related in spirit but it’s a much broader idea, covering study design, data handling and analysis not a statistical measure of measurement spread.

Running an R&R study on parts that are too similar. If your ten parts barely differ from each other, PV comes out small, and %GRR looks terrible even with a perfectly good gauge. The parts in a study should span the range of real production variation, or the percentage is meaningless.

Forgetting that operators know they’re being watched. People measure more carefully during a study than during a Tuesday afternoon shift. Randomise the order, keep parts anonymous, and don’t let operators see their previous readings — otherwise your reproducibility number is better than reality.


Frequently Asked Questions

What is the difference between repeatability and reproducibility?

Repeatability is the variation you get when the same person measures the same item with the same instrument in the same place over a short time. Reproducibility is the variation when those conditions deliberately change different operators, different instruments, different laboratories or different days. The item measured and the method stay the same in both.

Which is always larger, repeatability or reproducibility?

Reproducibility, or at minimum equal. Reproducibility conditions contain all the variation present under repeatability conditions plus additional between-operator and between-laboratory variation. A result showing reproducibility smaller than repeatability points to a problem with the study design or the calculation.

What does Gauge R&R measure?

It measures both properties together and splits measurement system variation into equipment variation (EV, the repeatability) and appraiser variation (AV, the reproducibility). These combine as GRR = √(EV² + AV²), and the result is expressed as a percentage of total variation or of the part tolerance.

What is an acceptable %GRR value?

Under 10% is generally acceptable, 10–30% is conditionally acceptable depending on how critical the measurement is and what improvement would cost, and above 30% means the measurement system is not fit for purpose. The number of distinct categories should also be 5 or more.

Does high repeatability mean a measurement is accurate?

No. Repeatability describes only how tightly readings cluster together. An instrument with a systematic offset can produce highly repeatable readings that are all consistently wrong by the same amount. Accuracy requires calibration against a traceable reference standard, which is a separate exercise from an R&R study.

What is intermediate precision?

It’s the middle tier defined in ISO 5725, sitting between repeatability and reproducibility. It covers variation within a single laboratory when something deliberately changes a different day, analyst, or instrument while the method and location stay the same. Most routine laboratory precision estimates are actually intermediate precision rather than true reproducibility.

How many operators and parts should a Gauge R&R study use?

The common convention is three operators, ten parts and three trials each, giving 90 measurements. The parts should span the range of variation seen in real production, since parts that are too similar to each other will make the measurement system look worse than it is.


Conclusion

The distinction comes down to a single question: what were you allowed to change?

Hold everything steady and you’re measuring repeatability the floor, the best your system can manage. Let the operator, the gauge, the day and the location vary, and you’re measuring reproducibility, which is closer to what your process actually experiences. The gap between those two numbers is where most measurement problems live, and it’s invisible if you only ever run the narrow study.

For anyone working in quality, the sharpest tool here isn’t the %GRR percentage at all it’s comparing EV against AV. A gauge problem and a procedure problem produce similar-looking bad results and need completely different responses, and that one comparison tells you which you’re dealing with before you spend anything.

For students, the thing worth carrying forward is subtler. Neither of these says your measurement is right. A perfectly repeatable, perfectly reproducible measurement system can be wrong by a mile, consistently, in a way no amount of repetition will ever reveal. That’s what calibration is for, and it’s a different conversation entirely.

Leave a Reply