1.5 Probability of Detection, Reliability, and Technique Validation
Key Takeaways
- ASME Section V, Article 7 states that MT sensitivity is greatest for surface discontinuities and diminishes rapidly with increasing subsurface depth, which makes MT capability inherently probabilistic rather than absolute.
- A POD curve plots detection probability against flaw size; a50 is the 50% detection size, a90 is a point estimate of the 90% detection size, and a90/95 adds a 95% statistical confidence bound and is the value damage-tolerance analysts use.
- MIL-HDBK-1823 treats conventional MT as hit/miss data because a magnetic particle indication has no calibrated amplitude, and hit/miss analysis needs on the order of 60 characterized flawed targets to bound a90/95.
- A Ketos ring or QQI shim check is a system performance verification against a known baseline, not a POD study and not a procedure qualification; the three prove different things at different frequencies.
- A defensible MT technical demonstration uses specimens with independently documented real cracks rather than notches, runs blind where possible, and records hits, misses, and false calls together.
1.5 Probability of Detection, Reliability, and Technique Validation
The current ASNT NDT Level III Magnetic Particle Testing exam topics place "assist in MT POD studies (e.g., reliability, sensitivity)" and "prepare MT technical demonstrations" inside MT Principles and Theory. That placement is deliberate: a Level III is expected to reason about MT as a statistical detection process, not as a pass/fail ritual. A technician asks "did the indication form?" A Level III asks "what fraction of cracks of this size, in this material, under this technique, would form an indication that this inspector would call?"
Why MT Capability Is a Probability, Not a Guarantee
ASME Section V, Article 7 states the physical boundary plainly: "the sensitivity is greatest for surface discontinuities and diminishes rapidly with increasing depth of subsurface discontinuities below the surface." That single sentence is a capability statement — MT is a surface and near-surface method whose detection rate falls off steeply with depth.
Detection is probabilistic because every link in the chain varies:
| Variable Class | Specific Sources of Variation | Effect on Detection |
|---|---|---|
| Flaw character | Length, depth, width (tightness), orientation to flux, whether the crack is open or closed by compressive residual stress or smeared by grinding | A tight, smeared, shallow crack may produce no visible leakage field even at correct field strength |
| Field | Tangential field strength (commonly controlled to 30 to 60 G / 2.4 to 4.8 kA/m), direction, waveform, continuous vs. residual application | A flaw within 30 degrees of parallel to the flux may be missed entirely |
| Medium | Dry vs. wet, visible vs. fluorescent, particle size distribution, bath concentration, contamination | Concentration drift inside the allowed band still moves detection rate |
| Viewing | UV-A irradiance, ambient white light, dark adaptation, filter condition, inspector visual acuity | Fluorescent contrast is the limiting factor for the smallest indications |
| Human factors | Scan speed, search pattern discipline, time on task, expectation and prior information, fatigue | Miss rate rises sharply with lot size and shift length |
Because all of these vary, the honest expression of method capability is a POD curve: detection probability plotted against flaw size, rising from near zero at small sizes toward an asymptote below 100%.
Reading a POD Curve: a50, a90, and a90/95
Three points on the curve carry contractual weight:
- $a_{50}$ — the flaw size detected 50% of the time. It is the midpoint of the curve and a poor engineering basis, because half the flaws of that size are missed.
- $a_{90}$ — the flaw size detected 90% of the time. A point estimate from one data set.
- $a_{90/95}$ — the flaw size for which detection is 90% probable at 95% statistical confidence. This is the value damage-tolerance analysts use to set the assumed initial flaw size for inspection intervals, because it accounts for the finite size of the demonstration sample.
Level III reasoning trap: a demonstration that found 9 of 10 cracks does not establish $a_{90}$. With 10 targets the 95% lower confidence bound on a 90% observed rate sits near 60%, which is why capability demonstrations require dozens of characterized flaws rather than a handful.
Two Analysis Models
MIL-HDBK-1823, Nondestructive Evaluation System Reliability Assessment, is the reference most aerospace customers invoke for POD demonstrations, and it distinguishes two data types:
- Hit/miss analysis — each target is recorded only as found or not found. This is the natural model for MT, because a magnetic particle indication has no calibrated amplitude. Hit/miss analysis is statistically inefficient, so the handbook guidance calls for a substantially larger target set (on the order of 60 flawed sites) to support a defensible $a_{90/95}$.
- Signal response ("$\hat{a}$ versus $a$") analysis — each target yields a measured signal amplitude that is regressed against flaw size. Fewer targets are needed, but the method requires a quantitative signal, which eddy current and ultrasonic testing provide and conventional MT does not.
A Level III should also insist that the demonstration record the false call rate alongside POD. A technique tuned to maximum sensitivity by over-magnetizing will show high POD and an unusable false-call rate from background furring; reporting one without the other is meaningless.
System Performance Verification Is Not a POD Study
This is the distinction most often confused on the examination:
| Activity | What It Proves | Typical Tools | Frequency |
|---|---|---|---|
| System performance check | That today's equipment, bath, and lighting reproduce a known baseline response | Ketos/Betz ring (ASME V T-766, SAE AS 5282), QQI notched shims (SAE AS 5371) | Daily, per shift, or weekly per the governing document |
| Procedure qualification / technique demonstration | That this written technique detects the smallest rejectable flaw on representative hardware | Specimens with documented real cracks; ASME V T-721.2 and Mandatory Appendix I for coated surfaces | On procedure issue and after any essential-variable change |
| POD study | The statistical capability of the method as a function of flaw size | Large sets of characterized natural flaws, blind trials, multiple inspectors | Programme-level; supports damage-tolerance and inspection-interval decisions |
A QQI shim indication proves the field is adequate at the shim. It says nothing about whether a 0.020 in. deep tight fatigue crack in the part's fillet radius would be found.
Designing an MT Technical Demonstration
"Prepare MT technical demonstrations" is an explicit Level III task, and demonstrations are what auditors, customers, and design authorities actually witness. A defensible demonstration package contains:
- Objective statement — the specific claim being demonstrated (for example, "detection of surface-breaking cracks of 0.060 in. length or greater through 0.004 in. of epoxy primer using the AC yoke technique").
- Representative specimens — same alloy, heat treatment, surface finish, and geometry as production hardware, containing documented flaws whose size has been established by an independent means (fracture, sectioning, replication, or a calibrated secondary method). Notches are not cracks; a notch demonstration over-states capability.
- Controlled process — the exact written procedure under demonstration, with field verification, lighting readings, bath concentration, and current settings recorded as data, not as assumptions.
- Blind conditions where possible — the inspector should not know flaw locations or counts. A demonstration in which the operator is pointing at a known crack proves only that the crack is visible.
- Recorded outcome — hits, misses, false calls, and the identity and qualification of the personnel performing and witnessing the work. ASME Section V Mandatory Appendix I requires exactly this kind of record for the coated-surface AC yoke demonstration: identification of the qualification, the personnel performing and witnessing, equipment and materials, illumination levels, and results.
- Disposition — an explicit statement of what the demonstration does and does not qualify, signed by the Level III.
Worked Level III Scenario: An Over-Claimed Capability
Scenario: A turbine overhaul shop's MT procedure states that wet fluorescent MT "detects all surface cracks greater than 0.030 in. long." The customer's damage-tolerance engineer asks for the $a_{90/95}$ value supporting a 2 000-cycle inspection interval. The shop offers its daily Ketos ring logs (three holes at 1 400 A for the last two years) and a photograph of a 0.035 in. crack found on a QQI-verified setup.
Level III analysis:
- The claim is unsupported. "Detects all" is an absolute statement no NDE method can support. The procedure language must be corrected to a probabilistic statement or removed.
- The Ketos ring evidence is the wrong evidence. Ring performance is a system baseline check under ASME V T-766; it demonstrates day-to-day reproducibility of the bath and power supply, not flaw-size capability on turbine hardware.
- The single photograph is anecdote, not data. One hit provides no confidence bound.
- Correct path forward. Commission a hit/miss POD demonstration on retired blades containing characterized service fatigue cracks, with a target set large enough to bound $a_{90/95}$, multiple certified inspectors, blind presentation, and recorded false calls. Until that data exists, the Level III must tell the customer that the shop can support a procedure qualification claim but not an $a_{90/95}$ value — and must not let an interval be set on a number the data cannot carry.
- Document the limitation. The honest interim statement is that the technique has been qualified by demonstration to detect the smallest rejectable indication defined in the acceptance criteria, with capability below that size uncharacterized.
A customer asks a Level III for the a90/95 crack length supporting a wet fluorescent MT inspection interval. What does that value actually represent?
Why is hit/miss analysis, rather than signal response (a-hat versus a) analysis, the appropriate POD model for conventional magnetic particle testing?
A facility offers two years of passing daily Ketos (Betz) ring records as evidence of its MT probability of detection on production forgings. What is the correct Level III response?
Which specimen choice most seriously over-states the capability shown by an MT technique demonstration?