12.4 Sampling Strategies & Generating Findings
Key Takeaways
- AI system populations are usually sampled judgmentally with risk-based priority — high-impact, novel, recently changed, or poorly evidenced systems first, not only the organization’s showcase models — while statistical sampling applies when large homogeneous record populations exist (tickets, logs, approvals).
- Findings evaluate evidence against criteria and classify as conformity, nonconformity, or observation/OFI with clear condition–criteria linkage.
- Working papers and daily team meetings protect consistency, manage missing evidence, and prevent unsupported conclusions.
- Sampling bias (convenience, success-only, guide-steered) is an exam and practice trap that can falsely signal AIMS conformity.
- The benefit of the doubt applies only where evidence is inconclusive: conformity stands until objective evidence proves otherwise, but absent required documented information is a nonconformity, not doubt.
12.4 Sampling Strategies & Generating Findings
Auditor focus: You will almost never audit 100% of AI systems or records. Sampling must be planned, risk-based, and free of success-story bias. Findings must then state facts against criteria—conformity, nonconformity, or observation—supported by working papers strong enough for peer review.
Why sampling matters in AIMS audits
Organizations may run many models, pipeline runs, predictions, and reviews. ISO 19011 treats sampling as procedures applied to less than 100% of a population; sampling risk is concluding differently than a full examination would. AIMS populations include systems/use cases, impact/risk files, release tickets, evaluations, oversight cases, incidents/drift alerts, competence records, and third-party suppliers. A dozen high-impact systems is not sampled like thousands of homogeneous tickets.
Judgmental vs statistical sampling
| Approach | Basis | Best used when | Limitations |
|---|---|---|---|
| Judgmental (non-statistical) | Auditor competence, risk, prior knowledge, anomalies | Heterogeneous AI systems; process walks; selecting systems for deep dive | Conclusions not statistically projected; bias if judgment is weak |
| Statistical | Probability, sample size formulas, random/systematic selection | Large homogeneous record populations (approvals, overrides, training completions) | Needs defined population and attributes; can miss rare high-impact events if not stratified |
Most system-level AIMS sampling is judgmental and risk-based. Statistical methods shine for attributes like “percentage of sampled release tickets with required evaluation attached” once you define the population of releases in the period.
Hybrid pattern (common in practice)
- Judgmentally select high-priority AI systems for end-to-end lifecycle testing.
- Statistically or systematically sample records within those systems (e.g., last 90 days of promotions, random 25 oversight cases).
- Extend if errors appear (enlarge sample, follow related systems sharing the same pipeline).
Risk-based sample selection for AI systems
Prioritize systems that combine impact, uncertainty, and change:
| Priority signal | Why sample |
|---|---|
| High impact / high risk (rights, safety, livelihood, large populations) | Harm magnitude if controls fail |
| New or significantly changed since last audit or Stage 1 | Control design may lag delivery |
| Complex third-party / foundation components | Opacity and supply-chain risk |
| Prior incidents, complaints, or audit NCs | Known weakness areas |
| Weak documentation or missing inventory entries | Potential uncontrolled AI |
| Heavy automation with limited human oversight | Failure modes may be silent |
| Business-critical revenue systems under production pressure | Pressure to bypass gates |
Deliberately avoid a sample composed only of mature showcase systems with dedicated AI governance staff. Include at least one “awkward” system: a shadow IT model brought under AIMS late, a rapidly iterated GenAI assistant, or a vendor model embedded in a SaaS workflow.
Document why each system was selected. Selection rationale is itself audit evidence of a risk-based approach.
Within-system record sampling
Define the population (e.g., promotions in a quarter), choose random/periodic/stratified methods, size samples for frequency and risk, and record population, method, IDs, and pass/fail results. Systemic failure in a small sample (e.g., 4/5 releases missing required impact re-assessment) can support an NC if the condition is fairly characterized—enlarge when results are mixed.
Generating findings
Audit findings result from evaluating evidence against audit criteria. Typical classifications in management system audits:
Conformity
Evidence shows requirements are met for the area sampled. Write positively but precisely: which criterion, which evidence IDs, which period. Avoid “everything is fine” without anchors.
Nonconformity (NC)
Non-fulfillment of a requirement. Good NC statements include:
- Condition — what was found (facts, sample results, IDs)
- Criteria — which requirement (ISO/IEC 42001 clause, Annex A control as applicable, or organizational procedure)
- Evidence — pointers to working papers
- Often consequence — risk or impact of the gap (without inventing harm)
Example shape (illustrative):
Of five production promotions for high-impact model claims-triage-v4 between February and April (tickets CT-220…CT-224), four lacked an updated AI system impact assessment after the intended-use expansion to automatic denial recommendations. Procedure AIA-02 and the organization’s application of impact assessment requirements require reassessment before material use-case changes. Condition indicates operational planning and impact controls are not consistently implemented for this system.
Severity (major/minor) follows the scheme and effect on AIMS outcomes—often refined at closing; in fieldwork, capture facts first.
Observation / Opportunity for Improvement (OFI)
Not an NC, but a noteworthy risk or improvement chance. Do not park clear requirement failures as OFIs to stay “friendly,” and do not inflate OFIs into NCs without a breached requirement.
The Benefit of the Doubt
PECB names this principle explicitly, and it decides a large share of scenario items. The benefit of the doubt means that where evidence is genuinely inconclusive — not absent, not contradicted, merely insufficient to prove nonconformity — the auditor does not record a nonconformity. Conformity stands until objective evidence shows otherwise, because a finding must rest on evidence rather than suspicion.
| Situation | Evidence state | Correct call |
|---|---|---|
| Drift monitoring is demonstrated, but one month's report cannot be retrieved during the visit | Control shown for 11 of 12 months; the single gap is plausibly administrative | Benefit of the doubt on conformity; note it, request the record, consider an OFI |
| An interviewee gives a confused answer, but records and a second interviewee confirm the process | Testimony weak, documentary evidence strong | Conformity — weight the more reliable evidence |
| No AI system impact assessment exists for a high-impact production system | The requirement applies and the evidence is absent | Nonconformity — absence of required evidence is not doubt |
| The auditee promises to send evidence "after the audit" for a control central to the scope | Nothing verified within the audit | Nonconformity or unresolved issue; you cannot conclude on unseen evidence |
Two failure modes sit on either side of the principle. Over-applying it turns the audit into a courtesy visit: missing required records, an unjustified Statement of Applicability exclusion, or a control the auditee simply cannot demonstrate are nonconformities, and "we probably do it" is not doubt. Under-applying it produces findings the auditee will successfully appeal, damaging both the certification body's credibility and yours.
A usable test. Before writing a nonconformity, ask whether a reasonable, informed auditor reading your working papers would reach the same conclusion from the same evidence. If the honest answer is "they might not," either gather more evidence or give the benefit of the doubt and record an observation — then write down which you did and why. That note is what protects the finding in an appeal.
Exam trap: the benefit of the doubt concerns inconclusive evidence only. It is never justified by time pressure, a cooperative auditee, or a wish to keep the finding count low, and it never applies where the standard requires documented information that does not exist.
Working papers
Working papers must let another competent auditor reconstruct plan vs. performed work, populations/samples, interviews/observations, confidential evidence, finding derivation, and open items. Index to systems and criteria; protect model secrets, personal data, and proprietary prompts per opening rules.
Daily meetings and missing evidence
Team huddles align themes, avoid duplicate interviews, decide sample extensions, and pull specialist competence when needed. Auditee debriefs confirm facts and deadlines for missing files—they are not NC classification negotiations.
When evidence is missing (migrations, departed staff, vendor portals): request with deadline; offer alternatives (read-only session, secondary system); document limitations if still absent. Do not assume conformity from polish, and do not invent records from interviews when procedures or the standard require documented information. Persistent gaps for high-impact systems often support NCs on operational control, documented information, or related Annex A themes.
Exam scenarios: sampling bias traps
Lead Auditor exams love bias stories. Recognize these patterns:
| Bias | What happens | Corrective instinct |
|---|---|---|
| Convenience sampling | Only systems hosted in the HQ office / only English-speaking teams | Include remote, multi-language, or business-unit systems in scope |
| Success-only sampling | Auditee steers you to the award-winning model | Insist on inventory-based risk selection including weak docs |
| Guide-steered sampling | Escorts pre-select “ready” interviewees and tickets | Randomize within agreed populations; request full lists |
| Recency bias | Only last week’s perfect release after remediation theater | Sample across the full period, including before the cleanup |
| Survivorship bias | Only models still in production; ignore failed launches and rollbacks | Include retired/rolled-back systems for change and incident learning |
| Metric tunnel vision | Sample only accuracy dashboards, ignore impact and oversight | Balance lifecycle, data, impact, use, and third-party evidence |
If the auditee refuses access to a high-risk system “because it is sensitive,” that is not a free pass—it is a scope/access issue to escalate, potentially a limitation or NC depending on rules and reason.
From sample to conclusion
As fieldwork ends, reconcile plan vs. performed work, finalize findings with evidence references, separate clear NCs from items needing lead review, and prepare closing inputs (classified findings, conclusion direction, limitations). Strong sampling and disciplined finding statements are the heart of conducting the AIMS audit—and of certificate credibility.
When selecting AI systems for deep-dive sampling in an ISO/IEC 42001 Stage 2 audit, which approach best reflects risk-based practice?
Which situation is most suitable for statistical (or systematic attribute) sampling rather than pure system-level judgmental selection alone?
Which statement best describes a well-written nonconformity from AIMS fieldwork?
An auditee provides only last week’s release tickets after a special cleanup, while the audit period is the prior six months. What sampling bias risk is present, and what should the auditor do?
An auditee demonstrates monthly drift-monitoring reports for 11 of the past 12 months and cannot locate the twelfth during the visit. Separately, it has no AI system impact assessment at all for a high-impact production model. How should the audit team apply the benefit of the doubt?