10.3 Evidence-Based & Risk-Based Audit Approaches
Key Takeaways
- Audit evidence must be sufficient (enough quantity) and appropriate (relevant and reliable quality) to support findings and conclusions
- Common AIMS evidence methods include document review, interviews, observation, and technical artifact examination (model cards, lineage, logs, impact assessments)
- Sampling is normal: auditors design risk-based samples that concentrate on high-risk AI systems, weak controls, prior findings, and areas of significant interested-party impact rather than attempting 100% verification of every AI system
- Professional judgment is required because evidence has limitations — an audit delivers reasonable assurance on a defined scope and sample, never absolute assurance or a guarantee that no nonconformity exists, and materiality decides which deficiencies are significant enough to affect the conclusion
- Inherent risk and control risk belong to the auditee and are assessed, not changed; detection risk — the chance the audit misses a real nonconformity — is the only one the auditor controls, through sample size, selection basis, and method.
10.3 Evidence-Based & Risk-Based Audit Approaches
Two ISO 19011 principles shape day-to-day AIMS audit technique more than any others: the evidence-based approach and the risk-based approach. Together they answer: What will we look at, how much is enough, and how do we know our conclusions are trustworthy?
Audit Evidence: Sufficient and Appropriate
Audit evidence is records, statements of fact, or other information which are relevant to the audit criteria and verifiable. Evidence supports findings; findings support conclusions.
Two quality dimensions matter:
| Dimension | Meaning | Weak AIMS example | Stronger AIMS example |
|---|---|---|---|
| Sufficient | Enough quantity to support the finding | One interview that “impact assessments exist” | Multiple completed assessments + process records + interview triangulation |
| Appropriate | Relevant and reliable quality | Marketing slide claiming “ethical AI” | Versioned model card, evaluation report, change ticket, monitoring alert history |
Evidence should be verifiable. “Trust me” statements, undated screenshots without context, and draft policies never approved are weak. Corroboration across methods increases reliability.
Hierarchy of confidence (practical, not absolute):
- Direct observation of process in operation + system-generated records
- Independent documented records created as part of normal operations
- Interviews corroborated by documents or observation
- Uncorroborated interviews or self-prepared summaries
Methods of Collecting Evidence
ISO 19011 describes methods that can be used alone or in combination. In AIMS audits, expect a blended toolkit:
1. Review of documented information
Policies, AI system inventories, risk and impact assessments, SoA, procedures, training records, supplier contracts, data governance records, model documentation.
AIMS tip: A beautiful AI policy without operational records for high-risk systems is design evidence only—not effectiveness evidence.
2. Interviews
Process owners, data scientists, MLOps engineers, risk/compliance, product managers, third-party managers, and (where relevant) human-oversight operators.
AIMS tip: Ask “show me the last time this control ran,” not only “do you have a process?”
3. Observation
Watch change-approval workflows, human review of model outputs, data labeling quality checks, or incident response tabletop evidence where live observation is feasible.
4. Examination of technical and operational artifacts
Model cards / system cards, dataset datasheets, lineage and feature-store records, training/evaluation metrics, bias test results, monitoring dashboards, drift alerts, access logs, deployment tickets, rollback history, red-team or evaluation reports.
EVIDENCE TRIANGULATION FOR AN AIMS CONTROL
(e.g., AI impact assessment)
Documents Interviews Artifacts
───────── ────────── ─────────
Procedure + Process owner + Completed AIA
templates model owner records linked
to system IDs
↘ ↓ ↙
Corroborated finding
Sampling in AIMS Audits
Audits almost always use sampling because it is impractical to examine every AI system, every dataset, every commit, and every decision log. Sampling may be judgmental (professional judgment guided by risk) and/or more structured depending on the audit type and available data.
Risk-based sampling drivers for AIMS:
- AI systems with high impact (rights, safety, financial, legal)
- Systems processing sensitive personal data
- New or significantly changed models since last audit
- Prior nonconformities or incidents
- Heavy third-party model or data dependence
- Weak or missing monitoring signals
- Boundary processes between development and production
Example sample plan excerpt:
- 2 high-risk credit models (full lifecycle and impact assessment trail)
- 1 hiring-screening system (fairness evaluation + human oversight)
- 1 low-risk internal productivity LLM (lighter sample for contrast)
- 1 critical third-party foundation-model supplier control set (A.10)
- Cross-cutting samples of competence records, internal audit, management review
Exam trap: Sampling is not a defect. Failure to design sampling with risk and to disclose limitations can be a defect. Claiming 100% verification when only samples were tested violates fair presentation.
Risk-Based Planning of Audit Activities
Risk-based auditing means audit planning and execution consider risks and opportunities so that resources focus on significant matters. In AIMS practice:
At program level
More frequent or deeper coverage for business units running high-risk AI; coordination with related ISMS/QMS audits; attention after major model launches or regulatory changes.
At individual audit level
- Use Stage 1 results to reshape Stage 2 depth
- Allocate senior AI-competent auditors to complex technical areas
- Sequence activities so critical high-risk systems are covered early (time buffer for follow-up)
- Adjust the plan when new information appears (undocumented production model, major incident, missing logs)
Scenario: Stage 1 shows strong policy documentation but incomplete impact assessments for customer-facing generative AI and no inventory of shadow AI tools. A risk-based Stage 2 plan spends less time re-reading the polished AI policy and more time on system inventory completeness, impact assessment effectiveness, data for AI (A.7), information for interested parties (A.8), and third-party tools (A.10).
Risk-based does not mean ignoring low-risk areas entirely. Some low-risk samples still test whether the management system works consistently. It means proportion, not blind spots by design without justification.
Inherent, Control & Detection Risk
The risk-based approach becomes usable only when you split risk into the three components PECB names. Together they decide how much testing an area needs.
| Risk | Question it answers | AIMS example | Can the auditor change it? |
|---|---|---|---|
| Inherent risk | How exposed is this area before any control? | A generative chatbot giving unscripted advice to consumers is inherently riskier than an internal document classifier | No — it is a property of the activity. You assess it; you do not reduce it |
| Control risk | How likely is it that the organization's own controls fail to prevent or detect the problem? | Bias testing exists, but it is run manually by the model's own developer with no independent review | No, but you assess control design and operation; weak controls raise the testing you must do |
| Detection risk | How likely is it that your audit misses a nonconformity that is really there? | Sampling only the three systems the auditee volunteered, all of them low-impact | Yes — this is the one you control |
The working relationship: high inherent risk plus high control risk means detection risk must be driven down — larger samples, a different selection basis, a second evidence method, or a technical expert on the team. Where inherent and control risk are both low, lighter testing is a defensible professional judgment, and you record why.
Applied to an AIMS: an automated credit-decisioning model is high inherent risk. If the organization's only control is a quarterly self-assessment signed by the data-science lead, control risk is high too. Interviewing that same lead and reading their own summary leaves detection risk high — three highs stacked, and a clean conclusion would be indefensible. Reduce detection risk by pulling the production model ID, re-performing a slice-performance check against the retained evaluation set, and tracing two live declined applications to their override records.
Escalate rather than paper over. If you cannot get detection risk to an acceptable level in the audit time available — for example the AI inventory is demonstrably incomplete, so you cannot define a population to sample — that is a feasibility and planning problem to raise with the audit team leader and the client, not a gap to absorb with a favourable conclusion.
Materiality & Reasonable Assurance
Materiality is the threshold at which a deficiency matters enough to affect the audit conclusion. In an AIMS audit it is rarely monetary. The practical drivers are the number and vulnerability of people affected, the reversibility of the harm, whether the AI decision is automated or advisory, regulatory exposure, and whether the failure is systemic or isolated. One missing signature on one low-risk model card is immaterial; the same gap across every high-impact system reveals a process failure and is material. Materiality is therefore what stops an audit report from being a list of trivia, and what stops a serious pattern from being graded as an administrative slip.
Reasonable assurance is the level of confidence an audit can honestly deliver. An audit samples; it does not examine every record, so it cannot promise the AIMS is free of every nonconformity. What the auditor concludes is that, on sufficient and appropriate evidence, the AIMS conforms to the criteria and is effective — stated against a defined scope, period, and sampling basis. Absolute assurance would require testing the whole population, which is neither affordable nor required by ISO 19011 or ISO/IEC 17021-1.
| Concept | What it governs | Wrong answer to avoid |
|---|---|---|
| Materiality | Whether a deficiency is significant enough to report and to affect the conclusion | "Every deviation is a nonconformity of equal weight" |
| Reasonable assurance | How much confidence the conclusion may claim | "The audit guarantees there are no nonconformities" |
| Sufficient evidence | Quantity — enough to support the conclusion | "One interview is enough because the interviewee was senior" |
| Appropriate evidence | Quality — relevant and reliable for the requirement tested | "A screenshot of a dashboard proves the control operated all year" |
Exam trap: options claiming an audit "guarantees", "certifies that no nonconformity exists", or "provides absolute assurance" are always wrong. So are options that use reasonable assurance as cover for thin sampling — reasonable assurance is a conclusion standard, and reaching it is precisely what forces sample sizes up when inherent and control risk are high.
Limitations of Evidence
Every AIMS audit faces evidence limitations. Common sources:
| Limitation | Why it happens in AI environments | Auditor response |
|---|---|---|
| Incomplete historical logs | Short vendor retention; tool changes | Report limitation; seek alternatives; raise control weakness if retention required |
| Black-box third-party models | IP restrictions; limited explainability | Assess organizational controls, contractual rights, evaluation evidence, residual risk handling |
| Ephemeral cloud environments | Dynamic infrastructure | Rely on IaC records, CI/CD logs, access governance |
| Sampling uncertainty | Finite sample | Avoid over-generalization; state confidence carefully |
| Interview bias | Social desirability, fear | Corroborate; interview multiple levels |
| Timing | Snapshot vs continuous operation | Consider period covered; surveillance continuity |
Limitations do not automatically invalidate an audit, but material limitations must influence conclusions and be communicated (fair presentation). If limitations are so severe that objectives cannot be met, the audit may need scope change, extension, or—in extreme cases—termination/replan.
Professional Judgment
Professional judgment is the application of relevant training, knowledge, experience, and ethical standards to decisions about planning, evidence, findings, and conclusions. It is not personal preference or “auditor gut feel” disconnected from criteria.
Judgment is required to:
- Decide what constitutes sufficient evidence for a high-risk AI control vs a low-risk control
- Interpret whether an exclusion in the SoA is justified
- Weigh conflicting evidence (engineers say monitoring works; dashboards show weeks of missing alerts)
- Classify findings (major/minor nonconformity vs observation) based on criteria and risk impact
- Determine when team competence is inadequate and technical experts must be added
AIMS judgment example: An organization produces fairness metrics quarterly for a hiring model, but metrics cover only one demographic axis and ignore a known high-risk subgroup identified in the impact assessment. Documents “exist,” yet judgment may support a nonconformity against effectiveness of impact-assessment outcomes and related controls—not a clean pass.
Judgment must still be evidence-based and criteria-referenced. “I don’t like this model” is not a finding. “Criteria require X; evidence shows X is missing/ineffective; impact is Y” is a finding.
Linking Evidence-Based and Risk-Based Approaches
Think of them as orthogonal axes:
RISK-BASED (where to focus)
▲
│ High-risk AI systems,
│ weak controls, prior NCs
│
─────────────────────┼─────────────────────▶ EVIDENCE-BASED
│ (how to conclude)
│ Sufficient &
│ appropriate evidence
│
- Risk-based without evidence-based → opinionated tours of scary systems without proof
- Evidence-based without risk-based → meticulous proof on trivial low-risk areas while high-impact AI remains unexamined
Lead Auditors must do both: focus where it matters and prove what they claim.
Worked Mini-Case: Monitoring Control for a High-Risk Model
Criteria: ISO/IEC 42001 operational control / Annex A monitoring-related expectations as applicable via SoA; organization’s monitoring procedure.
Risk-based choice: Select the production credit model with automated adverse decisions.
Evidence methods:
- Document: monitoring procedure, thresholds, escalation matrix
- Artifact: 90 days of drift/performance dashboards, alert tickets
- Interview: MLOps lead and credit risk owner
- Observation: walkthrough of alert triage board
Finding logic: Procedure requires alerts within 24 hours for performance drops below threshold. Tickets show three breaches with 5–11 day delays and no root-cause records. Interviews confirm “we were busy with a launch.” Evidence is sufficient and appropriate to support a nonconformity on implementation/effectiveness—not merely an observation that “documentation could be clearer.”
Exam Focus
Be ready to define sufficient vs appropriate evidence; list methods with AI artifacts; explain why sampling is expected; show how risk directs sample selection toward high-risk AI; and describe how limitations and professional judgment interact without abandoning criteria or verifiable evidence.
Which pair best describes the two quality attributes of audit evidence emphasized in management-system auditing?
An auditor has only two days on site and must cover an AIMS with 40 AI systems. Which sampling approach best reflects a risk-based audit approach?
A foundation-model vendor refuses to provide internal weight files but supplies evaluation summaries and contractual audit rights short of full model IP disclosure. What is the most professional auditor response?
Why is professional judgment still required even when the auditor follows an evidence-based approach?
An AIMS audit covers an automated loan-decisioning model (high inherent risk) whose only control is a quarterly self-assessment signed by the model's own developer (high control risk). Which action correctly addresses the risk position?
A closing meeting attendee asks the audit team leader to confirm in the report that the organization's AI management system contains no nonconformities. What is the correct response?