6.1 Clause 9 Performance Evaluation
Key Takeaways
- Clause 9.1 requires determining what to monitor and measure for the AIMS and AI systems, methods, timing, and who analyzes results—including performance, fairness, drift, and incidents where relevant.
- Clause 9.2 internal audit is a first-party program that must be planned, objective, risk-based, and reported to management; it is not the same as third-party certification audit.
- Clause 9.3 management review needs AIMS-specific inputs (AI objectives, risk/impact results, external issues, audit and monitoring results) and decisions on improvement, resources, and changes.
- Lead auditors judge monitoring design adequacy—whether metrics can detect AI failure modes for scoped systems—not merely whether any dashboard exists.
- Shared ISMS evaluation forums support an integrated AIMS only when AI metrics, AI audit coverage, and AI management-review agenda items are actually evidenced.
6.1 Clause 9 Performance Evaluation
Auditor focus: Clause 9 asks whether the organization knows if its AIMS and AI systems are working. “We have Grafana” is not conformity. Test what is monitored, whether methods fit AI failure modes, who acts on the data, and whether internal audit and management review close the loop to leadership decisions.
Clause 9 Performance evaluation is the Check stage of the AIMS PDCA cycle. In ISO/IEC 42001:2023 it comprises monitoring, measurement, analysis and evaluation (9.1); internal audit (9.2); and management review (9.3). Together they answer: Is the AIMS effective? Are AI systems performing as intended and remaining trustworthy? Does top management see evidence and decide? For Lead Auditors, this is where polished policies often fail—undefined metrics, internal audit that skips high-risk models, or management review that never discusses AI objectives.
9.1 Monitoring, Measurement, Analysis and Evaluation
The organization shall determine what needs to be monitored and measured, the methods (to ensure valid results), when monitoring and measuring occur, when results are analyzed and evaluated, and who performs that work. It shall evaluate AIMS performance and effectiveness and retain documented information of the results.
What / how / when / who
| Dimension | Intent | Weak evidence | Strong evidence |
|---|---|---|---|
| What | AI and AIMS topics tied to intended outcomes | Only GPU/API uptime | Quality, fairness/slices, drift, overrides, incidents, objective progress, control effectiveness |
| Methods | Valid, consistent, reproducible | Ad-hoc notebook checks | Defined metrics, thresholds, data sources, sampling rules, versioned dashboards |
| When | Frequency and triggers fit risk/lifecycle | “When we have time” | Continuous in production; periodic offline; event-driven after retrain/incident |
| Who | Competent measurement and analysis owners | Unowned Slack channel | Named MLOps/product/risk owners; escalation to AIMS owner |
AI-relevant metric families
Generic IT service metrics are necessary but not sufficient:
| Family | Examples | Why auditors care |
|---|---|---|
| Performance / quality | Accuracy, precision/recall, error rate, inference SLOs | Silent degradation after data or model change |
| Fairness / bias | Group disparity, slice error rates | Equity treatments are monitored, not only promised |
| Drift / data quality | Feature drift, prediction drift, train–serve skew | Core production ML failure mode |
| Incidents & safety | AI incidents, overrides, rollbacks | Links Clause 8 operation to Clause 10 improvement |
| AIMS process metrics | % systems with current impact assessment; finding closure time | Measures the management system, not only the model |
| Objectives | Status vs Clause 6.2 AI objectives | Feeds 9.3 management review |
Design adequacy: Judge whether monitoring can detect foreseeable failure modes for scoped systems. A high-risk credit model with no fairness or drift monitoring is a 9.1 design gap even at 99.9% uptime. Over-instrumenting a low-risk internal tool while ignoring a customer-facing high-impact system also signals weak risk-based design.
Sampling probes: Pick 2–3 systems (prefer high risk). Request metrics, owners, thresholds, last analysis. Trace one breach: notification, decision, retained records. Map residual risks from risk/impact assessments to metrics. Check that retrain/redeploy triggers re-evaluation.
Trap: Offline development validation is not the same as ongoing operational evaluation. Clause 9.1 expects continuing evaluation over time, not a single go-live accuracy report.
9.2 Internal Audit
The organization shall conduct internal audits at planned intervals to determine whether the AIMS conforms to its own requirements and ISO/IEC 42001 and is effectively implemented and maintained. It shall maintain an audit program covering frequency, methods, responsibilities, planning, and reporting—considering process importance, change, and prior results. Auditors shall be selected to ensure objectivity and impartiality. Results go to relevant management; documented information evidences the program and results.
Internal audit vs certification audit
| Aspect | Internal audit (9.2) | Certification (third-party) |
|---|---|---|
| Party | First-party program (may use contracted auditors under client control) | Independent conformity assessment body |
| Purpose | Self-assurance; feed review and improvement | Independent certification decision |
| Impartiality | Do not audit own work; manage conflicts | CAB independence (e.g., ISO/IEC 17021-1 context) |
| Frequency | Risk-based planned intervals | Certification cycle (initial, surveillance, recert) |
| Substitution? | No—certification does not cancel 9.2 | No—internal audit does not replace certification |
Program expectations: rolling coverage of AIMS scope and high-risk AI systems; AI-competent auditors; objectivity safeguards; findings reported to people who can resource action; open NCs linked to Clause 10.
Scenario: Annual “IT security” internal audit is claimed as AIMS coverage. Reports never sample impact assessments, model monitoring, or AI policy conformity. Treat ISMS internal audit as partial evidence of an integrated program only if AIMS criteria appear in the program and working papers. Otherwise raise 9.2: AIMS internal audit missing or ineffective.
9.3 Management Review
Top management shall review the AIMS at planned intervals for continuing suitability, adequacy and effectiveness. Reviews consider prior actions; external/internal issues; performance information (nonconformities and corrective actions, monitoring results, audit results, improvement opportunities); and resource adequacy. Outputs include decisions on continual improvement and any need for AIMS changes.
AIMS-specific inputs and outputs
| Inputs (expect AI flavor) | Outputs |
|---|---|
| Progress on AI objectives (6.2) | Revise objectives, KPIs, resourcing |
| AI risk and impact assessment results | New treatments, controls, scope changes |
| External issues (regulation, model markets, customer AI requirements) | Policy/process updates; interested-party communication |
| 9.1 monitoring (drift, fairness, incidents) | Operational changes, retrain priorities, sunset decisions |
| Audit results (internal and external) | CA prioritization; program adjustments |
| Resource/competence gaps for AI roles | Hiring, training, tooling, vendor decisions |
Evidence: agenda, minutes/records, top management participation (or effective feedback loop), actions with owners and dates. Combined QMS/ISMS/AIMS reviews work only when AI items and decisions are explicit.
Scenario: Integrated management review covers security incidents and quality CAPA but never AI objectives, model performance, impact assessment status, or AI NCs—despite production clinical models in scope. Weekly data-science standups do not replace Clause 9.3. Raise 9.3 (and possibly 5.1) for failure to review AIMS with required information.
Evaluating Monitoring Design Adequacy
- Scope & inventory — Any high-impact system unmonitored?
- Risk & impact linkage — Residual risks mapped to metrics or justified alternatives?
- Method validity — Reproducible methods; controlled evaluation data and fairness definitions?
- Timeliness — Frequency vs change rate; post-incident and post-change evaluation?
- Analysis & escalation — Owners, thresholds, retained results?
- 9.2 / 9.3 integration — Does internal audit test monitoring? Does management review see trends?
Common traps: treating certification as internal audit; security-only KPIs for AIMS; model cards without production monitoring; “AI strategy” reviews without performance data; vendor metrics with no client-side AIMS evaluation.
Clause 9 closes when measurement informs decisions. Clause 10 (next) covers nonconformity and improvement when evaluation shows gaps.
Under ISO/IEC 42001 Clause 9.1, which monitoring portfolio best demonstrates AI-relevant performance evaluation for a high-risk production decisioning model?
How should a lead auditor treat a claim that the annual ISO/IEC 27001 internal audit fully satisfies ISO/IEC 42001 Clause 9.2?
Which set best represents AIMS-specific inputs expected in a Clause 9.3 management review?
What is the primary difference between Clause 9.2 internal audit and a certification audit against ISO/IEC 42001?