9.2 A.9 Use of AI Systems
Key Takeaways
- Annex A.9 (three controls, typically A.9.2–A.9.4) governs responsible use: processes and objectives for use, intended use versus foreseeable misuse, user guidance, and human oversight during operation.
- Intended use must be explicit; foreseeable misuse (prompt leakage of secrets, over-reliance on generative output, using decision support as fully automated decisions) needs controls, not hope.
- Human oversight is ineffective without competence, time, authority to override, and monitoring of override/acceptance rates—rubber-stamp review fails auditor tests.
- Employee generative-AI tools and decision-support systems are high-yield audit samples: acceptable use, data handling, verification rules, and escalation paths must be evidenced.
- A.9 connects to Clause 8 operation, A.6 lifecycle intended use, A.8 user information, and Clause 7 competence—use controls are not “IT acceptable use only.”
9.2 A.9 Use of AI Systems
Auditor focus: A.9 asks whether the organization uses AI systems as intended—and prevents or detects misuse—after design and deployment. Pre-production validation does not prove responsible use. Sample how people actually work with the system: prompts, overrides, confidential paste, silent acceptance of bad recommendations, and shadow tools outside the inventory.
Annex A.9 Use of AI systems comprises three controls (commonly A.9.2–A.9.4). The objective addresses processes for responsible use, objectives related to use, intended purpose boundaries, user guidance, and human oversight during operation. For Lead Auditors, A.9 is the runtime "people + system" layer—distinct from A.6 lifecycle gates and pure infrastructure monitoring.
Control themes under A.9
| Theme | Auditor question | Strong evidence | Weak evidence |
|---|---|---|---|
| Responsible-use process | Are people authorized and expected to use AI in a managed way? | AUP, role-based access, exceptions process | "Common sense" only |
| Objectives for use | Are use outcomes defined and checked? | Output quality samples, leakage metrics, compliance checks | Only chatbot "adoption rate" |
| Intended-use boundaries | Is intended use documented and enforced? | Intended-use statements, blocked workflows | "Research only" card used in production credit |
| User guidance | Do users know how to verify and escalate? | In-product tips, SOPs, job aids | 40-page PDF never opened |
| Human oversight in use | Can humans intervene meaningfully—and do they? | HITL design, authority, override telemetry | Checkbox "reviewed" in 3 seconds |
SoA: Include A.9 whenever humans or workflows operate AI in scope—including purchased tools. Excluding A.9 because "we only use vendor SaaS" is usually weak: use is where deployer responsibility concentrates.
Intended use vs foreseeable misuse
Intended use covers purpose, context, users, data types, and decision stakes. Foreseeable misuse includes reasonably predictable stretch or abuse even without malice.
| Intended use | Foreseeable misuse | Control response |
|---|---|---|
| Draft internal meeting summaries | Paste regulated customer data into public LLM | Enterprise tenant, DLP, banned-data list |
| Decision support for underwriters | Treat score as automatic approval | Policy + forced reason codes; QA of overrides |
| Coding assistant (non-prod) | Ship unverified AI code to production | PR checks, human review gate |
| Content moderation assist | Auto-publish all "low risk" labels | Thresholds, spot checks, escalation SLAs |
| HR screening assist | Proxy features encoding prohibited discrimination | Feature policy, adverse-impact monitoring |
Method: From the intended-use statement, brainstorm misuse, then ask what control exists. "Users are professionals" is inadequate for high residual risk.
Human oversight in use
| Pattern | Human role | Failure mode to sample |
|---|---|---|
| Human-in-the-loop | Approval before consequence | Queue pressure → auto-accept culture |
| Human-on-the-loop | Monitor and intervene | Alerts ignored; no on-call |
| Human-in-command | Sets constraints; can stop system | Kill-switch unknown to operators |
Effectiveness tests: (1) Authority to refuse without penalty; (2) Competence on limitations and automation bias; (3) Time and information to decide; (4) Telemetry—override rates, agreement rates, time-to-decision; (5) Consequences when humans never challenge the model.
Scenario: Policy requires underwriter review of AI scores. Logs show median review of two seconds and 97% acceptance; KPIs reward only cases per hour. Finding: oversight is nominal—A.9 (and often 9.1). Fix process and metrics, not a reminder email.
User guidance and objectives for use
Guidance should be task-complete: allowed inputs, how to interpret outputs, when to escalate, how to record rationale, and prohibitions. Feature-only help ("click Generate") fails A.9 intent.
Objectives for use make responsibility measurable—e.g., max share of decisions without documented rationale; max confidential leakage events from AI tools; minimum quality sample pass rate for external generative outputs.
| Guidance element | Why it matters |
|---|---|
| Allowed / prohibited prompt data | Confidentiality and privacy |
| Verification rules for generative output | Stops hallucinations in customer/legal channels |
| Escalation when AI is uncertain | Makes oversight real |
| Tool allow-list | Controls shadow AI |
| Misuse examples | Improves awareness |
High-yield auditor scenarios
Employees using generative AI
Sample approved tools vs discovery of shadow SaaS; AUP on PII/source code/client secrets; technical enforcement (SSO enterprise instance, DLP, logging)—not PDF-only; training and interviews ("What would you paste?"); leakage incidents.
Common NC: Policy bans public LLMs; employees use personal accounts; AIMS owner assumes IT blocks them—no evidence.
Decision-support and automation bias
Sample RACI for final decision; UI forcing disagreement reasons; QA of accepted vs overridden cases; whether metrics reward speed over correct challenge. Common NC: System labeled "advisory only" but operationally drives automatic outcomes.
Scope creep after deployment
A model approved for fraud triage later freezes accounts automatically. Intended use and impact assessment not updated; no new guidance. Findings span A.9, lifecycle/change, and impact assessment.
Multi-clause linkage and evidence list
| Observation | Primary A.9 angle | Often also |
|---|---|---|
| No AUP for GenAI | Responsible use process | 7.3, A.2 |
| Users lack limitation knowledge | User guidance | A.8 |
| Rubber-stamp HITL | Oversight in use | 9.1, 7.2 |
| Use outside intended purpose | Intended use | 8.1, 6.1 |
| Unvetted personal AI tools | Unauthorized use | Inventory; A.10 if SaaS |
Request: responsible-use procedure; intended-use statements; point-of-use guidance; training records; allow-lists; override logs and quality samples; misuse/exception records; use-related management metrics.
Trap: Accepting ISO 27001 acceptable use as full A.9. Classic AUP covers assets and malware; A.9 needs AI-specific intended use, output verification, oversight, and misuse modes (hallucination, prompt injection via pasted content, discriminatory reliance).
A.9 conformity means use is governed: boundaries are known and enforced proportionally, oversight works in practice, and the organization checks that use remains responsible as tools evolve.
Which situation best illustrates a failure of Annex A.9 human oversight in use, even if a “human review” step exists on paper?
Employees paste customer PII into a consumer generative-AI website despite an AI policy forbidding it, and there is no enterprise tool or technical control. What is the strongest A.9-oriented audit issue?
A credit model’s documented intended use is “decision support for underwriters,” but the workflow auto-issues low-score declines without human action. How should the lead auditor treat this?