7.2 Control Objective A.7: Data for AI Systems
Key Takeaways
A.7 contains five controls, A.7.2 through A.7.6.
A.7.2 governs data-management processes for development and enhancement; A.7.3 addresses acquisition and selection details.
A.7.4 sets data-quality requirements and ensures development and operational data meet them.
A.7.5 records provenance across data and AI system lifecycles.
A.7.6 defines criteria for selecting data preparations and preparation methods.
Annex A.7: data for AI systems
A.7 ensures the organization understands the role and impacts of data throughout development, provision, and use. It contains five controls. An inaccurate three-control summary that combines training partitions, provenance, and quality misses acquisition and preparation and uses the wrong numbers.
A.7.2 Data for development and enhancement
The organization defines, documents, and implements data-management processes related to development of AI systems.
A data-management process can cover ownership, collection, access, storage, labeling, partitioning, quality, change, protection, retention, disposal, and links to model versions. The scope depends on the system. A knowledge-based system may rely more on curated knowledge resources than on a large training corpus.
The control does not mandate a 70/15/15 split, inter-annotator kappa, or a specific platform. Those can be useful techniques when relevant.
A.7.3 Acquisition of data
The organization determines and documents details about acquisition and selection of data used in AI systems.
Relevant details can include source, collection method, selection rationale, time period, population, licensing or contractual constraints, consent or lawful basis where applicable, transformations, and limitations.
Public accessibility does not automatically establish permission for every use. However, ISO/IEC 42001 is not itself a copyright or privacy statute. The organization identifies applicable legal and contractual requirements through context and interested-party analysis and reflects them in acquisition controls.
A.7.4 Quality of data for AI systems
The organization defines and documents data-quality requirements and ensures data used to develop and operate the AI system meet those requirements.
Quality is fitness for purpose, not abstract perfection. Dimensions can include accuracy, completeness, consistency, timeliness, representativeness, relevance, uniqueness, validity, and label reliability. The appropriate dimensions and thresholds depend on intended use.
Operational data matters as well as development data. A system can be trained on adequate data and later receive inputs outside defined conditions. Quality controls can therefore include monitoring and response criteria.
Fairness analysis may depend on data-quality evidence, but A.7.4 does not prescribe one bias metric or remediation technique. Removing a sensitive feature can leave proxies; collecting the feature can create privacy constraints. The organization makes a justified, lawful decision.
A.7.5 Data provenance
The organization defines and documents a process for recording provenance of data used in AI systems across the lifecycles of the data and AI system.
Provenance supports traceability: where data came from, when and how it was obtained, what transformations or labels were applied, which versions were used, and where it flowed. It can also support incident investigation, reproducibility, rights management, and supplier oversight.
Cryptographic hashes, metadata catalogues, feature stores, and lineage graphs are possible implementations, not universal requirements.
A.7.6 Data preparation
The organization defines and documents criteria for selecting data preparations and methods to be used.
Preparation can include cleaning, normalization, imputation, deduplication, labeling, augmentation, encoding, balancing, tokenization, filtering, or partitioning. Criteria should follow requirements and avoid introducing hidden leakage or bias.
Data leakage example
If normalization statistics are calculated on a future test period before a forecasting model is trained, information leaks across the evaluation boundary and can inflate apparent performance. A suitable preparation process defines chronological splits and fits transformations only on permitted data.
That is a useful technical example, not the complete content of A.7.6. Data preparation applies more broadly than train/test splitting.
How the controls connect
| Control | Governing question |
|---|---|
| A.7.2 | What overall data-management processes govern development and enhancement? |
| A.7.3 | Where did selected data come from and why was it chosen? |
| A.7.4 | What quality does the use require, and is the data fit? |
| A.7.5 | Can the organization trace origin and transformation over time? |
| A.7.6 | How and why is data prepared for the system? |
A single record can support several controls, but their purposes remain distinct.
Third-party and synthetic data
Third-party data does not transfer all responsibility to a broker. The organization should obtain enough information to apply its acquisition, quality, provenance, and preparation processes. Contractual warranties can support evidence but may not be sufficient alone.
Synthetic data can reduce some constraints or expand coverage, but it can reproduce source bias, reduce diversity, or leak training examples. Relevant records can include generator, source data, method, parameters, validation, and limitations. These are context-dependent examples.
Data roles across an AI value chain
A producer may acquire and prepare training data. A provider may package a model with documentation. A user may supply operational inputs or fine-tuning examples. Responsibility allocation under A.10 should identify who supplies which data evidence. The user still evaluates whether operational data fits its context.
Evidence and proportionality
Evidence can include acquisition records, data specifications, quality reports, lineage metadata, preparation procedures, access controls, and change histories. Not every dataset requires the same detail. The degree should reflect risk, impact, applicable requirements, and feasibility while still meeting selected controls.
Tip
Remember the sequence: management process, acquisition, quality, provenance, preparation—A.7.2 through A.7.6.
Which control addresses acquisition and selection details for data?
A.7.1
A.7.3
A.7.5
A.7.6
What is the core of A.7.4?
Publishing all datasets
Mandating a 70/15/15 split
Defining data-quality requirements and ensuring development and operational data meet them
Using only synthetic data
Which statement best distinguishes provenance from preparation?
They are identical
Preparation concerns supplier contracts only
Provenance is A.6 and preparation is A.8
Provenance traces origin and transformations; preparation defines criteria and methods applied to data
Sections you finish are checked off in the contents.