12.1 Criteria for Evidence Quality

Key Takeaways

  • GIAS Standard 14.1 requires information that is relevant (bears on the objective and evaluation criteria, inside scope), sufficient (persuasive quantity so a prudent, informed, competent person could reach the same conclusion), and reliable (factual, current, trustworthy source)
  • CIA Part 2 B2a is an AND test: failing any one attribute means you do not conclude on that packet — volume does not repair irrelevance, and original documents do not repair a wrong question
  • Relevance fails when the evidence answers a nearby topic (matched invoices filed for a vendor-master dual-authorization objective) or uses the wrong period, population, or criterion
  • Sufficiency fails when nature, extent, or timing cannot support the intended conclusion — one observed Friday review does not prove a year of operating effectiveness
  • If relevant evidence cannot be obtained, identify a finding or scope limitation; do not launder the gap by calling leftover documents close enough
Last updated: August 2026

12.1 Criteria for Evidence Quality

Quick Answer: CIA Part 2 B2a tests whether you can apply suitable criteria to evidence: it must be relevant (it bears on the engagement objective and evaluation criteria), sufficient (enough persuasive quantity to support the conclusion), and reliable (quality you can trust given the source and how you obtained it). Global Internal Audit Standards (GIAS) Standard 14.1 requires all three before you analyze and evaluate. Failing any one attribute means you gather more, change procedures, or — if relevant evidence cannot be obtained — consider a finding. Volume does not repair irrelevance.

Chapter 11 identified sources of information: interviews, observations, walk-throughs, data analysis, policies, checklists, questionnaires, and self-assessments. This chapter is the quality filter you apply while you gather. Section B2 of the 2025 syllabus is proficient-level: evaluate relevance, sufficiency, and reliability of evidence gathered to support engagement objectives. You are in GIAS Principle 14 — Conduct Engagement Work — not still choosing design procedures (Chapter 9) and not yet aggregating findings into an engagement conclusion (Chapter 16).

The 2019 IPPF Performance Standard 2310 listed four attributes: sufficient, reliable, relevant, and useful. GIAS 2024 and the 2025 CIA Part 2 B2 bullets use three. If an item offers “useful to the organization” as a required fourth quality test for B2, treat it as a leftover, not the current exam criterion. Usefulness still matters as a professional outcome; it is not the B2a scoring key.

Why three attributes, not one score

Evidence quality is not a blended gut feel that the file looks thick. Each attribute answers a different question. The exam will give you a packet that is excellent on two attributes and broken on the third, then ask whether you may conclude.

AttributeQuestion it answersFailure modeTypical repair
RelevanceDoes this information bear on the objective and the evaluation criteria?Right documents, wrong questionChange what you collect; do not enlarge a wrong population
SufficiencyIs there enough persuasive quantity that a prudent, informed, competent person could reach the same conclusion?Thin sample, one snapshot, inquiry onlyExpand extent, add procedures, cover the period
ReliabilityIs the information factual and current, from a source and method you can trust?Biased source, untested IPE, oral claim, weak systemObtain it yourself, corroborate, test IPE, go to an independent source

GIAS Standard 14.1 states the same three ideas in standard language:

  • Relevant — consistent with engagement objectives, within the scope of the engagement, and contributes to the development of engagement results.
  • Reliable — factual and current. Internal auditors use professional skepticism (GIAS Standard 4.3) to evaluate whether information is reliable. Section 12.2 unpacks the three official strengtheners: obtained directly or from an independent source, corroborated, and gathered from a system with effective governance, risk management, and control processes.
  • Sufficient — enables internal auditors to perform analyses and complete evaluations, and can enable a prudent, informed, and competent person to repeat the engagement work program and reach the same conclusions as the internal auditor.

Sufficiency therefore has two faces: enough for you to analyze, and enough for an outsider to reperform. Section 12.3 develops the outsider test as a documentation standard (GIAS 14.6). Here, treat it as the quantity and persuasiveness test.

If evidence is not relevant, reliable, or sufficient, Standard 14.1 requires you to determine whether to gather additional information. If relevant evidence cannot be obtained — records destroyed, a system with no audit trail, a scope block you cannot lift — identify that as a finding or a scope limitation. Scope limitations were Chapter 2. Here the skill is refusing to launder a gap by calling leftover documents close enough.

Relevance: it must bear on the objective and the criterion

Relevance is a relationship test, not a popularity test. Information is relevant when it helps you compare condition to evaluation criteria for the stated objective (GIAS 13.4 and 14.2). A perfect bank confirmation is irrelevant to an objective about physical inventory existence at a warehouse you never counted. A complete, well-controlled HR timesheet extract is irrelevant to whether three-way match blocked unmatched invoices before payment.

Ask three questions of every item you file:

  1. Objective — Does this information answer the engagement objective, or a nearby interesting topic?
  2. Criterion — Does it speak to the standard you selected (policy, law, configuration, contract, COSO principle), or to a different standard?
  3. Scope — Is it inside the activities, locations, systems, and period named in the scope?

Worked example — relevance fails. Objective: determine whether vendor bank-account changes in the ERP required dual authorization during the fiscal year. Criterion: the documented dual-control workflow. The auditor obtains 40 original, matched vendor invoices from a system with strong IT general controls. The invoices are reliable and plentiful. They do not bear on bank-account master-data changes. Filing them as the primary support for the objective is a relevance failure. Repair: obtain the vendor-master change log, workflow history, and who-approved extracts for the period — then apply sufficiency and reliability tests to those items.

Relevance also fails when the period is wrong (last year’s SOC report used as the sole support for this year’s access recertification), when the population is wrong (testing only domestic vendors when the objective includes the shared-service center’s offshore payees), or when the auditor tests a compensating control and then concludes on the original control without saying so.

Sufficiency: quantity and persuasiveness, not page count

Sufficiency is not “we printed a lot.” It is whether the persuasive weight of relevant, reliable evidence can support the conclusion you intend to draw. A high-risk, high-frequency manual control supported by two invoices from the last week of fieldwork is insufficient even if both invoices are authentic originals. A low-risk annual control may be sufficient with a deep test of that one occurrence.

Persuasiveness comes from nature × extent × timing, the same idea you used when planning operating-effectiveness procedures in Section 9.2:

  • Nature — inquiry is the least persuasive; observation is a snapshot; inspection and reperformance carry more weight. Inquiry alone is almost never sufficient to support an assurance conclusion.
  • Extent — how much of the population, how many procedures, how many independent sources.
  • Timing — evidence from throughout the period, including peak and cutoff, not a single comfortable month.

Sufficiency is conclusion-dependent. The evidence needed to say “we found no exceptions in a walk-through of design” is smaller than the evidence needed to say “the control operated effectively all year.” The evidence needed for a potential finding you will rate significant is larger than the evidence needed to drop a trivial observation. Do not over-file low-risk noise and under-file the issue that will drive the engagement conclusion.

Worked example — sufficiency fails. Objective: operating effectiveness of the Friday unmatched-invoice review. The auditor watches the supervisor perform the review once, on the Tuesday internal audit is in the building. The observation is relevant and, as far as it goes, reliable. It is not sufficient to conclude the control operated throughout the year. Repair: inspect signed reports across the period, test IPE completeness of the report, and reperform a sample of investigations. One observed performance is a snapshot, not a year.

Reliability: quality and source (preview of 12.2)

Reliability asks whether the information is factual and current and whether you obtained it in a way that professional skepticism would accept. A document can be relevant and numerous and still be unreliable: a photocopy the process owner printed the night before, an unsigned spreadsheet with no tie-out to the subledger, oral assurance from the person whose work you are evaluating, or an exception report nobody tested for completeness (IPE — information produced by the entity).

Worked example — reliability fails. Objective: existence of inventory at a third-party warehouse. The warehouse manager emails a listing the manager prepared in Excel, with no receiving records, no independent count, and no auditor observation. The listing is relevant (it names the SKUs in scope) and voluminous (12,000 lines). It is not reliable as the sole support for existence. Repair: observe or independently count, obtain receiving and shipping records from a system you test, or confirm with an independent party — then corroborate.

All three, or you do not conclude

Treat the attributes as AND gates. Relevant plus sufficient but unreliable is advocacy, not assurance. Relevant plus reliable but insufficient is a hunch with nice exhibits. Sufficient plus reliable but irrelevant is a well-supported answer to the wrong question.

PacketRelevant?Sufficient?Reliable?May you conclude on the objective?
40 original matched invoices, strong ITGCs, objective is vendor bank-account dual authorizationNon/aYesNo — wrong evidence
One observed Friday review, objective is year-long operating effectivenessYesNoYes, for that dayNo — expand
Process-owner spreadsheet of “all exceptions closed,” weak controls, no corroborationYesMaybe by volumeNoNo — corroborate or replace
Auditor-obtained change log plus workflow IDs plus independent HR termination list, period coverage, IPE testedYesYesYesYes — proceed to analysis (14.2)

Exam traps

  • Equating volume with sufficiency, or original documents with relevance.
  • Concluding on inquiry alone.
  • Using last period’s evidence as if it were current (reliability and relevance both suffer).
  • Treating a 2019-style “useful” fourth attribute as the B2a key.
  • Filing evidence that supports a different criterion than the one you selected in planning (Chapter 3).
  • Skipping the Standard 14.1 duty to gather more — or to report that relevant evidence could not be obtained.

B2a is complete when you can look at a packet and name which attribute fails, why, and what you do next — not when you can recite three definitions.

Loading diagram...
GIAS 14.1 three-attribute gate for engagement evidence
Illustrative persuasiveness by nature (1–4; still apply relevance and reliability — not IIA-published scores)
Test Your Knowledge

Which statement best applies the GIAS 14.1 / CIA Part 2 B2a criteria for evidence quality?

A
B
C
D
Test Your Knowledge

An engagement objective is to determine whether vendor bank-account changes required dual authorization during the year. The auditor files 40 original three-way-matched invoices from an ERP with effective IT general controls. What is the strongest evaluation of that packet?

A
B
C
D
Test Your Knowledge

The auditor observes the Friday unmatched-invoice review once during fieldwork and concludes the control operated effectively all year. Which quality attribute most clearly fails?

A
B
C
D