9.2 Procedures to Test Operating Effectiveness

Key Takeaways

  • Operating effectiveness asks whether a soundly designed control operated as designed throughout the period, by people with the authority and competence to perform it, with evidence that it happened
  • Nature of tests, from least to most persuasive: inquiry, observation, inspection of evidence, reperformance — inquiry alone is never enough to conclude a control operated
  • Extent is conceptual: frequency, manual versus automated, risk, and period coverage drive how much you test — this chapter is not a statistics textbook; sampling as a B1 information method comes later
  • Automated application controls are tested mainly through configuration plus IT general controls and limited instances; manual controls need items spread across the period, including busy and cutoff windows
  • Information produced by the entity (IPE) used in a control or as evidence must be tested for completeness and accuracy; reviewing an incomplete exception report is not an effective review
Last updated: August 2026

9.2 Procedures to Test Operating Effectiveness

Quick Answer: Operating effectiveness asks whether a control that is adequately designed actually operated as designed throughout the period, performed by people with the authority and competence to do it, with evidence that it happened. CIA Part 2 A6b tests whether you can identify the right mix of inquiry, observation, inspection, and reperformance, scale extent to manual versus automated controls, and test information produced by the entity (IPE) before you rely on it. Inquiry alone does not get you there.

Section 9.1 decided whether the control could work. This section assumes you are past that gate — or that you are testing a compensating control whose design you have already evaluated. If design is broken, do not dress up a large operating sample as professionalism. If design is sound, A6b is the planning skill of choosing procedures that will later let you conclude the control functioned (COSO) during the engagement period.

You are still determining procedures for the work program (GIAS 13.6), not writing a statistics chapter and not judging whether the overall program is adequate (A6d). Execution of these procedures in fieldwork is GIAS Principle 14 (information gathering and analysis). Sampling as a source of information is also a Section B1 topic later; here you need the conceptual drivers of extent, not formulas.

The operating-effectiveness question

Did the control operate as designed, throughout the period, by the right people, with the right evidence?

Four parts, four ways candidates drop a piece:

  • As designed — you are testing the control you already concluded could work, not a workaround the team invented in week two of fieldwork.
  • Throughout the period — a control that ran in January and November but was waived all of March during a system conversion has an operating issue for that window. A sample taken only in the last week of fieldwork does not cover the year.
  • Right people — delegation of authority, segregation of duties, and competence. An unauthorized clerk who clicks “approve” is not the designed control, even if a box is checked.
  • Right evidence — if the designed control requires a documented, contemporaneous review, “we always look at it” is inquiry, not inspection. A stack of signatures dated the day internal audit arrived is a reliability problem, not comfort.

Operating failure means a sound design was not followed (or not followed by the right person, or not evidenced). It is not a synonym for “we dislike how slow this is” (efficiency, 9.3) and not a synonym for “the control could never have worked” (design, 9.1).

Nature: inquiry, observation, inspection, reperformance

ProcedureWhat you actually doRelative persuasivenessLimit you must plan for
InquiryAsk how, when, and by whom the control is performedLowestPeople describe the ideal; never sufficient alone to conclude operating effectiveness
ObservationWatch the control being performedSnapshotStaff may perform better while watched; one Tuesday does not prove the year
InspectionExamine documents, logs, tickets, emails, workflow history, access extractsStrongerEvidence can be backdated; the report being reviewed may be incomplete (IPE)
ReperformanceThe auditor independently executes the control (re-matches, re-calculates, re-runs the query)Highest among the fourCostly; still needs period coverage and IPE if you started from the entity’s extract

A competent work program combines procedures. Classic pattern: inquire to understand, observe or inspect to see it happening, reperform on a subset, and inspect user IDs against the authority list. The exam item that says the auditor only asked the supervisor whether reviews occur, then concluded the control is effective, is testing whether you know inquiry’s limit.

Observation is underrated and overused. It is excellent for physical or real-time controls (count of goods, dual custody of a token, a live approval in the workflow). It is a weak sole procedure for a year-long detective review.

Inspection is only as good as what you inspect. Initials on a report are evidence that someone marked the report, not that the report contained every exception. That is the IPE problem below.

Reperformance is the closest you get to proving the control can produce the right outcome. Reperforming three-way match on sampled invoices, or re-running the duplicate-payment query with independently obtained parameters, is a different procedure from reading a manager’s tick marks.

Extent: sample sizes conceptually

This is not a statistics textbook. You will see more on choosing information-gathering methods, including sampling, in Section B. For A6b, know what drives how much you plan to test — not a magic n the IIA publishes as law.

Control typeWhat drives extentConceptual approach
Automated application control with effective IT general controlsThe system does the same thing until someone changes itInspect configuration; evaluate access and change-management ITGCs; test a limited instance that the control actually fired (for example, an unmatched invoice was blocked)
Automated control with weak ITGCsConfig or access may have changed unnoticedDo not treat it like a once-and-done application control; plan more instance testing and log review across the period
High-frequency manual (daily)Human error can occur any dayItems spread across the period, including peak volume and cutoff — not 25 invoices all from last Thursday
Low-frequency (weekly / monthly / quarterly)Few occurrences existTest a large fraction or all occurrences rather than a tiny sample of a tiny population
Annual controlOne eventTest that one occurrence in depth (who, evidence, completeness of the population recertified)

Risk still matters. A manual control over vendor bank-account changes deserves more persuasive procedures than a manual check of office-supply coding, even if both are “daily.” Higher risk means stronger nature (more inspection and reperformance) and more attention to period coverage — not a license to test 200 low-risk items and call it quality (that waste is an efficiency issue in 9.3 and a work-program issue in Chapter 10).

Do not invent an IIA-required sample of 25, 40, or 60. Review-course folklore is not the syllabus. The exam tests whether your planned extent matches frequency, automation, and risk, and whether you covered the period.

Automated versus manual versus IT-dependent

Manual control. A person performs a step each time (sign a receiving report, compare a vendor invoice to a PO printout). Plan items across time. Competence and backup coverage matter: the backup who works only at month-end may not know the designed procedure.

Automated application control. The system enforces a rule (match required, tolerance, blocked posting). If ITGCs over logical access and program change are effective, the same configured rule should fire consistently. Design inspection plus limited operating evidence is usually the efficient, sufficient pair. If ITGCs are not effective, you cannot assume the rule was stable.

IT-dependent manual control. A person reviews a system report or dashboard. You must test both the human review and the report. Skipping IPE is the highest-yield A6b trap in this family.

Chapter 10 will go deeper on IT, cyber, and finance methodologies. Here you only need to identify that automated, manual, and hybrid controls are not tested with the same procedure mix.

IPE: information produced by the entity

IPE is a report, extract, query result, spreadsheet, or system listing the organization produces that you either (1) use as audit evidence or (2) rely on because a control uses it (the manager reviews the unmatched-invoice report; the reconciler uses the GL extract).

Before you conclude the review control operated effectively, you need a basis that the IPE is complete and accurate for the control’s purpose:

  • Complete — every item that should be on the report is on it (vendors added mid-month, invoices on hold, users who left last week).
  • Accurate — amounts, dates, IDs, and status flags are right.

Procedures that support IPE reliability include: reconciling the report to a source population (subledger, payment-run file, HR termination list); inspecting report parameters and query logic; comparing record counts; reperforming the extraction with independently obtained criteria; and testing a sample of line items to source documents.

If the unmatched-invoice report omits invoices parked in a second company code, the manager can initial it every Friday and the control still does not operate effectively over the full risk population. Inspection of initials without IPE work overstates effectiveness.

IPE testing is not a fishing expedition through every spreadsheet in the department. Plan it for reports that matter to the control or to your evidence. That is procedure selection, not a data-analytics methodology catalog (Section B).

Right people, right evidence, right period

Inspect user IDs on workflow history against the delegation-of-authority table and against HR hire/termination dates. A terminated approver’s ID that still released payments is an operating (and often access-design) problem.

Inspect timing of evidence. A reconciliation completed before close supports a detective control intended to support reporting. The same reconciliation completed 90 days later may fail timeliness — which can be design if the control could never detect in time, or operating if the designed timetable was simply missed this quarter. Read the designed due date first.

Cover changes in the period: new system, new shared-service team, waiver programs, “touchless” processing go-lives. Operating effectiveness is not an average of two happy months.

Worked example: exception-report review

Designed control: each Friday the AP supervisor reviews the ERP unmatched-invoice exception report, investigates items over $5,000, and signs the report. Design was evaluated: the report exists, the threshold matches policy, and the supervisor’s role cannot also change vendor master.

Operating procedures you should identify in the work program:

  1. Inquire how the Friday review is done and what the supervisor does with old items (inquiry — not sufficient alone).
  2. Inspect signed reports across the period, including a month with peak volume and the month of the ERP patch (inspection + period coverage).
  3. Test IPE: reconcile one report’s population to the unmatched-invoice table, including parked and intercompany items; inspect parameters.
  4. Reperform the investigation for a sample of items over $5,000 (reperformance).
  5. Inspect that the signer was the supervisor on the authority list, not a clerk using shared credentials (right people).

If step 3 shows the report drops a company code, you do not conclude “operating effective because signatures exist.”

Exam traps

  • Concluding effectiveness on inquiry only.
  • Treating one observation as period coverage.
  • Sampling only from the last week of fieldwork.
  • Testing an automated control like a daily manual control (or the reverse) without considering ITGCs.
  • Ignoring IPE completeness and accuracy.
  • Running a large operating sample after a design failure.
  • Confusing a slow but timely detective control with operating ineffectiveness (that fight is Section 9.3).

Operating-effectiveness procedures are complete when the work program can show, for each key control that survived design evaluation, how you will prove it ran all period, by the right people, on reliable information.

Loading diagram...
Nature of procedures for operating effectiveness
Relative persuasiveness for operating effectiveness (illustrative 1–4; combine procedures)
Test Your Knowledge

An auditor inquires of the accounts-payable supervisor, who states that the unmatched-invoice exception report is reviewed every Friday. No other procedure is planned. What is the strongest evaluation of that approach?

A
B
C
D
Test Your Knowledge

A three-way match is configured in the ERP as an automated application control. IT general controls over access and program change have been evaluated as effective. Which mix best tests operating effectiveness of the match?

A
B
C
D
Test Your Knowledge

The AP supervisor initials the Friday unmatched-invoice report each week. Testing shows the report omits invoices parked in a second company code. Which conclusion follows?

A
B
C
D