11.2 Data Analysis as an Information Source

Key Takeaways

  • B1a includes data analyses as a method of obtaining information, alongside interviews, observations, and walk-throughs.
  • In the preliminary survey, full-population scans (duplicates, round-dollar amounts, weekend or holiday postings, sequence gaps) are used to focus later tests — not to issue final conclusions.
  • Diagnostic, prescriptive, and predictive analytics as an engagement product are Chapter 14; do not import that catalog into the survey.
  • Information produced by the entity (IPE) must be complete and accurate enough for the scan's purpose; analyzing an incomplete extract aims later tests at the wrong population.
  • Survey output is a list of areas of focus, refined procedures, and populations worth examining — not a finding that fraud occurred because 18 duplicates appeared on a first-pass report.
Last updated: August 2026

Data analysis is a B1a method, not a closing argument

Quick Answer: Data analyses are the remaining B1a method of obtaining information, with interviews, observations, and walk-throughs. In the preliminary survey, you use data to focus later tests: full-population exception scans for duplicates, round-dollar amounts, weekend postings, and similar anomalies. That output is a map of where to look. Diagnostic and predictive analytics as analysis methods are Chapter 14. Do not analyze garbage: if the extract is incomplete, the exception list is not a picture of risk.

A walk-through showed you one unmatched invoice that still paid. Data analysis asks a different survey question: how common is that pattern in the whole file? You do not need a data-scientist title to use the method. You need a defined objective, a population that matches the process you just walked, and the humility to treat the first report as a lead list.

GIAS Standard 14.1 still sits in the background: you gather information in order to analyze and evaluate. This section is only the source. Chapter 12 is whether that source is relevant, sufficient, and reliable enough to support a finding. Chapter 14 is when to use diagnostic, prescriptive, predictive, anomaly, and text methods, and how to run the analytics process (objectives, obtain, normalize, analyze, communicate). If an item asks which survey procedure will best aim AP tests at duplicate payments, the answer is a population scan, not a neural net.

Survey analytics versus later analytics

Survey data analysis (this section)Later / Chapter 14 analytics
JobObtain information that focuses proceduresAnalyze information to support findings; choose diagnostic, prescriptive, predictive, anomaly, or text methods
Typical workFull-population exception scans; simple sorts and matches; counts of red-flag typesRoot-cause and effect analysis; models; dashboards as an engagement product
Unit of success"18 exact duplicates and 9 weekend postings — test those populations""Duplicates cluster in vendor 4412 after a bank-detail change; condition vs criteria; significance"
DangerCalling the scan a conclusion ("fraud") or skipping IPERunning a model on a file you never reconciled to the process

Survey analytics are usually descriptive and exception-seeking. You are allowed to notice a pattern. You are not allowed to skip the investigation that turns a pattern into a condition. Eighteen duplicate invoice numbers can be a system-assign glitch, a re-post after a failed payment, or a second payment to a real vendor. The scan tells you which rows to walk and which vendors to interview. It does not tell the audit committee that fraud occurred.

Full-population exception scans that earn their keep

When the process posts to a system you can extract, a full-population scan in the survey is often cheaper than guessing which 40 items to sample. You are still in B1 (source of information), not in sampling theory. Common survey scans:

Duplicates. Exact matches on vendor + invoice number + amount catch the obvious double-pay. Near-duplicates (same vendor and amount, invoice number off by one character; same invoice number, two vendor IDs) catch splits and master-file clones. The survey output is a list, a dollar total, and a decision: examine all of them, or stratify and test the large ones plus a slice of the rest.

Round-dollar amounts. Invoices ending in 000 or 00.00 are not illegal. They are over-represented in manual journals, estimated accruals, and "just pay it" overrides. A survey count that shows 41 round-dollar vendor invoices over $10,000 in a process that is supposed to be three-way matched is a reason to walk those items and inspect override user IDs. It is not a finding that round-dollar equals fictitious.

Weekend and holiday postings. If the designed review happens on business-day afternoons, postings time-stamped Saturday 02:14 either bypassed the review or used a shared ID. Pair the scan with the interview hint about "the weekend upload." Nine weekend postings in a quarter is a focus area for later tests of the upload job, access, and whether the Friday proposal review ever saw those invoices.

Other survey-useful scans. Sequence gaps in check or invoice numbers; vendors with the same bank account as an employee; payments just under an approval threshold; users who both created a vendor and released a payment. Pick scans that match this engagement's risks, not a generic "run Benford because we have ACL." A payroll engagement cares about duplicate bank accounts and hours over a plausible week; a treasury engagement cares about beneficiary changes and off-hours wires.

ScanWhat it can flag in surveyWhat it cannot prove by itself
Exact duplicate vendor + invoice + amountPossible double payment or re-postIntent, loss, or that the control never operated
Round-dollar over a thresholdManual or override-heavy itemsThat round amounts are fictitious
Weekend / holiday posting timesActivity outside the designed review windowThat every weekend posting is unauthorized
Same employee and vendor bank accountA conflict worth walkingPayroll fraud without HR and payment-file work
Amounts just under approval limitsPossible splittingThat splitting occurred on every item near the limit

Worked example: AP survey scan

You extract the AP invoice register for the quarter after reconciling record count and total to the subledger (IPE, below). Scans return 18 exact duplicates ($612,000), 41 round-dollar invoices over $10,000, 9 weekend postings, and 12 same-vendor same-invoice-number pairs with different amounts. You do not draft four findings. You:

  1. Put duplicate and weekend populations on the work program as examine-all or high-priority tests.
  2. Walk two round-dollar overrides to see whether the same user ID completed match and override (tying back to 11.1).
  3. Interview the clerk who mentioned weekend uploads, now with nine timestamps on the table.
  4. Leave predictive scoring of "likely fraudulent vendor" for a Chapter 14 discussion if the function even does that work — it is not the survey product.

The chart below is the survey output: counts that aim hours. Dollar investigation, root cause, and significance wait until you have evidence Chapter 12 would respect.

Illustrative AP survey exception counts (focus later tests; do not conclude yet)
Loading diagram...
Survey data analysis produces focus areas, not engagement conclusions

IPE: do not analyze garbage

Information produced by the entity (IPE) is the report, extract, query, or spreadsheet the organization produces that you are about to treat as the population. Chapter 9.2 introduced IPE when a control is a human review of a system report. Here the IPE is the source you are scanning. If company code 200 is missing, your "no weekend postings" result is a statement about company code 100. If the date parameter is invoice-entry date and you think it is posting date, weekend flags are fiction.

Before you trust a survey scan:

  • Complete — record count and total dollars reconcile to the subledger or process file that the walk-through used; parked, held, intercompany, and voided items are included or excluded on purpose.
  • Accurate — a sample of lines traces to source documents for vendor, amount, and date; user IDs are the live IDs, not display names that several people share.
  • Parameters — period, company codes, document types, and status flags match the activity under review. Screenshot or save the query.

You do not need Chapter 12's full reliability rubric to refuse a bad file. You need the survey discipline: wrong population in, wrong focus out. If IT cannot get a complete extract this week, say so in the work program, use another method (walk-throughs, observations) for those days, and do not pretend a partial scan covered the plant you never received.

How far IPE work has to go before a finding can rest on the scan is an evidence-quality question (Chapter 12). This section only stops you from aiming the engagement at a hole in the extract.

Planning output: areas of focus, not final conclusions

Write the survey analytics so a reviewer sees:

  • The objective the scan served (focus duplicate-payment tests; see whether weekend activity exists).
  • The population and IPE steps.
  • Exception counts and dollars as leads.
  • The procedure changes: examine-all duplicates; add a walk-through of the weekend job; inspect override IDs on round-dollar items; drop a low-value scan that matched no engagement risk.

Do not write "management's AP process is ineffective" from a first-pass duplicate list. Do not skip interviews and walk-throughs because the spreadsheet is colorful. Do not run a full predictive model in week one and call B1a complete — that is a different syllabus bullet.

Exam traps

  • Treating a survey exception scan as proof of fraud or as a substitute for period testing.
  • Analyzing an extract you never reconciled to the process (garbage in).
  • Importing Chapter 14's diagnostic/predictive catalog into a B1a method-choice item.
  • Skipping data analysis entirely on a high-volume automated process because "we always sample 40."
  • Using scans that do not match this activity's risks (Benford on a 12-invoice grant, weekend flags on a plant that posts only on Saturdays by design).

Survey data analysis is done when later procedures have a named population to hit, and when you can explain why that population is the live file, not a convenient slice of it.

Test Your Knowledge

During the preliminary survey, a full-population scan finds 18 exact-duplicate AP invoices totaling $612,000. What is the appropriate use of that result?

A
B
C
D
Test Your Knowledge

Why must the auditor test information produced by the entity before relying on a survey exception report?

A
B
C
D
Test Your Knowledge

How should data analysis in the preliminary survey differ from later diagnostic or predictive analytics?

A
B
C
D