1.2 Sources of Data, Big Data & Data Capture Costs
Key Takeaways
- Modern management data originates across three primary technical conduits: machine/sensor devices (IoT, telemetry), transactional processing systems (ERP, invoicing), and human/social interactions (sentiment, surveys).
- Data is classified into internal versus external sources, and primary (collected first-hand for a specific purpose) versus secondary (repurposed pre-existing datasets).
- The economic rule of management accounting requires that data only be captured and processed if its expected marginal value exceeds the full cost of collection and storage.
- Big Data is defined by the 5 Vs: Volume, Velocity, Variety, Veracity, and Value, encompassing structured, semi-structured, and unstructured data architectures.
- Organizational data analytics progresses across four analytical tiers: descriptive (what happened), diagnostic (why it happened), predictive (what will happen), and prescriptive (what should be done).
Sources of Data, Big Data & Data Capture Costs
Core Principle: Information is an economic resource that carries substantial acquisition, processing, and maintenance costs. A well-designed management accounting system must identify the most reliable data conduits—balancing machine, transactional, and human sources—while rigorously enforcing the rule that data should only be captured when its expected decision-making value exceeds its total cost.
1. Three Primary Sources of Data by Technical Origin
In modern digitized enterprise environments, raw data flows into accounting and enterprise resource planning (ERP) systems from three primary technical conduits:
┌─────────────────────────────────┐
│ Data Conduits in the Firm │
└────────────────┬────────────────┘
│
┌───────────────────────────────┼───────────────────────────────┐
▼ ▼ ▼
┌───────────────────┐ ┌───────────────────┐ ┌───────────────────┐
│ Machine / Sensor │ │ Transactional │ │ Human / Social │
│ • RFID inventory │ │ • Sales invoices │ │ • Customer reviews│
│ • GPS fleet logs │ │ • Purchase orders │ │ • Staff surveys │
│ • Machine sensors │ │ • Payroll ledgers │ │ • Social sentiment│
└───────────────────┘ └───────────────────┘ └───────────────────┘
1. Machine and Sensor Data
Generated automatically by hardware, physical instruments, and internet-connected devices without human intervention:
- Internet of Things (IoT) & Smart Sensors: Temperature monitors in refrigerated food supply chains, vibration sensors on factory milling machines, and smart energy meters tracking kilowatt-hour consumption by department.
- Radio Frequency Identification (RFID) & Optical Scanners: Automated warehouse pallets scanning across dock doors, updating inventory balances in real time.
- Telematics & GPS Tracking: Real-time route, mileage, and fuel consumption logs transmitted continuously by commercial transport fleets.
- Operational Characteristics: Characterised by immense volume, continuous high velocity, and consistent machine formatting. It provides objective, real-time operational transparency with minimal transcription error.
2. Transactional Data
Generated directly from commercial exchanges and formal business events recorded within the organisation's core accounting and operational ledgers:
- Sales & Revenue Transactions: Point-of-sale (POS) terminal receipts, e-commerce order checkouts, credit sales invoices, and customer credit notes.
- Purchasing & Operational Payables: Purchase orders, goods received notes (GRNs), supplier invoices, and electronic payment records.
- Payroll & Human Capital: Clock-card biometric logs, hourly wage runs, piece-work bonus tickets, and statutory payroll deductions.
- Operational Characteristics: Highly structured, historically accurate, auditable, and stored in relational database tables following double-entry rules. It forms the backbone of traditional cost accounting.
3. Human and Social Data
Originates from human communication, subjective perceptions, behavioral interactions, and external social discourse:
- Customer Interaction Channels: Online reviews, customer service call audio recordings, support ticket text chats, and complaint logs.
- Internal Human Feedback: Annual employee engagement surveys, exit interviews, suggestion scheme submissions, and performance review qualitative notes.
- Social Media & Market Discourse: Brand sentiment scores on social networks, industry forum discussions, and consumer feedback on competitor offerings.
- Operational Characteristics: Predominantly unstructured or qualitative, rich in commercial context, but challenging to standardize, verify, and quantify for formal cost accounting ledgers.
2. Classification of Data: Origin and Collection Method
Beyond technical origin, management accountants categorize data across two crucial intersecting dimensions:
Data Classification Matrix
[ Origin Dimension ]
Internal External
┌──────────────┬──────────────┐
Primary │ Internal │ External │
│ Primary Data │ Primary Data │
[ Purpose ────────┼──────────────┼──────────────┤
Dimension ] │ Internal │ External │
Secondary │ Secondary │ Secondary │
│ Data │ Data │
└──────────────┴──────────────┘
Internal versus External Data Sources
- Internal Data: Generated entirely within the organisation's administrative, production, and commercial boundaries. Examples include material requisition slips, factory machine logs, production job cards, and historical sales ledgers.
- Advantages: Readily accessible, confidential, directly tailored to the business's specific chart of accounts, and low incremental collection cost.
- Limitations: Inward-looking; fails to provide context regarding competitor actions, macroeconomic threats, customer shifts, or industry-wide cost pressures.
- External Data: Originates outside the enterprise. Examples include published supplier price indices, central bank interest rate announcements, competitor financial filings, and industry association salary surveys.
- Advantages: Provides vital macro intelligence, external benchmarking standards, and market demand context necessary for strategic planning.
- Limitations: Can be expensive to license, may suffer from publication lag, and often uses classifications or definitions that do not match the company's internal cost centres.
Primary versus Secondary Data Sources
- Primary Data: Data collected first-hand for the specific, explicit investigation or decision currently facing management. Examples include conducting a bespoke time-and-motion study on an assembly line to establish new direct labour standard times, or running a targeted customer survey to assess willingness to pay for a new product feature.
- Advantages: Precisely tailored to the immediate analytical question, fresh, and fully controlled by the investigator.
- Limitations: Resource-intensive; incurs substantial expenditure of time and money to design, gather, and analyse.
- Secondary Data: Pre-existing data that was originally collected for another purpose (either by another department or an external agency) and is now repurposed for a new decision. Examples include using published government census data to estimate regional demand, or using historical statutory accounts to benchmark overhead ratios.
- Advantages: Instantly accessible, inexpensive, and immediately usable.
- Limitations: May be outdated, may not align cleanly with the required units of analysis, and may reflect sampling biases that cannot be verified.
3. Published Data and Internet Information: Uses and Limitations
Management accountants increasingly incorporate published external data and open internet sources into forecasting, standard cost setting, and variance benchmarking.
Core Uses in Management Accounting
- Establishing Benchmark Standards: Utilizing commodity exchange indices (e.g., London Metal Exchange spot and futures prices) to negotiate supplier contracts and establish realistic direct material standard prices.
- Labour Cost Projections: Incorporating published national statistics on inflation (CPI, RPI) and regional wage indices when projecting annual factory labour rate revisions.
- Competitor Benchmarking: Examining publicly disclosed annual reports (under IFRS) of listed industry competitors to evaluate their gross margins, inventory turnover days, and research and development spending ratios.
Limitations of Published and Internet Data
- Veracity and Authenticity Risks: Open internet data often lacks rigorous editorial or academic peer-review. Sources may exhibit commercial sponsorship bias or publish unverified claims.
- Definitional Mismatches: External datasets may define terms differently. For example, an industry benchmark for "cost of sales" may include distribution freight, whereas the company treats carriage outwards as a non-production selling expense.
- Historical Time Lags: Government census or economic survey results are frequently published months or years after the data was captured, limiting their predictive reliability in volatile markets.
4. The Economics of Information: Capture Costs vs. Decision Value
Information is not free. Incurring costs to obtain marginal refinements in data accuracy is only rational if the improved data produces superior commercial outcomes.
The Cost Components of Information
Capturing, processing, and maintaining management information incurs costs across four categories:
- Capture and Ingestion Costs: Hardware installation (barcode scanners, RFID readers, industrial sensors), bespoke market research fees, software API licenses, and manual clerical data-entry wages.
- Processing and Storage Costs: Enterprise cloud hosting infrastructure, relational database server clustering, data warehousing subscriptions, and IT cybersecurity protocols.
- Analysis and Presentation Costs: Salaries of management accountants, business intelligence specialists, data scientists, and reporting visualization software licenses.
- Opportunity and Delay Costs: The commercial losses incurred while delaying a critical market decision while waiting for an elaborate data report to be compiled.
Valuing Information: The Cost-Benefit Equilibrium
The value of information is defined as the incremental financial benefit (additional revenue generated or cost avoided) resulting from a decision made with the information compared to a decision made without it:
If refining an inventory tracking model costs $50,000 in software integration but only reduces stockholding obsolescence by an expected $12,000 annually, the capture of that additional granular data destroys shareholder value and should be rejected.
5. Big Data in Management Accounting: The 5 Vs
The emergence of ubiquitous computing, cloud architecture, and digital commerce has created Big Data—datasets so immense, fast-moving, and structurally diverse that conventional relational database systems cannot process them efficiently.
The 5 Vs of Big Data
┌──────────────────────────────────┬──────────────────────────────────┐
│ Volume │ Velocity │
│ • Terabytes, petabytes of logs │ • Real-time continuous streaming │
│ • Millions of daily clicks │ • Sub-second pricing adjustments │
├──────────────────────────────────┼──────────────────────────────────┤
│ Variety │ Veracity │
│ • Structured (SQL ledgers) │ • Noise filtering, data hygiene │
│ • Semi-structured (JSON/XML) │ • Mitigating data bias & errors │
│ • Unstructured (Video, audio) │ │
├──────────────────────────────────┴──────────────────────────────────┤
│ Value │
│ • Actionable commercial insights, cost reductions, optimized margins│
└─────────────────────────────────────────────────────────────────────┘
1. Volume
The sheer scale and magnitude of data generated. Where traditional management accounting dealt with thousands of monthly general ledger rows, Big Data systems process petabytes ($10^{15}$ bytes) containing billions of individual records, including clickstream logs, sensor readings, and mobile device pings.
2. Velocity
The unprecedented speed at which new data is generated, collected, and required to be processed. Rather than relying on periodic month-end reconciliation runs, modern systems process streaming data continuously to enable real-time dynamic pricing, instantaneous inventory replenishment, and live fraud detection.
3. Variety
The structural diversity of data formats arriving at the organization:
- Structured Data: Highly organized, tabular data residing in fixed fields within relational databases (e.g., standard cost cards, general ledger balances, customer invoice records).
- Semi-Structured Data: Data containing internal markers, tags, or hierarchical schemas without a rigid relational table structure (e.g., JSON payloads, XML feeds, web server log files).
- Unstructured Data: Data with no predefined conceptual structure, comprising over 80% of modern enterprise data (e.g., customer service phone call audio, surveillance camera video feeds, free-form email inquiries, PDF invoices).
4. Veracity
The truthfulness, reliability, accuracy, and noise level within the data. Big Data feeds are frequently messy, containing duplicate records, fake reviews, transmission glitches, and measurement noise. Management accountants must establish robust data governance and cleansing procedures to prevent flawed data from corrupting costing algorithms.
5. Value
The ultimate business imperative. Immense volumes of high-velocity data remain an expensive liability unless converted into actionable insights that enhance operating profit, reduce waste, improve customer retention, or uncover cost efficiencies.
6. The Four Stages of Organizational Analytics
Management accountants apply Big Data and business intelligence across four progressively sophisticated analytical stages:
| Analytics Tier | Core Question | Management Accounting Application |
|---|---|---|
| Descriptive | "What happened?" | Historical variance reporting; monthly standard cost reconciliations; customer profitability summaries. |
| Diagnostic | "Why did it happen?" | Root-cause analysis; drilling into adverse material price variances to correlate with specific supplier batches or freight carrier delays. |
| Predictive | "What will happen?" | Econometric regression modeling for forecasting electricity costs; machine learning algorithms predicting customer default probabilities and bad debts. |
| Prescriptive | "What should we do?" | Algorithmic dynamic pricing that automatically adjusts e-commerce prices based on real-time competitor stockouts; linear programming models solving complex production bottleneck constraints. |
A multinational logistics firm equips its delivery fleet with engine telemetry sensors and GPS transponders that continuously transmit speed, idle time, and fuel consumption logs to its operations server. How should this data stream be classified?
A company is considering purchasing an external industry analytics subscription costing $40,000 per year. The management accountant estimates that having access to this data will improve inventory scheduling and prevent stockouts, yielding an expected annual cost saving of $28,000. Applying the cost versus value principle, what should management do?
In the context of the 5 Vs of Big Data, which dimension specifically addresses data hygiene, noise filtering, and the trustworthiness of incoming unstructured feeds?