5.2 Conceptual, Logical, and Physical Data Models and Data Management
Key Takeaways
Enterprise Architecture operates across three levels of data abstraction: Conceptual Data Models (business entities and subject areas), Logical Data Models (normalized entities, attributes, and business rules), and Physical Data Models (database schemas and storage mechanisms).
Conceptual and Logical data models fall squarely under Enterprise Architecture governance to establish shared business vocabularies and canonical schemas, whereas Physical Data Models are typically delegated to implementation teams unless constrained by cross-cutting non-functional requirements.
Master Data Management (MDM) identifies, aggregates, cleanses, and synchronizes core non-transactional business entities (Customer, Product, Vendor, Asset) into an authoritative 'Golden Record' to eliminate siloed duplication.
Master data architectures are deployed through four primary patterns: Registry (federated indexing), Consolidation (analytical reporting hub), Coexistence (synchronized bidirectional hub), and Centralized (single authoritative operational system).
Enterprise data security classification (Public, Internal, Confidential, Restricted/PII) combined with data sovereignty regulations (such as GDPR and HIPAA) dictate mandatory geographical, storage, and lifecycle constraints.
5.2 Conceptual, Logical, and Physical Data Models and Data Management
Data modeling is the foundational technique used in Phase C Data Architecture to describe, structure, and govern information assets. On the TOGAF Enterprise Architecture Practitioner examination, candidates must demonstrate a nuanced understanding of the three levels of data abstraction, recognize the boundary separating enterprise architecture from implementation engineering, and understand data management patterns such as Master Data Management (MDM) and regulatory sovereignty.
The Three Levels of Data Abstraction
To balance high-level strategic alignment with technical precision, enterprise data architecture employs three distinct tiers of data modeling:
1. Conceptual Data Model (CDM)
- Primary Purpose: Establishes a shared business vocabulary and high-level structural understanding of enterprise information for business executives, domain owners, and enterprise architects.
- Content: Represents major Subject Areas and core Business Entities (such as Customer, Account, Product, Agreement, and Claim) along with the fundamental business relationships connecting them (e.g., "Customer holds Account", "Claim references Policy").
- Technology Independence: Completely independent of software packages, database management systems (DBMS), data storage paradigms (relational, document, graph), and implementation technologies.
- ADM Context: Directly translates the business concepts identified during Phase B (Information Mapping and Business Capability Assessment) into structured architectural constructs.
2. Logical Data Model (LDM)
- Primary Purpose: Defines the detailed structure, relationships, business rules, and attributes of enterprise data to guide system integration and software development.
- Content: Fully normalized data entities (typically adhering to 3rd Normal Form or domain-driven entity boundaries), entity attributes, business keys (natural keys), foreign key relationships, and cardinality constraints (1:1, 1:N, M:N).
- Technology Independence: Independent of any specific database vendor (such as Oracle, PostgreSQL, MongoDB, or Snowflake) and hardware platform, but structured with sufficient precision to define canonical message schemas (JSON Schema, XML Schema), API payload contracts, and enterprise data exchange standards.
- ADM Context: Forms the core technical deliverable of Phase C Data Architecture. It defines the canonical information model that bridges business functions with application interfaces.
3. Physical Data Model (PDM)
- Primary Purpose: Specifies the exact technical implementation of data structures optimized for specific database platforms, storage engines, and performance profiles.
- Content: Physical database tables, columns, physical data types (VARCHAR2, INT4, TIMESTAMP), primary key constraints, foreign key indexes, secondary indexes (B-Tree, GIN, Hash), partitioning schemes, sharding keys, tablespaces, and storage parameters.
- Technology Independence: Completely dependent on the target technology platform and hardware environment.
The Enterprise Architecture Boundary: Where Does EA Stop?
In practice, enterprise-level data architecture focuses primarily on conceptual and logical data models; TOGAF's Phase C data artifacts include Conceptual Data and Logical Data diagrams.
Physical Data Models belong almost exclusively to the domain of Solution Architecture, database administration, and software engineering during system design and Phase G (Implementation Governance). An enterprise architect who spends weeks designing individual database table indexes or defining column length constraints is operating outside the proper level of enterprise abstraction—a classic anti-pattern frequently tested on the practitioner exam.
Critical Exceptions Where Physical Concerns Become EA Concerns
While physical schemas are generally delegated, physical data architecture becomes an enterprise architecture concern when it intersects with cross-cutting non-functional requirements (NFRs) or enterprise-level constraints:
- Ultra-Low Latency and High Throughput: When global systems require multi-region distributed databases with sub-millisecond replication (e.g., active-active CockroachDB or AWS Spanner clusters) to meet core business continuity goals.
- Regulatory Data Sovereignty: When physical data residency laws dictate that data must physically reside within specific geographic national borders.
- Enterprise Storage Standardization: When an enterprise establishes a strategic standard mandating cloud-native managed databases and decommissioning on-premises relational databases to reduce total cost of ownership (TCO).
Comparison of Data Modeling Levels
The following table outlines the architectural distinctions across the three modeling tiers:
| Modeling Dimension | Conceptual Data Model (CDM) | Logical Data Model (LDM) | Physical Data Model (PDM) |
|---|---|---|---|
| Primary Audience | Business Executives, Domain Leaders, Enterprise Architects | Data Architects, System Analysts, Integration Engineers | Database Administrators, Software Developers, Cloud Engineers |
| Scope & Breadth | Enterprise-wide or Major Line of Business | Enterprise Domain or Functional Capability | Specific Software System or Database Instance |
| Entity Detail | High-level business entities only | Normalized entities with full attribute lists | Physical tables, collections, views, and schemas |
| Key Constructs | Subject areas and high-level relationships | Business keys, foreign keys, cardinalities | Primary keys, foreign keys, indexes, clustering keys |
| Technology Dependency | Completely technology-neutral | Platform-neutral; software-independent | Vendor-specific (e.g., PostgreSQL, Oracle, MongoDB) |
| Primary TOGAF Artifact | Data Entity / Business Function Matrix | Data Entity / Application Matrix, Canonical Schemas | Database DDL Scripts, Storage Configurations (Phase G) |
| Governance Level | Governed by Architecture Board & Data Owners | Governed by Domain Data Architects & Stewards | Governed by Project Engineering Teams & DBAs |
Master Data Management (MDM) and the Golden Record
In complex enterprises, data is fragmented across dozens of disparate transactional systems. A customer may be recorded in an e-commerce platform, a retail point-of-sale system, a CRM database, and an ERP billing engine—frequently with conflicting names, addresses, and identifiers. Master Data Management (MDM) is the architectural discipline that defines, consolidates, and manages the core non-transactional entities of an organization.
Understanding Data Categories
- Master Data: The core business entities that provide context to business transactions (such as Customer, Product, Supplier, Location, and Chart of Accounts). Master data is characterized by high reuse across multiple systems and relatively low update frequency.
- Transactional Data: Records of business operations and events (such as Sales Orders, Invoices, Credit Card Charges, and Shipments). Characterized by high transaction volume, time-stamped entries, and operational immutability once finalized.
- Reference Data: Semi-static classification tables and standardized code sets (such as ISO Country Codes, Currency Codes, State Abbreviations, and Unit of Measure tables).
The Golden Record (Single Version of the Truth)
The Golden Record is the single, authoritative, cleansed, and deduplicated master record for an entity. When master data is federated across multiple operational platforms, the MDM system applies probabilistic matching algorithms, survivorship rules (e.g., "Trust CRM for phone numbers, but trust Billing for physical addresses"), and data enrichment to assemble the definitive Golden Record.
Core MDM Architectural Patterns
Practitioners must evaluate four distinct MDM deployment patterns based on enterprise trade-offs:
- Registry Pattern:
- How it Works: The MDM hub stores only an index of global identifiers and cross-system pointers. Source systems retain their original data. Queries search the registry to retrieve federated records from source systems.
- Trade-offs: Highly non-intrusive, rapid implementation, and lowest capital cost; however, provides read-only search capabilities and does not clean up data within source systems.
- Consolidation Pattern:
- How it Works: Source systems periodically push their master data via batch or streaming pipelines into a central MDM hub. The hub consolidates, cleanses, and dedupes the data for enterprise analytics and reporting.
- Trade-offs: Excellent for business intelligence, data warehouses, and reporting; however, operational source systems do not receive cleansed master records and continue operating with fragmented data.
- Coexistence Pattern:
- How it Works: Master data is created in source systems, synchronized into a central MDM hub where the Golden Record is generated, and updates are synchronized bidirectionally back to all participating operational systems.
- Trade-offs: Harmonizes operational and analytical environments while maintaining legacy autonomy; however, introduces high integration complexity and potential distributed synchronization conflicts.
- Centralized (Transactional) Pattern:
- How it Works: Master data is authored, updated, and validated exclusively within the central MDM hub. Operational applications must call the MDM hub via APIs or services to interact with master entities.
- Trade-offs: Guarantees highest data quality, complete transactional integrity, and zero duplication; however, requires substantial organizational change and poses a potential single point of failure.
Data Classification, Security, and Lifecycle Management
Enterprise data architecture must incorporate security and regulatory controls into the logical model through formal classification schemas:
- Public: Information approved for unrestricted public distribution (e.g., marketing brochures, press releases). Loss or disclosure presents zero risk.
- Internal: Standard operational business data intended for internal staff (e.g., corporate policies, organizational charts, intranet content). Unauthorized disclosure causes minor disruption.
- Confidential: Sensitive business data whose unauthorized disclosure could cause financial, competitive, or operational damage (e.g., vendor contracts, unannounced product pricing, financial forecasts).
- Restricted / Regulated: Highly sensitive data subject to statutory legal penalties, including Personally Identifiable Information (PII), Protected Health Information (PHI), and Payment Card Data (PCI-DSS). Requires mandatory encryption at rest, encryption in transit, strict role-based access control (RBAC), and immutable audit logging.
The Enterprise Data Lifecycle
Data architecture governs information across six explicit lifecycle phases:
- Acquisition / Creation: Initial capture via user interfaces, external feeds, or IoT sensors.
- Ingestion & Storage: Validation, encryption, and persistence in structured/unstructured repositories.
- Processing & Enrichment: Normalization, business rule execution, and analytical transformation.
- Dissemination & Sharing: Distribution via APIs, event streams, or ETL pipelines.
- Archival: Migration of cold data to low-cost, immutable, compliant storage tiers.
- Destruction & Purging: Cryptographic erasure or physical destruction adhering to statutory retention schedules.
Data Sovereignty and Cross-Border Constraints
Modern cloud-based data architectures operate in a heavily regulated global environment. Enterprise architects must account for Data Sovereignty—the legal concept that digital data is subject to the laws and governance of the geographic jurisdiction where it is physically located or where the subject resides:
- European Union GDPR: Imposes strict extraterritorial mandates regarding the processing of EU citizen data, requiring valid legal transfer mechanisms (such as Standard Contractual Clauses), enforcing data minimization, and granting the "Right to be Forgotten".
- US HIPAA / HITECH: Mandates strict technical and administrative safeguards for electronic health records, including dedicated audit logs and business associate agreements.
- Geopolitical Residency Laws: Multiple nations (including Canada, Germany, India, and China) enforce localized residency mandates requiring citizen financial, health, or personal data to be stored exclusively on physical servers located within national borders.
Architectural Impact: These constraints prevent architects from deploying a single, centralized global database. Instead, architects must design partitioned multi-region topologies, deploy regionalized cloud tenants, implement bring-your-own-key (BYOK) encryption schemes, and use localized data pipelines.
Common Exam Traps & Pitfalls
- Trap 1: Placing Physical DDL Design in Enterprise Architecture: Exam questions often present scenarios where an enterprise architect is spending weeks tuning database indexing algorithms or writing SQL scripts. The correct architectural assessment is that this represents an anti-pattern; physical design belongs to Solution Architecture and engineering teams.
- Trap 2: Confusing Master Data with High-Volume Transaction Data: A distracter option may classify customer invoices, clickstream logs, or sales transactions as "Master Data". These are transactional records that reference master data entities.
- Trap 3: Overlooking Data Sovereignty in Cloud Migrations: In cloud migration scenarios, proposing to consolidate all global regional databases into a single centralized US-based cloud data lakehouse without addressing international data privacy regulations will receive zero points. Legal sovereignty overrides architectural simplicity.
An enterprise architect is tasked with defining a common canonical data model to standardize message exchanges between heterogeneous applications across several business units. Which level of data modeling abstraction is most appropriate for this deliverable?
Physical Data Model, because it specifies the exact column storage lengths and primary key indexing algorithms each application must use
Conceptual Data Model, because it omits all attributes and relationship cardinalities so that business executives can read and approve the message formats
Logical Data Model, because it defines normalized entities, keys, attributes, and relationships independent of any database technology
Hardware Infrastructure Model, because integration formats are determined by the network transport hardware
Which Master Data Management (MDM) architectural pattern cleanses, reconciles, and maintains an authoritative Golden Record in a central hub while synchronizing updates bidirectionally back to operational source systems?
Coexistence Pattern
Registry Pattern
Consolidation Pattern
Analytical Staging Pattern
When designing a target cloud data architecture for a global corporation operating across the European Union and the United States, what primary architectural constraint restricts customer personal data from being consolidated into a single central US cloud data warehouse?
The physical bandwidth constraints of transatlantic undersea optical cables
The Architecture Content Framework metamodel relationship rules in Phase C
The inability of relational database management systems to partition tables across different regions
Data sovereignty and legal regulations such as the European Union General Data Protection Regulation (GDPR)
Sections you finish are checked off in the contents.