5.2 Conceptual, Logical, and Physical Data Models and Data Management

Key Takeaways

  • Enterprise Architecture operates across three levels of data abstraction: Conceptual Data Models (business entities and subject areas), Logical Data Models (normalized entities, attributes, and business rules), and Physical Data Models (database schemas and storage mechanisms).

  • Conceptual and Logical data models fall squarely under Enterprise Architecture governance to establish shared business vocabularies and canonical schemas, whereas Physical Data Models are typically delegated to implementation teams unless constrained by cross-cutting non-functional requirements.

  • Master Data Management (MDM) identifies, aggregates, cleanses, and synchronizes core non-transactional business entities (Customer, Product, Vendor, Asset) into an authoritative 'Golden Record' to eliminate siloed duplication.

  • Master data architectures are deployed through four primary patterns: Registry (federated indexing), Consolidation (analytical reporting hub), Coexistence (synchronized bidirectional hub), and Centralized (single authoritative operational system).

  • Enterprise data security classification (Public, Internal, Confidential, Restricted/PII) combined with data sovereignty regulations (such as GDPR and HIPAA) dictate mandatory geographical, storage, and lifecycle constraints.

Last updated: October 2026

5.2 Conceptual, Logical, and Physical Data Models and Data Management

Data modeling is the foundational technique used in Phase C Data Architecture to describe, structure, and govern information assets. On the TOGAF Enterprise Architecture Practitioner examination, candidates must demonstrate a nuanced understanding of the three levels of data abstraction, recognize the boundary separating enterprise architecture from implementation engineering, and understand data management patterns such as Master Data Management (MDM) and regulatory sovereignty.


The Three Levels of Data Abstraction

To balance high-level strategic alignment with technical precision, enterprise data architecture employs three distinct tiers of data modeling:

1. Conceptual Data Model (CDM)

  • Primary Purpose: Establishes a shared business vocabulary and high-level structural understanding of enterprise information for business executives, domain owners, and enterprise architects.
  • Content: Represents major Subject Areas and core Business Entities (such as Customer, Account, Product, Agreement, and Claim) along with the fundamental business relationships connecting them (e.g., "Customer holds Account", "Claim references Policy").
  • Technology Independence: Completely independent of software packages, database management systems (DBMS), data storage paradigms (relational, document, graph), and implementation technologies.
  • ADM Context: Directly translates the business concepts identified during Phase B (Information Mapping and Business Capability Assessment) into structured architectural constructs.

2. Logical Data Model (LDM)

  • Primary Purpose: Defines the detailed structure, relationships, business rules, and attributes of enterprise data to guide system integration and software development.
  • Content: Fully normalized data entities (typically adhering to 3rd Normal Form or domain-driven entity boundaries), entity attributes, business keys (natural keys), foreign key relationships, and cardinality constraints (1:1, 1:N, M:N).
  • Technology Independence: Independent of any specific database vendor (such as Oracle, PostgreSQL, MongoDB, or Snowflake) and hardware platform, but structured with sufficient precision to define canonical message schemas (JSON Schema, XML Schema), API payload contracts, and enterprise data exchange standards.
  • ADM Context: Forms the core technical deliverable of Phase C Data Architecture. It defines the canonical information model that bridges business functions with application interfaces.

3. Physical Data Model (PDM)

  • Primary Purpose: Specifies the exact technical implementation of data structures optimized for specific database platforms, storage engines, and performance profiles.
  • Content: Physical database tables, columns, physical data types (VARCHAR2, INT4, TIMESTAMP), primary key constraints, foreign key indexes, secondary indexes (B-Tree, GIN, Hash), partitioning schemes, sharding keys, tablespaces, and storage parameters.
  • Technology Independence: Completely dependent on the target technology platform and hardware environment.

The Enterprise Architecture Boundary: Where Does EA Stop?

In practice, enterprise-level data architecture focuses primarily on conceptual and logical data models; TOGAF's Phase C data artifacts include Conceptual Data and Logical Data diagrams.

Physical Data Models belong almost exclusively to the domain of Solution Architecture, database administration, and software engineering during system design and Phase G (Implementation Governance). An enterprise architect who spends weeks designing individual database table indexes or defining column length constraints is operating outside the proper level of enterprise abstraction—a classic anti-pattern frequently tested on the practitioner exam.

Critical Exceptions Where Physical Concerns Become EA Concerns

While physical schemas are generally delegated, physical data architecture becomes an enterprise architecture concern when it intersects with cross-cutting non-functional requirements (NFRs) or enterprise-level constraints:

  1. Ultra-Low Latency and High Throughput: When global systems require multi-region distributed databases with sub-millisecond replication (e.g., active-active CockroachDB or AWS Spanner clusters) to meet core business continuity goals.
  2. Regulatory Data Sovereignty: When physical data residency laws dictate that data must physically reside within specific geographic national borders.
  3. Enterprise Storage Standardization: When an enterprise establishes a strategic standard mandating cloud-native managed databases and decommissioning on-premises relational databases to reduce total cost of ownership (TCO).

Comparison of Data Modeling Levels

The following table outlines the architectural distinctions across the three modeling tiers:

Modeling DimensionConceptual Data Model (CDM)Logical Data Model (LDM)Physical Data Model (PDM)
Primary AudienceBusiness Executives, Domain Leaders, Enterprise ArchitectsData Architects, System Analysts, Integration EngineersDatabase Administrators, Software Developers, Cloud Engineers
Scope & BreadthEnterprise-wide or Major Line of BusinessEnterprise Domain or Functional CapabilitySpecific Software System or Database Instance
Entity DetailHigh-level business entities onlyNormalized entities with full attribute listsPhysical tables, collections, views, and schemas
Key ConstructsSubject areas and high-level relationshipsBusiness keys, foreign keys, cardinalitiesPrimary keys, foreign keys, indexes, clustering keys
Technology DependencyCompletely technology-neutralPlatform-neutral; software-independentVendor-specific (e.g., PostgreSQL, Oracle, MongoDB)
Primary TOGAF ArtifactData Entity / Business Function MatrixData Entity / Application Matrix, Canonical SchemasDatabase DDL Scripts, Storage Configurations (Phase G)
Governance LevelGoverned by Architecture Board & Data OwnersGoverned by Domain Data Architects & StewardsGoverned by Project Engineering Teams & DBAs

Master Data Management (MDM) and the Golden Record

In complex enterprises, data is fragmented across dozens of disparate transactional systems. A customer may be recorded in an e-commerce platform, a retail point-of-sale system, a CRM database, and an ERP billing engine—frequently with conflicting names, addresses, and identifiers. Master Data Management (MDM) is the architectural discipline that defines, consolidates, and manages the core non-transactional entities of an organization.

Understanding Data Categories

  • Master Data: The core business entities that provide context to business transactions (such as Customer, Product, Supplier, Location, and Chart of Accounts). Master data is characterized by high reuse across multiple systems and relatively low update frequency.
  • Transactional Data: Records of business operations and events (such as Sales Orders, Invoices, Credit Card Charges, and Shipments). Characterized by high transaction volume, time-stamped entries, and operational immutability once finalized.
  • Reference Data: Semi-static classification tables and standardized code sets (such as ISO Country Codes, Currency Codes, State Abbreviations, and Unit of Measure tables).

The Golden Record (Single Version of the Truth)

The Golden Record is the single, authoritative, cleansed, and deduplicated master record for an entity. When master data is federated across multiple operational platforms, the MDM system applies probabilistic matching algorithms, survivorship rules (e.g., "Trust CRM for phone numbers, but trust Billing for physical addresses"), and data enrichment to assemble the definitive Golden Record.

Core MDM Architectural Patterns

Practitioners must evaluate four distinct MDM deployment patterns based on enterprise trade-offs:

  1. Registry Pattern:
    • How it Works: The MDM hub stores only an index of global identifiers and cross-system pointers. Source systems retain their original data. Queries search the registry to retrieve federated records from source systems.
    • Trade-offs: Highly non-intrusive, rapid implementation, and lowest capital cost; however, provides read-only search capabilities and does not clean up data within source systems.
  2. Consolidation Pattern:
    • How it Works: Source systems periodically push their master data via batch or streaming pipelines into a central MDM hub. The hub consolidates, cleanses, and dedupes the data for enterprise analytics and reporting.
    • Trade-offs: Excellent for business intelligence, data warehouses, and reporting; however, operational source systems do not receive cleansed master records and continue operating with fragmented data.
  3. Coexistence Pattern:
    • How it Works: Master data is created in source systems, synchronized into a central MDM hub where the Golden Record is generated, and updates are synchronized bidirectionally back to all participating operational systems.
    • Trade-offs: Harmonizes operational and analytical environments while maintaining legacy autonomy; however, introduces high integration complexity and potential distributed synchronization conflicts.
  4. Centralized (Transactional) Pattern:
    • How it Works: Master data is authored, updated, and validated exclusively within the central MDM hub. Operational applications must call the MDM hub via APIs or services to interact with master entities.
    • Trade-offs: Guarantees highest data quality, complete transactional integrity, and zero duplication; however, requires substantial organizational change and poses a potential single point of failure.

Data Classification, Security, and Lifecycle Management

Enterprise data architecture must incorporate security and regulatory controls into the logical model through formal classification schemas:

  • Public: Information approved for unrestricted public distribution (e.g., marketing brochures, press releases). Loss or disclosure presents zero risk.
  • Internal: Standard operational business data intended for internal staff (e.g., corporate policies, organizational charts, intranet content). Unauthorized disclosure causes minor disruption.
  • Confidential: Sensitive business data whose unauthorized disclosure could cause financial, competitive, or operational damage (e.g., vendor contracts, unannounced product pricing, financial forecasts).
  • Restricted / Regulated: Highly sensitive data subject to statutory legal penalties, including Personally Identifiable Information (PII), Protected Health Information (PHI), and Payment Card Data (PCI-DSS). Requires mandatory encryption at rest, encryption in transit, strict role-based access control (RBAC), and immutable audit logging.

The Enterprise Data Lifecycle

Data architecture governs information across six explicit lifecycle phases:

  1. Acquisition / Creation: Initial capture via user interfaces, external feeds, or IoT sensors.
  2. Ingestion & Storage: Validation, encryption, and persistence in structured/unstructured repositories.
  3. Processing & Enrichment: Normalization, business rule execution, and analytical transformation.
  4. Dissemination & Sharing: Distribution via APIs, event streams, or ETL pipelines.
  5. Archival: Migration of cold data to low-cost, immutable, compliant storage tiers.
  6. Destruction & Purging: Cryptographic erasure or physical destruction adhering to statutory retention schedules.

Data Sovereignty and Cross-Border Constraints

Modern cloud-based data architectures operate in a heavily regulated global environment. Enterprise architects must account for Data Sovereignty—the legal concept that digital data is subject to the laws and governance of the geographic jurisdiction where it is physically located or where the subject resides:

  • European Union GDPR: Imposes strict extraterritorial mandates regarding the processing of EU citizen data, requiring valid legal transfer mechanisms (such as Standard Contractual Clauses), enforcing data minimization, and granting the "Right to be Forgotten".
  • US HIPAA / HITECH: Mandates strict technical and administrative safeguards for electronic health records, including dedicated audit logs and business associate agreements.
  • Geopolitical Residency Laws: Multiple nations (including Canada, Germany, India, and China) enforce localized residency mandates requiring citizen financial, health, or personal data to be stored exclusively on physical servers located within national borders.

Architectural Impact: These constraints prevent architects from deploying a single, centralized global database. Instead, architects must design partitioned multi-region topologies, deploy regionalized cloud tenants, implement bring-your-own-key (BYOK) encryption schemes, and use localized data pipelines.


Common Exam Traps & Pitfalls

  • Trap 1: Placing Physical DDL Design in Enterprise Architecture: Exam questions often present scenarios where an enterprise architect is spending weeks tuning database indexing algorithms or writing SQL scripts. The correct architectural assessment is that this represents an anti-pattern; physical design belongs to Solution Architecture and engineering teams.
  • Trap 2: Confusing Master Data with High-Volume Transaction Data: A distracter option may classify customer invoices, clickstream logs, or sales transactions as "Master Data". These are transactional records that reference master data entities.
  • Trap 3: Overlooking Data Sovereignty in Cloud Migrations: In cloud migration scenarios, proposing to consolidate all global regional databases into a single centralized US-based cloud data lakehouse without addressing international data privacy regulations will receive zero points. Legal sovereignty overrides architectural simplicity.
Loading diagram...
Data Modeling Abstraction Tiers and Master Data Management Hub Patterns
Test Your Knowledge

An enterprise architect is tasked with defining a common canonical data model to standardize message exchanges between heterogeneous applications across several business units. Which level of data modeling abstraction is most appropriate for this deliverable?

A

Physical Data Model, because it specifies the exact column storage lengths and primary key indexing algorithms each application must use

B

Conceptual Data Model, because it omits all attributes and relationship cardinalities so that business executives can read and approve the message formats

C

Logical Data Model, because it defines normalized entities, keys, attributes, and relationships independent of any database technology

D

Hardware Infrastructure Model, because integration formats are determined by the network transport hardware

Test Your Knowledge

Which Master Data Management (MDM) architectural pattern cleanses, reconciles, and maintains an authoritative Golden Record in a central hub while synchronizing updates bidirectionally back to operational source systems?

A

Coexistence Pattern

B

Registry Pattern

C

Consolidation Pattern

D

Analytical Staging Pattern

Test Your Knowledge

When designing a target cloud data architecture for a global corporation operating across the European Union and the United States, what primary architectural constraint restricts customer personal data from being consolidated into a single central US cloud data warehouse?

A

The physical bandwidth constraints of transatlantic undersea optical cables

B

The Architecture Content Framework metamodel relationship rules in Phase C

C

The inability of relational database management systems to partition tables across different regions

D

Data sovereignty and legal regulations such as the European Union General Data Protection Regulation (GDPR)

Sections you finish are checked off in the contents.