7.2 Datasets, Profile Enablement & Ingestion Troubleshooting

Key Takeaways

  • All ingested dataset data can reside in the Data Lake, but only appropriately configured profile-enabled schema and dataset data contributes to Real-Time Customer Profile.

  • Enabling a schema for Profile does not automatically enable every dataset created from it.

  • Batch and streaming ingestion have different monitoring objects and timing; neither should be assigned an invented universal latency.

  • Troubleshooting starts with sandbox, source/connection, dataset, batch or dataflow status, schema validation, identity, and Profile enablement.

  • Dataset preview proves stored rows, while profile inspection proves that eligible data was combined into Profile.

Last updated: October 2026

7.2 Datasets, Profile Enablement, and Ingestion Troubleshooting

A dataset can contain correct rows and still be unusable in a journey. Follow the data through storage, validation, identity, Profile, audience evaluation, and the journey instead of diagnosing from one screen.

Data Lake versus Profile

When data is ingested into an Adobe Experience Platform dataset, it is stored in the Data Lake for supported platform uses. That alone does not make the data part of Real-Time Customer Profile.

For data to contribute to Profile:

  1. the schema must be eligible and enabled for Profile;
  2. the schema must have the required identity configuration;
  3. the dataset must also be enabled for Profile;
  4. ingested records must conform to schema and identity expectations;
  5. Profile processing must complete.

A schema's Profile toggle does not automatically enable every dataset created from it. This distinction is a frequent exam scenario.

Profile enablement is a consequential design choice. Confirm identity quality, data volume, merge behavior, consent, and retention before enabling a high-volume dataset.

Batch ingestion

Batch ingestion loads a bounded file or set of files. Monitor the batch status, record counts, failed records, and diagnostic messages. A completed batch means its accepted rows reached the dataset; Profile and audience processing can still follow afterward.

For CSV or delimited data, common issues include header mapping, delimiters, quoted values, date formats, empty required values, and datatype mismatches. For JSON, inspect nesting, arrays, and exact field paths.

Preview the dataset to confirm actual stored values. Do not infer success from the source system's export job alone.

Streaming ingestion

Streaming sends records or events continuously through a configured connection, datastream, or API. Monitor the source or dataflow and validate representative payloads. Processing latency depends on the ingestion route, schema, Profile, audience evaluation, and downstream service; avoid promising a fixed number of seconds.

A successful request can still contain data that fails validation or is routed to the wrong sandbox or dataset. Check both the source response and Platform monitoring.

Layer-by-layer troubleshooting

1. Environment: Confirm organization, sandbox, and region/context.

2. Source: Did the upstream system produce the record with the expected timestamp and identity?

3. Connection/dataflow: Is the mapping active and healthy? Are credentials or datastream settings current?

4. Dataset: Is the destination correct? Does preview show the record? Did the batch or stream report failures?

5. Schema: Do path and datatype match? Are required event fields present?

6. Identity: Is the value present under the intended namespace? Is the primary identity configured appropriately?

7. Profile enablement: Are both schema and dataset enabled when Profile use is required?

8. Profile: Does the Profile viewer show the attribute or event for the identity?

9. Audience: Has evaluation run, and is the profile a member?

10. Journey: Does the entry mechanism, namespace, schedule, and reentrance permit entry?

This sequence avoids changing a journey when the record never reached Profile.

Common examples

Rows in dataset, absent from Profile: Check dataset Profile enablement, schema eligibility, identity, validation errors, and processing time.

Profile visible, absent from batch audience: Confirm audience definition and that the latest batch evaluation completed.

Audience membership visible, no Read Audience entry: Check schedule, trigger-after-batch settings, namespace, incremental logic, and reentrance.

Event stored, no unitary entrance: Check event configuration, recognition rule, selected schema, identity namespace, and whether the event arrived in the same sandbox.

Some rows rejected: Download or inspect diagnostics and group failures by field/path instead of repeatedly re-uploading unchanged data.

Data quality and governance

Profile-enabled data can influence personalization and eligibility quickly, so correctness matters. Reject impossible dates, normalize enumerated values upstream, preserve source lineage, and avoid marking unstable device or shared values as person identities. Apply labels to governed fields and datasets as appropriate.

Monitor recurring imports and exports. A journey that worked yesterday can fail today because a scheduled file changed format, credentials expired, or the source stopped populating identity.

Warning

Dataset preview and Profile view answer different questions. Preview confirms storage; Profile view confirms that eligible data was incorporated under the identity and merge policy.

Reconciliation counts

Track counts at each stage: source rows, accepted dataset records, rejected records, identities incorporated into Profile, audience members, journey entrances, and channel-eligible recipients. Differences are expected, but every material difference should have an explanation. This funnel localizes a defect faster than repeatedly refreshing only the final audience or journey report.

Test Your Knowledge

A schema is Profile-enabled, but a dataset created from it is not. What should the practitioner expect?

A

Its rows automatically contribute to Profile anyway.

B

The data can be stored in the Data Lake but does not contribute through that dataset to Profile until the dataset is enabled.

C

The schema converts to an offer.

D

Every audience becomes edge-evaluated.

Test Your Knowledge

Dataset preview shows a row, but the profile does not. What should be checked next?

A

The email font

B

The exam price

C

Profile enablement, identity, validation, merge behavior, and processing

D

The fallback offer only

Test Your Knowledge

Why avoid promising one fixed streaming latency?

A

Streaming is always slower than batch.

B

Streaming never uses schemas.

C

Every streaming record becomes a message immediately.

D

End-to-end timing depends on ingestion, validation, Profile, audience, and journey processing.

Sections you finish are checked off in the contents.