Exploring the Lineage of Data with Cloud Data Fusion
This lab shows how to use Cloud Data Fusion to explore data lineage: the data's origins and its movement over time.
Cloud Data Fusion data lineage helps you:
- Detect the root cause of bad data events
- Perform an impact analysis prior to making data changes
Cloud Data Fusion provides lineage at the dataset level and field level, and is time-bound to show lineage over time.
Dataset level lineage shows the relationship between datasets and pipelines in a selected time interval.
Field level lineage shows the operations that were performed on a set of fields in the source dataset to produce a different set of fields in the target dataset.
For the purpose of this lab, you will use two pipelines that demonstrate a typical scenario in which raw data is cleaned then sent for downstream processing. This data trail from raw data to the cleaned shipment data to analytic output can be explored using the Cloud Data Fusion lineage feature.
Note: Currently, the Cloud Data Fusion Lineage feature is only available with the Cloud Data Fusion Enterprise Edition.
Join Qwiklabs to read the rest of this lab...and more!
- Get temporary access to the Google Cloud Console.
- Over 200 labs from beginner to advanced levels.
- Bite-sized so you can learn at your own pace.
Create a Cloud Data Fusion instance
Add Cloud Data Fusion API Service Agent role to service account
Import, Deploy and Run Shipment Data Cleansing pipeline
Import, Deploy, and Run the Delayed Shipments data pipeline