menu
arrow_back

Exploring the Lineage of Data with Cloud Data Fusion

Exploring the Lineage of Data with Cloud Data Fusion

1 hora 30 minutos 7 créditos

GSP812

Google Cloud Self-Paced Labs

Overview

This lab shows how to use Cloud Data Fusion to explore data lineage: the data's origins and its movement over time.

Cloud Data Fusion data lineage helps you:

  • Detect the root cause of bad data events
  • Perform an impact analysis prior to making data changes

Cloud Data Fusion provides lineage at the dataset level and field level, and is time-bound to show lineage over time.

  • Dataset level lineage shows the relationship between datasets and pipelines in a selected time interval.

  • Field level lineage shows the operations that were performed on a set of fields in the source dataset to produce a different set of fields in the target dataset.

For the purpose of this lab, you will use two pipelines that demonstrate a typical scenario in which raw data is cleaned then sent for downstream processing. This data trail from raw data to the cleaned shipment data to analytic output can be explored using the Cloud Data Fusion lineage feature.

Note: Currently, the Cloud Data Fusion Lineage feature is only available with the Cloud Data Fusion Enterprise Edition.

Únase a Qwiklabs para leer este lab completo… y mucho más.

  • Obtenga acceso temporal a Google Cloud Console.
  • Más de 200 labs para principiantes y niveles avanzados.
  • El contenido se presenta de a poco para que pueda aprender a su propio ritmo.
Únase para comenzar este lab
Puntuación

—/100

Create a Cloud Data Fusion instance

Ejecutar paso

/ 25

Add Cloud Data Fusion API Service Agent role to service account

Ejecutar paso

/ 25

Import, Deploy and Run Shipment Data Cleansing pipeline

Ejecutar paso

/ 25

Import, Deploy, and Run the Delayed Shipments data pipeline

Ejecutar paso

/ 25