Dataset Stages
A stage is a dataset component that takes in data, transforms it in some way, and outputs modified data. A stage performs a single computational function, such as joining data with a Join stage or importing data from the dataset source with an Import stage.
Adding a stage can change the number of records and fields in the dataset.
Prism Analytics supports these stage types:
Stage Type | Description |
|---|---|
Import | Workday creates this stage automatically as the first stage in derived datasets because they import a table or another dataset as the source. The first stage in every pipeline in a derived dataset is an Import stage. An Import stage imports data from a source, and for derived datasets that source is the output of a table or an existing dataset. You can't delete or add an Import stage. |
Parse | A Parse stage enables you to describe the source data in a tabular format. Workday creates this stage automatically for base datasets. Workday automatically displays the Parse stage when you create a dataset using external data. You can't delete or add a Parse stage.
If you're new to Workday, you don't have access to create or edit base datasets. |
Explode | Use an Explode stage to convert a multi-instance field into an instance field. Workday takes each instance value in the multi-instance field and creates a new record for each value. You can add or delete an Explode stage in the middle or at the end of a pipeline.
For details, see Reference: Explode Stages. |
Filter | Use a Filter stage to constrain the number records in the dataset based on a filter condition. Add Filter stages to limit the data in the dataset for analysis, such as on a particular region or year. You can add or delete a Filter stage in the middle or at the end of a pipeline.
For details, see: Reference: Filter Stages. |
Group By | Use a Group By stage to aggregate records of data into groups. You can summarize the values using a summarization type, such as MIN, MAX, or SUM. Add Group By stages to get data to the appropriate level that you need to join the data in one dataset with another dataset. You can add or delete a Group By stage in the middle or at the end of a pipeline.
For details, see: Reference: Group By Stages. |
Join | Use a Join stage to join the data in this dataset with the data that you import from another dataset or table based on the relationship between a field in each set of data.
A Join stage combines fields from 2 dataset pipelines based on common values that exist in each pipeline. Add Join stages to view and use related data from different datasets. You can add Join stages to derived datasets only. You can add or delete a Join stage anywhere in the Primary Pipeline. Deleting a Join stage disconnects but keeps the stage's pipeline. If you don't want to keep the disconnected pipeline, delete it or use it in another Join stage. For details, see: Reference: Join Stages. |
Manage Fields | Use a Manage Fields stage to view field changes, select fields, hide fields, or edit fields. You can add or delete a Manage Fields stage in the middle or at the end of a pipeline.
For details, see: Manage Dataset Fields. |
Union | Use a Union stage to combine the records from this dataset with the records from another dataset or table that you import.
A Union stage combines data from similar fields in different datasets into a single field. Add Union stages to combine datasets that have similar, but not identical schema, and to combine datasets with the same schema but with source data from different locations, such as different SFTP servers. You can add Union stages to derived datasets only. You can add or delete a Union stage anywhere in the Primary Pipeline. Deleting a Union stage disconnects but keeps the stage's pipeline. If you don't want to keep the disconnected pipeline, delete it or use it in another Union stage. For details, see: Reference: Union Stages. |
Unpivot | Use an Unpivot stage to convert fields (columns) to records (rows) in the dataset. Add an Unpivot stage to consolidate data from 2 or more similar fields into a pair of new fields. You can add Unpivot stages to derived datasets only. You can add or delete an Unpivot stage in the middle or at the end of a pipeline.
|
Validation | Use the Validation stage to ensure data integrity in your data pipeline. This stage uses predefined rules to catch duplicates or errors before they reach your production reports. For details, see Reference: Validation Stages. |