Versioning
What is reproducibility in a data pipeline?
Reproducibility means you can inspect the same inputs and column definitions and get the same cached outputs, or a known diff.
Updated · Part of What is time travel?
How it works
- Inputs are versioned rows.
- Expressions are the column definitions.
- Caches mean a rerun does not call the model again unless an input changed.
What it is not
It is not only setting a random seed in a training script.
reproducibility: this, and the thing it is confused with
| This | Not this | |
|---|---|---|
| Same inputs | Same cached outputs | A new sample from the model |
| Diff | Visible between versions | A different notebook result |
| Seed | Not the whole story | The usual ML lever |
Where Pixeltable fits
Pixeltable caches computed columns and versions the table, which is what makes a rerun comparable.
Questions
- How does reproducibility work?
- Inputs are versioned rows. Expressions are the column definitions. Caches mean a rerun does not call the model again unless an input changed.
- What is reproducibility often confused with?
- It is not only setting a random seed in a training script.