Versioning

What is reproducibility in a data pipeline?

Reproducibility means you can inspect the same inputs and column definitions and get the same cached outputs, or a known diff.

Updated · Part of What is time travel?

How it works

  • Inputs are versioned rows.
  • Expressions are the column definitions.
  • Caches mean a rerun does not call the model again unless an input changed.

What it is not

It is not only setting a random seed in a training script.

reproducibility: this, and the thing it is confused with

reproducibility: this, and the thing it is confused with
ThisNot this
Same inputsSame cached outputsA new sample from the model
DiffVisible between versionsA different notebook result
SeedNot the whole storyThe usual ML lever

Where Pixeltable fits

Pixeltable caches computed columns and versions the table, which is what makes a rerun comparable.

Questions

How does reproducibility work?
Inputs are versioned rows. Expressions are the column definitions. Caches mean a rerun does not call the model again unless an input changed.
What is reproducibility often confused with?
It is not only setting a random seed in a training script.