Documents

What is a document splitter?

A document splitter cuts a document into retrieval units — sentences, paragraphs, or pages — as rows over the source file.

Updated · Part of What is chunking?

How it works

  • The view iterates the document column.
  • Each unit is a row with the text and a pointer to the file.
  • A changed file rebuilds the units that came from it.

What it is not

It is not a one-off split in a notebook that never updates when the PDF changes.

document splitter: this, and the thing it is confused with

document splitter: this, and the thing it is confused with
ThisNot this
InputA document columnA string you split once
OutputPassage rowsA list in memory
On editUnits follow the fileYou rerun the notebook

Where Pixeltable fits

Pixeltable’s view sets iterator=document_splitter(...) on a pxt.Document column.

import pixeltable as pxt
from pixeltable.functions.document import document_splitter
TableModel = pxt.model_base()
class Docs(TableModel, name='docs'):
document: pxt.Document
class Chunks(
TableModel,
name='chunks',
base=Docs,
iterator=document_splitter(document=Docs.document, separators='sentence'),
):
pass

Questions

How does document splitter work?
The view iterates the document column. Each unit is a row with the text and a pointer to the file. A changed file rebuilds the units that came from it.
What is document splitter often confused with?
It is not a one-off split in a notebook that never updates when the PDF changes.