Documents
What is a document splitter?
A document splitter cuts a document into retrieval units — sentences, paragraphs, or pages — as rows over the source file.
Updated · Part of What is chunking?
How it works
- The view iterates the document column.
- Each unit is a row with the text and a pointer to the file.
- A changed file rebuilds the units that came from it.
What it is not
It is not a one-off split in a notebook that never updates when the PDF changes.
document splitter: this, and the thing it is confused with
| This | Not this | |
|---|---|---|
| Input | A document column | A string you split once |
| Output | Passage rows | A list in memory |
| On edit | Units follow the file | You rerun the notebook |
Where Pixeltable fits
Pixeltable’s view sets iterator=document_splitter(...) on a pxt.Document column.
import pixeltable as pxtfrom pixeltable.functions.document import document_splitterTableModel = pxt.model_base()class Docs(TableModel, name='docs'):document: pxt.Documentclass Chunks(TableModel,name='chunks',base=Docs,iterator=document_splitter(document=Docs.document, separators='sentence'),):pass
Questions
- How does document splitter work?
- The view iterates the document column. Each unit is a row with the text and a pointer to the file. A changed file rebuilds the units that came from it.
- What is document splitter often confused with?
- It is not a one-off split in a notebook that never updates when the PDF changes.