Retrieval
What is chunking?
Chunking splits a source into retrieval units — sentences, pages, token windows — small enough to embed and cite.
Updated
How it works
- Choose a unit that fits the question you expect.
- Each unit becomes a row that points at the source.
- The embedding is computed on the unit, not on the entire file.
What it is not
It is not embedding the whole file as one vector, and it is not throwing away the source after the split.
chunking: this, and the thing it is confused with
| This | Not this | |
|---|---|---|
| Unit | A passage you can quote | The whole PDF as one vector |
| Parent | Kept | Discarded after the split |
| Too-large units | Miss the local fact | A window that still cites a page |
Where Pixeltable fits
Pixeltable chunks with an iterator view, so each passage row still points at the document or string it came from.
Questions
- How does chunking work?
- Choose a unit that fits the question you expect. Each unit becomes a row that points at the source. The embedding is computed on the unit, not on the entire file.
- What is chunking often confused with?
- It is not embedding the whole file as one vector, and it is not throwing away the source after the split.