Documents
What is PDF RAG?
PDF RAG keeps the file, splits it into passages or pages, embeds those units, and generates an answer that can cite them.
Updated · Part of What is RAG?
How it works
- Store the PDF as a document.
- Split it into passages that point back at the file.
- Embed the passages and answer from the nearest ones.
What it is not
It is not uploading a PDF into a chat window with no index, and it is not OCR with nowhere to retrieve from.
PDF RAG: this, and the thing it is confused with
| This | Not this | |
|---|---|---|
| Source | The PDF row | A chat attachment that disappears |
| Units | Passages or pages | One vector for the whole file |
| Answer | Cites a passage | Paraphrases with no pointer |
Where Pixeltable fits
Pixeltable stores the document, splits it with a view, and indexes the passage text. The model call reads those hits.
Questions
- How does PDF RAG work?
- Store the PDF as a document. Split it into passages that point back at the file. Embed the passages and answer from the nearest ones.
- What is PDF RAG often confused with?
- It is not uploading a PDF into a chat window with no index, and it is not OCR with nowhere to retrieve from.