Retrieval
What is multimodal RAG?
Multimodal RAG retrieves pages, frames, or transcripts — not only plain-text passages — and passes those pieces to the model.
Updated · Part of What is RAG?
How it works
- Each modality is stored as a typed value.
- Retrieval units are passages, frames, or transcript windows.
- The model sees the retrieved piece and a pointer to its source.
What it is not
It is not text-only RAG with images stored as unused attachments.
multimodal RAG: this, and the thing it is confused with
| This | Not this | |
|---|---|---|
| Evidence | Page, frame, or transcript | Text chunks only |
| Unused media | A failure of the pipeline | Normal, if images are just attachments |
| Index | One per modality you intend to search | A single text index over filenames |
Where Pixeltable fits
Pixeltable keeps each modality on the table and retrieves from the view that holds the pieces. The model call is another column or a query.
Questions
- How does multimodal RAG work?
- Each modality is stored as a typed value. Retrieval units are passages, frames, or transcript windows. The model sees the retrieved piece and a pointer to its source.
- What is multimodal RAG often confused with?
- It is not text-only RAG with images stored as unused attachments.