Retrieval
What is cross-modal search?
Cross-modal search queries one modality with another — a sentence against frames, a still against a video library — because they share an embedding space.
Updated · Part of What is semantic search?
How it works
- One model places text and images in the same space.
- The query is embedded as text or as an image.
- The nearest frames or pictures are the hits, with their source rows.
What it is not
It is not transcript-only search, and it is not filename search.
cross-modal search: this, and the thing it is confused with
| This | Not this | |
|---|---|---|
| Query | Text or a still | Only the same file type, by name |
| Index | Frames or images in that shared space | A transcript |
| Hit | A picture that matches the phrase | A document that contains the words |
Where Pixeltable fits
Pixeltable uses a shared embedding such as CLIP on a frame or image column, then similarity with a string or an image.
Questions
- How does cross-modal search work?
- One model places text and images in the same space. The query is embedded as text or as an image. The nearest frames or pictures are the hits, with their source rows.
- What is cross-modal search often confused with?
- It is not transcript-only search, and it is not filename search.