Vision
What is CLIP?
CLIP is a model that maps images and text into one space, so a sentence can retrieve a matching picture or frame.
Updated
How it works
- The same model embeds a phrase and a picture.
- Nearby vectors are the matches.
- It does not, by itself, draw boxes around objects.
What it is not
It is not a captioning language model, and it is not a text-only sentence transformer.
CLIP: this, and the thing it is confused with
| This | Not this | |
|---|---|---|
| Query | Text or an image | Only keywords in a filename |
| Space | Shared by text and images | Text-only, or image-only |
| Does not do | Boxes and classes | Nearest picture |
Where Pixeltable fits
Pixeltable plugs CLIP in as the embedding function on an image or frame column.
Questions
- How does CLIP work?
- The same model embeds a phrase and a picture. Nearby vectors are the matches. It does not, by itself, draw boxes around objects.
- What is CLIP often confused with?
- It is not a captioning language model, and it is not a text-only sentence transformer.