Video
How does video search work?
Video search finds a moment inside a clip, not only a filename. The usual path is to sample frames, embed each frame, and rank those frames by similarity to a text phrase or a still image.
Updated
What it is
A video file is a sequence of pictures plus a timeline. Search that returns “beach.mp4” has matched a title. Search that returns a frame at 00:01:12 has matched the picture. The second kind is what people mean by video search, video retrieval, or finding a moment.
How it works
Three steps. The frame rate you choose is the tradeoff: more frames catch shorter moments and cost more embedding work.
- Sample. A view walks the video at a fixed rate, one row per frame, and keeps the frame index and the timestamp.
- Embed. A model such as CLIP turns each frame into a vector in a space where a text phrase and a picture can be neighbors.
- Query. The question is embedded the same way. The nearest frames come back with the source video and the time offset.
What it is not
Sending the whole file to a chat model on every question is not an index. It can describe one clip. It does not rank a library, and it does not get cheaper as the library grows.
A transcript search finds words that were spoken. It misses a red car that nobody named. Visual search and speech search answer different questions. A library often wants both, as two indexes on the same clip.
What you keep at each step of video search
| Step | What happens | What you can return |
|---|---|---|
| Sample frames | One row per frame at a chosen rate | Frame index and timestamp on the source video |
| Embed frames | A model maps each frame to a vector | Neighbors of a text phrase or a still image |
| Query | Rank frames by similarity | The moment, not only the filename |
Where Pixeltable fits
In Pixeltable a frame view uses frame_iterator on a Video column. An embedding index on the frame column is what the similarity query reads. The parent video and the timestamp stay on the frame row, so the hit is a moment. The step-by-step version is the video intelligence use case.
import pixeltable as pxtfrom pixeltable.functions.huggingface import clipfrom pixeltable.functions.video import frame_iteratorTableModel = pxt.model_base()image_embed = clip.using(model_id='openai/clip-vit-base-patch32')class Videos(TableModel, name='videos'):video: pxt.Videoclass Frames(TableModel,name='frames',base=Videos,iterator=frame_iterator(video=Videos.video, fps=1),):__indexes__ = [pxt.EmbeddingIndex(frame, image_embed=image_embed),]@pxt.querydef find_moment(query: str, limit: int = 5):sim = Frames.frame.similarity(string=query)return (Frames.order_by(sim, asc=False).limit(limit).select(Frames.video, Frames.frame_idx, Frames.pos_msec, sim))
Questions
- How many frames should I index?
- One frame per second is a common start. Higher rates catch shorter actions and create more rows to embed. Lower rates are cheaper and miss brief moments. The right rate is the shortest event you still need to find.
- Can I search video with a sentence?
- Yes, if the embedding model places text and images in one space. CLIP-style models do. The sentence and the frame are both vectors, and similarity ranks frames against the sentence.
- Is a transcript enough?
- A transcript finds what was said. It does not find a visual match that nobody described. Use a transcript index for speech and a frame index for pictures.