Multimodal data

What is a multimodal database?

A multimodal database stores structured values and typed media in one schema. An image, a video, an audio clip, or a document is a column the engine understands, not a file path hidden in a string.

Updated

What it is

A row can hold a title, a number, and a video at the same time. The catalog records that the video column is video. Bytes can live in a media store or in object storage you already run. Queries filter the structured fields and read the media the row points at.

Transforms belong on that schema too. Extracting frames, transcribing audio, or embedding a passage are columns that run when the row changes. An index on those outputs stays attached to the same row, so search does not depend on a second copy you update by hand.

How it works

Insert or update a row. The engine stores the structured values, records a typed pointer to the media, and runs the declared transforms for that row. Unchanged rows stay as they are. A later query can filter, play the media, or search an embedding that was computed from it.

  • The schema names the media type, so frame extraction or document splitting is a property of the column, not a script that hopes the path is a video.
  • Computed columns cache their outputs. A new file recomputes that row. Editing a prompt recomputes the columns that depend on it.
  • History can sit on the same table, so you can see what the row looked like before a model or a file changed.

What it is not

The phrase is easy to confuse with four other things. None of them let you insert a clip and keep the derived frames, transcripts, and embeddings consistent.

  • An academic benchmark. A “multimodal table” in a paper is often a dataset used to score models. It is not a system you insert into.
  • A path in a string column. A URL is not media. Filtering the path does not give you frames, pages, or a duration.
  • An untyped blob. If the engine does not know the value is video, you write the decoder yourself.
  • A split stack. Files in a bucket, metadata in a warehouse, embeddings in another database, and a job to keep them aligned. That is several systems, not one schema.

Where the media lives, and what the engine can do with it

Where the media lives, and what the engine can do with it
Path or blobSplit stackOne schema
Media in the rowOpaque path or untyped fileBytes in a bucket, ids in SQLA typed image, video, audio, or document column
Derived dataA script you remember to rerunA pipeline between servicesColumns that run when the row changes
SearchFilename or a separate index you loadEmbeddings copied into another storeAn index on a column of the same table
After an editRerun the scriptReconcile every storeRecompute the rows that changed

Where Pixeltable fits

Pixeltable is one multimodal database: Image, Video, Audio, and Document columns in a Python schema, with computed columns and embedding indexes on that schema. A longer essay on the table itself, as distinct from a benchmark or a file path, is What Is a Multimodal Data Table?

import pixeltable as pxt
TableModel = pxt.model_base()
class Library(TableModel, name='library'):
title: pxt.String
video: pxt.Video
notes: pxt.String | None

Questions

Is a multimodal database the same as a vector database?
No. A vector database stores embeddings and answers similarity queries. A multimodal database stores the media and the structured fields, and can keep an embedding index on those columns. The index is one part of the schema, not a second database you have to load.
Where do the video and image bytes live?
The catalog holds a typed pointer. The bytes live in the media store or in object storage. The column type is what lets later columns extract frames, transcribe audio, or split a document.
What is a multimodal data table?
A multimodal data table is the row-and-column form of this idea: structured values plus typed media, with computed columns and history on that schema. The essay is at /blog/what-is-a-multimodal-data-table.