Audio

What is Whisper?

Whisper is a speech-recognition model family used to transcribe audio and video soundtracks into text.

Updated · Part of What is speech-to-text?

How it works

  • You pass audio and a model name.
  • The result is transcript text, sometimes with segments.
  • A different model embeds or answers from that text.

What it is not

It is not a text embedding model, and it is not speaker diarization by itself.

Whisper: this, and the thing it is confused with

Whisper: this, and the thing it is confused with
ThisNot this
JobTranscriptionEmbedding or chat
InputAudioA sentence
SpeakersNot identified unless you add thatA diarization model

Where Pixeltable fits

Pixeltable calls Whisper through a transcription function on an audio column, for example openai.transcriptions.

Questions

How does Whisper work?
You pass audio and a model name. The result is transcript text, sometimes with segments. A different model embeds or answers from that text.
What is Whisper often confused with?
It is not a text embedding model, and it is not speaker diarization by itself.