Audio
What is Whisper?
Whisper is a speech-recognition model family used to transcribe audio and video soundtracks into text.
Updated · Part of What is speech-to-text?
How it works
- You pass audio and a model name.
- The result is transcript text, sometimes with segments.
- A different model embeds or answers from that text.
What it is not
It is not a text embedding model, and it is not speaker diarization by itself.
Whisper: this, and the thing it is confused with
| This | Not this | |
|---|---|---|
| Job | Transcription | Embedding or chat |
| Input | Audio | A sentence |
| Speakers | Not identified unless you add that | A diarization model |
Where Pixeltable fits
Pixeltable calls Whisper through a transcription function on an audio column, for example openai.transcriptions.
Questions
- How does Whisper work?
- You pass audio and a model name. The result is transcript text, sometimes with segments. A different model embeds or answers from that text.
- What is Whisper often confused with?
- It is not a text embedding model, and it is not speaker diarization by itself.