Audio
What is text-to-speech?
Text-to-speech synthesizes spoken audio from text.
Updated
How it works
- A column holds the text.
- A speech model writes audio.
- The audio can be stored on the same row as the script.
What it is not
It is not transcription, and it is not cloning a private recording without a speech model.
text-to-speech: this, and the thing it is confused with
| This | Not this | |
|---|---|---|
| Direction | Text to audio | Audio to text |
| Output | A waveform or file | A transcript |
| Stored with | The script row | A separate media bin |
Where Pixeltable fits
Pixeltable can assign a speech computed column that writes audio from a text column.
Questions
- How does text-to-speech work?
- A column holds the text. A speech model writes audio. The audio can be stored on the same row as the script.
- What is text-to-speech often confused with?
- It is not transcription, and it is not cloning a private recording without a speech model.