Orchestration
What is local inference?
Local inference runs a model on hardware you control, instead of calling a hosted model API.
Updated · Part of What is a model provider?
How it works
- A runtime such as Ollama, llama.cpp, or vLLM serves the model.
- A computed column calls that local endpoint.
- The row still caches the output.
What it is not
It is not free in electricity or GPUs, and it is not required for the open-source database engine.
local inference: this, and the thing it is confused with
| This | Not this | |
|---|---|---|
| Bill | Your machine | Per token to a vendor |
| Weights | On disk | Behind an API |
| Schema | Unchanged aside from the function | A different pipeline |
Where Pixeltable fits
Pixeltable columns can call Ollama, llama.cpp, or vLLM the same way they call a hosted provider.
Questions
- How does local inference work?
- A runtime such as Ollama, llama.cpp, or vLLM serves the model. A computed column calls that local endpoint. The row still caches the output.
- What is local inference often confused with?
- It is not free in electricity or GPUs, and it is not required for the open-source database engine.