Orchestration

What is local inference?

Local inference runs a model on hardware you control, instead of calling a hosted model API.

Updated · Part of What is a model provider?

How it works

  • A runtime such as Ollama, llama.cpp, or vLLM serves the model.
  • A computed column calls that local endpoint.
  • The row still caches the output.

What it is not

It is not free in electricity or GPUs, and it is not required for the open-source database engine.

local inference: this, and the thing it is confused with

local inference: this, and the thing it is confused with
ThisNot this
BillYour machinePer token to a vendor
WeightsOn diskBehind an API
SchemaUnchanged aside from the functionA different pipeline

Where Pixeltable fits

Pixeltable columns can call Ollama, llama.cpp, or vLLM the same way they call a hosted provider.

Questions

How does local inference work?
A runtime such as Ollama, llama.cpp, or vLLM serves the model. A computed column calls that local endpoint. The row still caches the output.
What is local inference often confused with?
It is not free in electricity or GPUs, and it is not required for the open-source database engine.