Ollama classification with probabilities, fully local
Run typed LLM classification on your own machine with Ollama: no key, no network hop, privacy strict allowed, and a cost of $0 per call.
Ollama classification with Curva gives you typed answers from a model on your own machine: one of your labels per question, with a probability for every option, and no API key, no network hop and no per-call price. You pull a small model with `ollama pull qwen3:4b`, set `CURVA_MODEL=@ollama/qwen3:4b`, and call `curva.decide`. Because the model's URL is on your machine, Curva also accepts `privacy: "strict"`, which turns "nothing leaves the machine" from a habit into a rule. This guide covers the setup, why strict works locally and fails on a remote provider, how to test a small model before you trust it, and how to move Ollama to another port or host.
Some data shouldn't leave the building: customer messages under contract, internal tickets, anything your legal team has opinions about. Some teams just don't want a per-call bill or a rate limit. For both, the answer is a model you run yourself. Curva is free to use under the Curva Free License, and a local model has no per-call price, so the running cost is your hardware.
ollama pull qwen3:4b and CURVA_MODEL
Ollama serves an OpenAI-compatible API on `http://127.0.0.1:11434/v1`, which is the default URL of Curva's built-in `ollama` provider. Two commands set it up:
ollama pull qwen3:4b export CURVA_MODEL=@ollama/qwen3:4b
Then install the Python package, which includes the Curva server: `pip install curva-ai`. That is the whole configuration. Model ids that start with `@ollama/` go to the built-in Ollama provider, which needs no key, so you don't set `OPENROUTER_API_KEY` or any other provider key. `qwen3:4b` is the docs' example; any instruct model you have pulled works the same way.
Your first Ollama classification with privacy strict
import curva
d = curva.decide("The app crashes on launch", {"team": ["billing", "technical"]}, privacy="strict")`curva.decide` starts a private Curva server for this Python process on a free localhost port, reuses it for later calls, and stops it when Python exits. The model call goes to Ollama on localhost. The answer is typed: `d.team` is `billing` or `technical`, and `d["team"]` holds the confidence and a probability for each option. `d["team"]` can also be `none_of_these`, the escape option every Choice gets, when the text fits neither team.
The local database lives at `~/.curva/curva.db`, so calibration learned in one run is there in the next. On disk, the audit log keeps only a salted hash of each state, never the state itself.
Why strict works locally and fails on a remote provider
`privacy: "strict"` routes the model call only to providers that neither store nor train on prompts. How Curva enforces it depends on the provider:
That last line is the point. If someone later points the model at a hosted API by mistake, a strict request fails loudly instead of quietly sending data out. Strict also returns extracted values (Text, Number, Integer) without storing them; the audit log shows them as `redacted`.
- 1curva.decide(..., privacy=strict)
- 2provider URL on this machine?
- 3call Ollama locally, typed answer
- 4422, nothing sent
Strict passes for a provider URL on this machine and is refused with 422 for a remote one.
Local models also change the bookkeeping. `cost_usd` is always 0 on your own machine. If you want decisions to carry your GPU cost, give a named provider a price with `CURVA_PROVIDER_<NAME>_PRICE` (input and output dollars per million tokens). And a provider URL on this machine has no rate limit by default.
Speed and accuracy: test a small local model first
A 4B model is not a frontier model. Measure it on your data before you route anything on it.
**Check the probabilities.** Curva reads probabilities in `logprobs` mode (one label token per question, probabilities from the token log-probabilities) or `verbal` mode (JSON limited to your labels, with a probability for each). The default, `auto`, tries logprobs first, falls back to verbal, and remembers the result per model id. Whether a local model returns logprobs varies, so check it:
curva spike @ollama/qwen3:4b
`curva spike` reports whether a model gives usable label probabilities. Either mode works. The response's `mode` field tells you which one answered.
**Measure accuracy.** Run `curva bench` on a labeled set with `--model @ollama/qwen3:4b` to see accuracy, calibration error, latency and cost, or replay your own logged traffic with `curva shadow`.
**Help the small model.** A few settings matter more locally:
from curva import Noul, Choice
urgent = Noul("Is this urgent?", examples=[
({"ticket": "The production server is down"}, True),
({"ticket": "Typo on the pricing page"}, False),
])
team = Choice("Which team?", ["billing", "technical"], examples=[
({"ticket": "Refund my last invoice"}, "billing"),
])Moving Ollama to another port or host
If Ollama listens somewhere else, move the built-in provider with one variable:
export CURVA_PROVIDER_OLLAMA_URL=http://127.0.0.1:9999/v1 # moves the built-in @ollama
The same works for a GPU box on your network: `CURVA_PROVIDER_GPU_URL=http://gpu-box:8000/v1` adds a provider named `gpu`, and model ids become `@gpu/<model>`. Remember that `privacy: strict` passes only when that URL is on this machine, so a model on another host is a remote provider as far as strict is concerned.
For more than one script, run one Curva server and point clients at it:
curva serve --model @ollama/qwen3:4b --db curva.db
Without API keys, `curva serve` listens only on localhost. To serve other machines, create a key with `curva keys create`, restart the server, and put a reverse proxy with TLS in front.
The same pattern works for the other local servers Curva knows:
When to add a hosted model
Some teams keep everything local. Others want a stronger model for the hard minority. A cascade asks the local model first and sends only the questions it is unsure about, below `escalate_below` (default 0.8), to the next model. A council can mix a local and a hosted model too. Be clear about what that means: those states leave the machine, and a strict request will refuse a named hosted provider. Decide per project which data may go out.
Next steps
The docs cover this in [model providers](https://itsmohitrohilla.github.io/curva-docs/guides/providers/) and [getting started](https://itsmohitrohilla.github.io/curva-docs/getting-started/). To compare a local model with hosted ones, read [choose a model for LLM classification](/blog/choose-a-model-for-llm-classification/). For the free hosted route, see [free LLM APIs for classification](/blog/free-llm-api-classification/), and to check logprobs support first, [check LLM logprobs support](/blog/check-llm-logprobs-support/).