One signed binary. Every feature compiled in. Free to run. Install Crowkis →
← back to the Roost
curva guidesOctober 3, 2026· 5 min read

Ollama classification with probabilities, fully local

Run typed LLM classification on your own machine with Ollama: no key, no network hop, privacy strict allowed, and a cost of $0 per call.

Ollama classification with Curva gives you typed answers from a model on your own machine: one of your labels per question, with a probability for every option, and no API key, no network hop and no per-call price. You pull a small model with `ollama pull qwen3:4b`, set `CURVA_MODEL=@ollama/qwen3:4b`, and call `curva.decide`. Because the model's URL is on your machine, Curva also accepts `privacy: "strict"`, which turns "nothing leaves the machine" from a habit into a rule. This guide covers the setup, why strict works locally and fails on a remote provider, how to test a small model before you trust it, and how to move Ollama to another port or host.

Some data shouldn't leave the building: customer messages under contract, internal tickets, anything your legal team has opinions about. Some teams just don't want a per-call bill or a rate limit. For both, the answer is a model you run yourself. Curva is free to use under the Curva Free License, and a local model has no per-call price, so the running cost is your hardware.

ollama pull qwen3:4b and CURVA_MODEL

Ollama serves an OpenAI-compatible API on `http://127.0.0.1:11434/v1`, which is the default URL of Curva's built-in `ollama` provider. Two commands set it up:

bash
ollama pull qwen3:4b
export CURVA_MODEL=@ollama/qwen3:4b

Then install the Python package, which includes the Curva server: `pip install curva-ai`. That is the whole configuration. Model ids that start with `@ollama/` go to the built-in Ollama provider, which needs no key, so you don't set `OPENROUTER_API_KEY` or any other provider key. `qwen3:4b` is the docs' example; any instruct model you have pulled works the same way.

Your first Ollama classification with privacy strict

python
import curva
d = curva.decide("The app crashes on launch", {"team": ["billing", "technical"]}, privacy="strict")

`curva.decide` starts a private Curva server for this Python process on a free localhost port, reuses it for later calls, and stops it when Python exits. The model call goes to Ollama on localhost. The answer is typed: `d.team` is `billing` or `technical`, and `d["team"]` holds the confidence and a probability for each option. `d["team"]` can also be `none_of_these`, the escape option every Choice gets, when the text fits neither team.

The local database lives at `~/.curva/curva.db`, so calibration learned in one run is there in the next. On disk, the audit log keeps only a salted hash of each state, never the state itself.

Why strict works locally and fails on a remote provider

`privacy: "strict"` routes the model call only to providers that neither store nor train on prompts. How Curva enforces it depends on the provider:

That last line is the point. If someone later points the model at a hosted API by mistake, a strict request fails loudly instead of quietly sending data out. Strict also returns extracted values (Text, Number, Integer) without storing them; the audit log shows them as `redacted`.

Ollama classification with privacy strict
  1. 1
    curva.decide(..., privacy=strict)
  2. 2
    provider URL on this machine?
  3. 3
    call Ollama locally, typed answer
  4. 4
    422, nothing sent

Strict passes for a provider URL on this machine and is refused with 422 for a remote one.

Local models also change the bookkeeping. `cost_usd` is always 0 on your own machine. If you want decisions to carry your GPU cost, give a named provider a price with `CURVA_PROVIDER_<NAME>_PRICE` (input and output dollars per million tokens). And a provider URL on this machine has no rate limit by default.

Speed and accuracy: test a small local model first

A 4B model is not a frontier model. Measure it on your data before you route anything on it.

**Check the probabilities.** Curva reads probabilities in `logprobs` mode (one label token per question, probabilities from the token log-probabilities) or `verbal` mode (JSON limited to your labels, with a probability for each). The default, `auto`, tries logprobs first, falls back to verbal, and remembers the result per model id. Whether a local model returns logprobs varies, so check it:

bash
curva spike @ollama/qwen3:4b

`curva spike` reports whether a model gives usable label probabilities. Either mode works. The response's `mode` field tells you which one answered.

**Measure accuracy.** Run `curva bench` on a labeled set with `--model @ollama/qwen3:4b` to see accuracy, calibration error, latency and cost, or replay your own logged traffic with `curva shadow`.

**Help the small model.** A few settings matter more locally:

python
from curva import Noul, Choice

urgent = Noul("Is this urgent?", examples=[
    ({"ticket": "The production server is down"}, True),
    ({"ticket": "Typo on the pricing page"}, False),
])

team = Choice("Which team?", ["billing", "technical"], examples=[
    ({"ticket": "Refund my last invoice"}, "billing"),
])

Moving Ollama to another port or host

If Ollama listens somewhere else, move the built-in provider with one variable:

bash
export CURVA_PROVIDER_OLLAMA_URL=http://127.0.0.1:9999/v1 # moves the built-in @ollama

The same works for a GPU box on your network: `CURVA_PROVIDER_GPU_URL=http://gpu-box:8000/v1` adds a provider named `gpu`, and model ids become `@gpu/<model>`. Remember that `privacy: strict` passes only when that URL is on this machine, so a model on another host is a remote provider as far as strict is concerned.

For more than one script, run one Curva server and point clients at it:

bash
curva serve --model @ollama/qwen3:4b --db curva.db

Without API keys, `curva serve` listens only on localhost. To serve other machines, create a key with `curva keys create`, restart the server, and put a reverse proxy with TLS in front.

The same pattern works for the other local servers Curva knows:

When to add a hosted model

Some teams keep everything local. Others want a stronger model for the hard minority. A cascade asks the local model first and sends only the questions it is unsure about, below `escalate_below` (default 0.8), to the next model. A council can mix a local and a hosted model too. Be clear about what that means: those states leave the machine, and a strict request will refuse a named hosted provider. Decide per project which data may go out.

Next steps

The docs cover this in [model providers](https://itsmohitrohilla.github.io/curva-docs/guides/providers/) and [getting started](https://itsmohitrohilla.github.io/curva-docs/getting-started/). To compare a local model with hosted ones, read [choose a model for LLM classification](/blog/choose-a-model-for-llm-classification/). For the free hosted route, see [free LLM APIs for classification](/blog/free-llm-api-classification/), and to check logprobs support first, [check LLM logprobs support](/blog/check-llm-logprobs-support/).