One signed binary. Every feature compiled in. Free to run. Install Crowkis →
← back to the Roost
curva guidesOctober 3, 2026· 5 min read

curva.local() vs Curva(): running the server from Python

Three ways to get a Curva client in Python: module-level decide, curva.local() with its own server, or Curva() for a shared one. What each does.

Curva runs a local LLM classification server from Python in three ways. The module-level `curva.decide` starts a private server on its first call and stops it when Python exits. `curva.local()` does the same but hands you the lifecycle: a client bound to a server on a free `127.0.0.1` port, its own database file, and a `with` block to stop it. `Curva()`, given a base URL, starts nothing and talks to a shared server you run with `curva serve` or Docker. All three run the same `curva` binary, which ships inside the Python wheel. This guide shows what each one does with your environment, your database and your keys, and when to move from one to the next.

A local LLM classification server from Python: three entry points

Three ways to run a local LLM classification server from Python
  1. 1
    curva.decide(...)
  2. 2
    private server, started on first call
  3. 3
    curva.local()
  4. 4
    private server on a free 127.0.0.1 port
  5. 5
    Curva(base_url)
  6. 6
    shared server: curva serve or Docker
  7. 7
    ~/.curva/curva.db or db=
  8. 8
    the server's own --db file

The first two start a private server inside your Python process's lifetime; the third connects to one you run yourself.

curva.decide: a private server on the first call, stopped when Python exits

python
import curva
d = curva.decide("I was charged twice, please refund me",
                 {"team": ["billing", "technical"], "refund": "Asks for a refund?", "total": float})
print(d.team, d.refund, d.total)           # billing True None

There is nothing to set up. The first `curva.decide` starts a private server for this Python process, reuses it for every later call, and stops it when Python exits. `curva.feedback` and `curva.calibration` at module level use the same server.

One switch changes that: with `CURVA_BASE_URL` set, the module-level functions talk to that server instead of starting their own. So the same script can run against a private server on a laptop and a shared one in production, with only an environment variable between them.

curva.local(): a free 127.0.0.1 port, startup_timeout=15.0 and the with block

python
curva.local(model=None, db=None, *, startup_timeout=15.0, **client_options)

`curva.local()` starts `curva serve` on a free `127.0.0.1` port and returns a client bound to it. Because it listens only on localhost, nothing outside your machine can reach it, and it needs no API key.

The server stops when Python exits, when you call `close()`, or at the end of a `with` block:

python
with curva.local() as client:
    ...

The full first decision from the docs uses it this way:

python
import curva
from curva import Choice, Score, Noul

client = curva.local()
d = client.decide(
    state={"ticket": "I was charged twice for order A-104. Please refund the duplicate!"},
    questions={
        "team": Choice("Which team should handle this?",
                       {"billing": "payments, refunds", "technical": "bugs", "sales": "pricing"}),
        "frustration": Score("How frustrated is the customer?", ["calm", "annoyed", "angry"]),
        "refund": Noul("The customer explicitly asks for a refund"),
    },
)

print(d["team"].choice, d["team"].confidence)   # billing 0.9999
print(d["frustration"].score)                   # 0.65
print(d["refund"].noul)                         # 0.999
print(d.mode, d.latency_ms, d.cost_usd, d.cached)

Where calibration lives: ~/.curva/curva.db and the db= argument

Calibration is only useful if it survives. `curva.local()` keeps its database at `~/.curva/curva.db` by default, so calibration learned in one run is there in the next. Feedback you send today fits a calibrator that tomorrow's script uses.

Pass `db=` to keep separate databases: one per experiment, a throwaway file in tests, or a fixed path in a container. Calibrators are also kept per `project` inside one database, so two use cases can share a file without mixing labels.

The server inherits your environment: OPENROUTER_API_KEY, CURVA_MODEL, CURVA_RPM

The private server is a child of your Python process, and it inherits the process's environment. So set provider variables before you start it:

bash
export OPENROUTER_API_KEY=sk-or-v1-...     # any OpenRouter key; free models work

Without `OPENROUTER_API_KEY`, the first key set among `OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, `GEMINI_API_KEY`, `GROQ_API_KEY` and a few others picks a small, fast model of that provider. `CURVA_MODEL` chooses one yourself, for example `CURVA_MODEL=@ollama/qwen3:4b` for a fully local setup. Server settings such as `CURVA_RPM` pass through the same way. With no key at all, `curva.decide` raises an error naming the variables to set.

Moving to a shared server with CURVA_BASE_URL and CURVA_API_KEY

A private server per process is right for a script. For a service, it means one database, one decision cache and one set of calibrators per process, and none of them shared. Move to one shared server:

bash
curva serve --addr 127.0.0.1:7777 --db curva.db
curva keys create --name my-service
python
from curva import Curva

client = Curva("http://your-server:7777")   # or set CURVA_BASE_URL

`Curva()` reads `CURVA_BASE_URL`, default `http://localhost:7777`, and `CURVA_API_KEY` when they aren't passed. Once the server has API keys, every request needs one. A shared server keeps one audit log, one set of calibrators per project, and one decision cache, so a repeat from any client comes back in about a millisecond. For production, the Docker image runs the same binary with its data in a volume.

Three mistakes to avoid

**A private server per request.** `curva.local()` starts a whole server. Calling it inside a web request handler starts one per request, each with its own cache and its own view of the database. Create one client when your process starts and reuse it, or point every process at one shared server with `Curva()`.

**Tests that write to your real calibration.** By default `curva.local()` writes to `~/.curva/curva.db`. A test suite that sends feedback would then fit calibrators your scripts later use. Give tests their own file with `db=`, and their own `project`.

**Several processes, one file, no server.** Two scripts that each start `curva.local()` on the same database file each run their own server against it. If several processes need the same calibrators and cache, that is the moment to run one `curva serve` and connect with `Curva()`. The server has one writer, one cache and one audit log, which is what a team needs.

Which wheels ship the curva binary: Linux, macOS and Windows

`curva.local()` and module-level `curva.decide` need the `curva` binary, and `pip install curva-ai` provides it. Platform wheels for Linux (x86_64, aarch64), macOS (Apple Silicon, Intel) and Windows include the server binary, so you don't need Rust. The SDK itself uses only the standard library, on Python 3.9 or newer. `curva --version` should print `curva 0.1.0`.

Next steps

The docs cover the clients in the [Python SDK reference](https://itsmohitrohilla.github.io/curva-docs/reference/python-sdk/) and [getting started](https://itsmohitrohilla.github.io/curva-docs/getting-started/). Start with the [Python LLM classification tutorial](/blog/python-llm-classification/), learn the shorthand in [Python LLM shorthand questions](/blog/python-llm-shorthand-questions/), and tune the client in [LLM client retries and timeouts in Python](/blog/llm-client-retries-timeouts-python/). When you outgrow a private server, read [self-host an LLM decision API](/blog/self-host-llm-decision-api/).