One signed binary. Every feature compiled in. Free to run. Install Crowkis →
← back to the Roost
curva guidesOctober 3, 2026· 5 min read

Python LLM classification with confidence and feedback

A start-to-finish Python tutorial: typed LLM labels with a probability each, abstain, feedback, calibration reports, async batches and errors.

To do Python LLM classification with a confidence you can act on, install `curva-ai`, declare your labels as typed questions, and call `decide`. You get one of your labels back with a probability for every option, never free text to parse. Send the true answer when you learn it, and after 30 labels Curva calibrates those probabilities on your data. This tutorial covers the whole loop: install, decide, abstain, feedback, the calibration report, async batches and error handling.

Most Python classification code built on an LLM looks the same today: a prompt that says "reply with one of: billing, technical, sales", a call, and a parser that hopes the reply is one of those words. Then someone adds "and how confident are you?", and the model writes a high number on almost everything, right or wrong. The steps below replace that pattern.

Install curva-ai

bash
pip install curva-ai
export OPENROUTER_API_KEY=sk-or-v1-...     # any OpenRouter key; free models work

The package includes the Curva server binary for Linux, macOS and Windows, so there is nothing else to install. It needs Python 3.9 or newer and uses only the standard library. Other providers work too: without `OPENROUTER_API_KEY`, the first of `OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, `GEMINI_API_KEY`, `GROQ_API_KEY` and a few more picks a small, fast model of that provider. Set `CURVA_MODEL` to choose one yourself. Curva is free to use; you pay only your model provider, and free models work.

Python LLM classification in three lines

python
import curva
d = curva.decide("I was charged twice, please refund me",
                 {"team": ["billing", "technical"], "refund": "Asks for a refund?", "total": float})
print(d.team, d.refund, d.total)           # billing True None

The first `curva.decide` starts a private server for this process and reuses it. The shorthand maps Python values to question types:

`d.team` is the plain answer. `d["team"]` is the full one, with `confidence` and `probabilities`. `d.to_dict()` gives all plain answers as a dict. For a yes/no question, the plain answer is `True` when P(yes) is at least 0.5; read `d["refund"].noul` when you want the probability itself. The shorthand is good for scripts and notebooks. For anything you will run in production, switch to explicit question objects, because they let you set thresholds, examples and coverage per question.

The full client with typed questions

For real code, use explicit question objects and a client:

python
import curva
from curva import Choice, Score, Noul

client = curva.local()      # or curva.Curva("http://your-server:7777")

QUESTIONS = {
    "team": Choice("Which team should handle this?",
                   {"billing": "payments, refunds", "technical": "bugs", "sales": "pricing"},
                   min_confidence=0.8),
    "frustration": Score("How frustrated is the customer?", ["calm", "annoyed", "angry"]),
    "refund": Noul("The customer explicitly asks for a refund"),
}

d = client.decide({"ticket": "I was charged twice for order A-104. Please refund the duplicate!"},
                  QUESTIONS, project="support")

print(d["team"].choice, d["team"].confidence)   # e.g. billing 0.9999
print(d["frustration"].score)                   # e.g. 0.65
print(d["refund"].noul)                         # e.g. 0.999
print(d.mode, d.latency_ms, d.cost_usd, d.cached)

A few details worth knowing:

Abstain: send unsure answers to a person

Because the Choice set `min_confidence=0.8`, answers below 0.8 come back with `abstain=True`:

python
answer = d["team"]
if answer.abstain:
    team = ask_a_human(ticket)
    client.feedback(d.id, "team", team)      # the person's answer teaches Curva
else:
    route(ticket, answer.choice)

`abstain` is only set on questions that have `min_confidence`. Until the question is calibrated, 0.8 is a guess about the model. Once calibrated, it means what it says. The thinking behind the threshold is in [LLM classification confidence scores you can act on](/blog/llm-classification-confidence-scores/).

Feedback and the calibration report

Keep `d.id` next to whatever you did with the answer. When you learn the truth (an agent closes the ticket, a reviewer fixes a label), send it:

python
client.feedback(decision_id, "team", "billing")     # Choice: the option key
client.feedback(decision_id, "frustration", 2)      # Score: the level index
client.feedback(decision_id, "refund", True)        # Noul: True or False

Sending feedback again for the same decision and question replaces the earlier label. After 30 labels for a question in a project, Curva fits a calibrator. It keeps it only when it clearly beats the raw probabilities on held-out labels in both log-loss and Brier score; then answers come back with `calibrated=True`. Otherwise the raw probabilities stay, because the model is already well calibrated. Check what happened:

python
report = client.calibration("team", project="support")
print(report["before"]["ece"], "→", report["after"]["ece"])
print(report["after"]["accuracy_when_automated"], report["after"]["automated"])

`after` is held out: each half of the labels is calibrated by a fit on the other half, so the number isn't flattered by testing on the data it was fitted to. Also send feedback for a sample of confident answers, not only the abstained ones, so calibration covers the whole range.

One rule to remember: calibration belongs to the exact wording of a question. Settle the wording before you collect labels.

Many records at once with AsyncCurva

For a batch, run the async client against a server:

python
import asyncio
from curva import AsyncCurva

async def classify(rows):
    async with AsyncCurva() as client:          # CURVA_BASE_URL, CURVA_API_KEY
        results = await client.decide_many(
            [({"ticket": r["text"]}, QUESTIONS) for r in rows],
            concurrency=16, project="support", return_exceptions=True,
        )
    return [
        {"id": r["id"], "error": str(d)} if isinstance(d, Exception)
        else {"id": r["id"], "decision_id": d.id, "team": d["team"].choice,
              "needs_human": bool(d["team"].abstain)}
        for r, d in zip(rows, results)
    ]

Results keep the input order. `return_exceptions=True` puts each failure in its slot, so one bad row doesn't fail the batch. `concurrency` defaults to 8. For files on disk, `curva map` on the command line is the other route: it resumes after a stop.

Error handling with CurvaError

The client retries 429 and 5xx responses (three times by default) and honours `Retry-After`. Anything it doesn't retry raises a `CurvaError` with `status`, `type` and `message`. Catch a subclass when you care about one case:

python
from curva import CurvaError, RateLimitError, InvalidRequestError

try:
    d = client.decide(state, QUESTIONS)
except RateLimitError as e:
    wait(e.retry_after)
except InvalidRequestError as e:
    log.error("bad question: %s", e.message)    # a 422 names the question
except CurvaError as e:
    log.error("%s %s", e.status, e.type)

Status 0 means the server wasn't reachable.

Move to a shared server

`curva.local()` is right for scripts and notebooks. For a service, run one server and point clients at it:

bash
curva serve --addr 127.0.0.1:7777 --db curva.db
curva keys create --name my-service
python
from curva import Curva
client = Curva()        # reads CURVA_BASE_URL and CURVA_API_KEY

A shared server keeps one audit log, one set of calibrators per project, and one decision cache. Identical requests come back from the cache in about a millisecond, at no cost.

What not to ask the model

Don't ask a model to count, add up or compare dates. Do that in Python and put the result in the state. Ask Curva the judgement questions: which team, how urgent, is this a refund request.

Next steps

Read the [Python SDK reference](https://itsmohitrohilla.github.io/curva-docs/reference/python-sdk/) and the [feedback guide](https://itsmohitrohilla.github.io/curva-docs/guides/feedback/) in the docs. To see why Curva asks every question in two orders, read [LLM position bias](/blog/llm-position-bias/). For a full worked use case, try [phishing detection with an LLM](/blog/phishing-detection-llm/).