Python LLM classification with confidence and feedback
A start-to-finish Python tutorial: typed LLM labels with a probability each, abstain, feedback, calibration reports, async batches and errors.
To do Python LLM classification with a confidence you can act on, install `curva-ai`, declare your labels as typed questions, and call `decide`. You get one of your labels back with a probability for every option, never free text to parse. Send the true answer when you learn it, and after 30 labels Curva calibrates those probabilities on your data. This tutorial covers the whole loop: install, decide, abstain, feedback, the calibration report, async batches and error handling.
Most Python classification code built on an LLM looks the same today: a prompt that says "reply with one of: billing, technical, sales", a call, and a parser that hopes the reply is one of those words. Then someone adds "and how confident are you?", and the model writes a high number on almost everything, right or wrong. The steps below replace that pattern.
Install curva-ai
pip install curva-ai export OPENROUTER_API_KEY=sk-or-v1-... # any OpenRouter key; free models work
The package includes the Curva server binary for Linux, macOS and Windows, so there is nothing else to install. It needs Python 3.9 or newer and uses only the standard library. Other providers work too: without `OPENROUTER_API_KEY`, the first of `OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, `GEMINI_API_KEY`, `GROQ_API_KEY` and a few more picks a small, fast model of that provider. Set `CURVA_MODEL` to choose one yourself. Curva is free to use; you pay only your model provider, and free models work.
Python LLM classification in three lines
import curva
d = curva.decide("I was charged twice, please refund me",
{"team": ["billing", "technical"], "refund": "Asks for a refund?", "total": float})
print(d.team, d.refund, d.total) # billing True NoneThe first `curva.decide` starts a private server for this process and reuses it. The shorthand maps Python values to question types:
`d.team` is the plain answer. `d["team"]` is the full one, with `confidence` and `probabilities`. `d.to_dict()` gives all plain answers as a dict. For a yes/no question, the plain answer is `True` when P(yes) is at least 0.5; read `d["refund"].noul` when you want the probability itself. The shorthand is good for scripts and notebooks. For anything you will run in production, switch to explicit question objects, because they let you set thresholds, examples and coverage per question.
The full client with typed questions
For real code, use explicit question objects and a client:
import curva
from curva import Choice, Score, Noul
client = curva.local() # or curva.Curva("http://your-server:7777")
QUESTIONS = {
"team": Choice("Which team should handle this?",
{"billing": "payments, refunds", "technical": "bugs", "sales": "pricing"},
min_confidence=0.8),
"frustration": Score("How frustrated is the customer?", ["calm", "annoyed", "angry"]),
"refund": Noul("The customer explicitly asks for a refund"),
}
d = client.decide({"ticket": "I was charged twice for order A-104. Please refund the duplicate!"},
QUESTIONS, project="support")
print(d["team"].choice, d["team"].confidence) # e.g. billing 0.9999
print(d["frustration"].score) # e.g. 0.65
print(d["refund"].noul) # e.g. 0.999
print(d.mode, d.latency_ms, d.cost_usd, d.cached)A few details worth knowing:
Abstain: send unsure answers to a person
Because the Choice set `min_confidence=0.8`, answers below 0.8 come back with `abstain=True`:
answer = d["team"]
if answer.abstain:
team = ask_a_human(ticket)
client.feedback(d.id, "team", team) # the person's answer teaches Curva
else:
route(ticket, answer.choice)`abstain` is only set on questions that have `min_confidence`. Until the question is calibrated, 0.8 is a guess about the model. Once calibrated, it means what it says. The thinking behind the threshold is in [LLM classification confidence scores you can act on](/blog/llm-classification-confidence-scores/).
Feedback and the calibration report
Keep `d.id` next to whatever you did with the answer. When you learn the truth (an agent closes the ticket, a reviewer fixes a label), send it:
client.feedback(decision_id, "team", "billing") # Choice: the option key client.feedback(decision_id, "frustration", 2) # Score: the level index client.feedback(decision_id, "refund", True) # Noul: True or False
Sending feedback again for the same decision and question replaces the earlier label. After 30 labels for a question in a project, Curva fits a calibrator. It keeps it only when it clearly beats the raw probabilities on held-out labels in both log-loss and Brier score; then answers come back with `calibrated=True`. Otherwise the raw probabilities stay, because the model is already well calibrated. Check what happened:
report = client.calibration("team", project="support")
print(report["before"]["ece"], "→", report["after"]["ece"])
print(report["after"]["accuracy_when_automated"], report["after"]["automated"])`after` is held out: each half of the labels is calibrated by a fit on the other half, so the number isn't flattered by testing on the data it was fitted to. Also send feedback for a sample of confident answers, not only the abstained ones, so calibration covers the whole range.
One rule to remember: calibration belongs to the exact wording of a question. Settle the wording before you collect labels.
Many records at once with AsyncCurva
For a batch, run the async client against a server:
import asyncio
from curva import AsyncCurva
async def classify(rows):
async with AsyncCurva() as client: # CURVA_BASE_URL, CURVA_API_KEY
results = await client.decide_many(
[({"ticket": r["text"]}, QUESTIONS) for r in rows],
concurrency=16, project="support", return_exceptions=True,
)
return [
{"id": r["id"], "error": str(d)} if isinstance(d, Exception)
else {"id": r["id"], "decision_id": d.id, "team": d["team"].choice,
"needs_human": bool(d["team"].abstain)}
for r, d in zip(rows, results)
]Results keep the input order. `return_exceptions=True` puts each failure in its slot, so one bad row doesn't fail the batch. `concurrency` defaults to 8. For files on disk, `curva map` on the command line is the other route: it resumes after a stop.
Error handling with CurvaError
The client retries 429 and 5xx responses (three times by default) and honours `Retry-After`. Anything it doesn't retry raises a `CurvaError` with `status`, `type` and `message`. Catch a subclass when you care about one case:
from curva import CurvaError, RateLimitError, InvalidRequestError
try:
d = client.decide(state, QUESTIONS)
except RateLimitError as e:
wait(e.retry_after)
except InvalidRequestError as e:
log.error("bad question: %s", e.message) # a 422 names the question
except CurvaError as e:
log.error("%s %s", e.status, e.type)Status 0 means the server wasn't reachable.
Move to a shared server
`curva.local()` is right for scripts and notebooks. For a service, run one server and point clients at it:
curva serve --addr 127.0.0.1:7777 --db curva.db curva keys create --name my-service
from curva import Curva client = Curva() # reads CURVA_BASE_URL and CURVA_API_KEY
A shared server keeps one audit log, one set of calibrators per project, and one decision cache. Identical requests come back from the cache in about a millisecond, at no cost.
What not to ask the model
Don't ask a model to count, add up or compare dates. Do that in Python and put the result in the state. Ask Curva the judgement questions: which team, how urgent, is this a refund request.
Next steps
Read the [Python SDK reference](https://itsmohitrohilla.github.io/curva-docs/reference/python-sdk/) and the [feedback guide](https://itsmohitrohilla.github.io/curva-docs/guides/feedback/) in the docs. To see why Curva asks every question in two orders, read [LLM position bias](/blog/llm-position-bias/). For a full worked use case, try [phishing detection with an LLM](/blog/phishing-detection-llm/).