One signed binary. Every feature compiled in. Free to run. Install Crowkis →
← back to the Roost
curva guidesOctober 3, 2026· 5 min read

Send feedback to an LLM classifier from Python

Send true labels to an LLM classifier from Python: store the decision id, use the right label per question type, and avoid 404 and 422 errors.

An LLM classifier feedback loop in Python has one call at its centre: `client.feedback(decision_id, question, label)`. Store the decision's `id` when you act on an answer, and when you learn the true answer, send it with the label format for that question type: an option key for a Choice, a level index for a Score, `True` or `False` for a yes/no, or the true value for an extraction question. Sending again replaces the earlier label. From 30 labels per question in a project, Curva calibrates the probabilities whenever that makes them more accurate. This guide covers the label formats, the 404 and 422 cases, and which answers to label.

In plain words: Keep the decision id, send the true answer in the format its question type expects, and label some confident answers too, not only the unsure ones.

Start the LLM classifier feedback loop in Python: keep d.id

Every decision has an `id` of the form `dec_…`, unique and sortable by time. Feedback needs it, so store it with whatever you did with the answer: the ticket, the order, the row in your queue.

python
d = client.decide(ticket, questions, project="support")
save(ticket_id, d.id)

Feedback usually arrives much later and from somewhere else: an agent closes a ticket, a reviewer fixes a label, a dispute resolves. If the `id` isn't on the record by then, you can't connect the outcome to the decision. Treat it like a foreign key.

The LLM classifier feedback loop in Python
  1. 1
    client.decide
  2. 2
    store d.id with the record
  3. 3
    act on the answer
  4. 4
    true answer learned later
  5. 5
    client.feedback(id, question, label)
  6. 6
    30+ labels: calibrator fitted if it helps

The decision id travels with the record until the true answer is known, then comes back as a label.

Label formats: option key, level index, True/False, the true value

The label must be a valid answer to the question. Each type has its own form:

python
client.feedback(decision_id, "team", "billing")         # Choice: option key
client.feedback(decision_id, "urgency", 2)              # Score: level index
client.feedback(decision_id, "refund", True)            # Noul: True or False
client.feedback(decision_id, "vendor", "ACME Inc.")     # Text: the true value
client.feedback(decision_id, "po_number", None)         # nullable: it really had none

For extraction questions, Curva compares your value with the one it answered and records right or wrong: text ignoring case and extra whitespace, numbers within a relative 1e-6. That is what the extraction confidence is calibrated against.

Note the Score case: the label is the level's index, not its description. If your levels are "Calm", "Mildly annoyed", "Frustrated but civil" and "Very angry", then "Frustrated but civil" is `2`.

Sending again replaces the earlier label

Sending feedback again for the same decision and question replaces the earlier label. It doesn't add a second one. That makes corrections safe: a reviewer who changes their mind, or an appeal that overturns a verdict, can send the new label and the old one is gone. It also means a retry after a network error can't double-count.

404 for unknown decisions, rule answers and skipped questions; 422 for invalid labels

Two errors cover almost every failed feedback call:

In Python, a 404 raises `NotFoundError` and a 422 raises `InvalidRequestError`, both subclasses of `CurvaError`. Handle them separately, because they mean different things:

python
from curva import CurvaError, NotFoundError, InvalidRequestError

try:
    client.feedback(decision_id, "team", label)
except NotFoundError:
    pass                          # rule answer, skipped question, or wrong id: nothing to learn
except InvalidRequestError as e:
    log.error("bad label for team: %s", e.message)   # fix the mapping in your code
except CurvaError as e:
    retry_later(decision_id, "team", label, e.status)

A 404 for a rule answer or a skipped question is expected. Check `d["team"].rule` before sending if you want to avoid the call: it is `None` when the model answered. A 422 is a bug in your code, usually a label that drifted from the option keys after someone edited the question.

Labels count per project and per exact question wording

Labels are kept per project and per exact question. The `project` you passed to `decide` is the calibration namespace, and the default is `"default"`. Use one project per use case, so feedback from one doesn't calibrate another.

Within a project, calibrators are keyed by a fingerprint of the question's type, wording and options. Rewording a question starts a fresh calibration, with its own 30-label minimum. That is deliberate: a model behaves differently when the wording changes. So settle the wording before you invest in labels.

Label a sample of confident answers, not only abstains

The usual loop sends unsure answers to a person and their answer back as feedback:

python
q = {"team": Choice("Which team?", {"billing": "", "technical": "", "sales": ""}, min_confidence=0.9)}
d = client.decide(ticket, q, project="support")

if d["team"].abstain:
    team = ask_a_human(ticket)                  # a person decides...
    client.feedback(d.id, "team", team)         # ...and Curva learns from it
else:
    route(ticket, d["team"].choice)

That loop on its own gives a skewed sample. Every label comes from below 0.9, so the calibrator learns nothing about how often the 0.95 answers are right, and those are exactly the answers you automate. Also send feedback for a random sample of automated answers. The docs put it plainly: label a sample of automated answers too, so the calibration covers the whole confidence range.

Reading the reply: labels, min_labels: 30, calibrated

Each feedback call returns a small dict:

json
{ "decision_id": "dec_…", "question": "department", "labels": 31, "min_labels": 30, "calibrated": true }

`labels` is how many labels this exact question has in the project. `min_labels` is 30. `calibrated` is `true` once a calibrator has been fitted and kept. It stays `false` past 30 labels if the calibrator doesn't beat the raw probabilities on held-out labels, which happens when the model is already well calibrated. That is not an error.

For the detail, call `client.calibration("team", project="support")`. It returns held-out accuracy, ECE, Brier score and reliability bins before and after calibration.

Next steps

Read the [feedback guide](https://itsmohitrohilla.github.io/curva-docs/guides/feedback/) in the docs and install with `pip install curva-ai`. For the full Python workflow, see [Python LLM classification](/blog/python-llm-classification/), then read the [Python calibration report](/blog/python-llm-calibration-report/). For agents that act on answers, see [an AI agent feedback loop with the decision id](/blog/ai-agent-feedback-loop-decision-id/). The thinking behind it all is in [LLM classification confidence scores](/blog/llm-classification-confidence-scores/).