One signed binary. Every feature compiled in. Free to run. Install Crowkis →
← back to the Roost
curva conceptsOctober 3, 2026· 7 min read

LLM calibration explained: when 0.9 means 90%

What LLM calibration means, how ECE and Brier measure it, and how Curva fits a calibrator from 30 feedback labels, with held-out before and after numbers.

LLM calibration means that a model's confidence matches how often it is right: of all the answers given at 0.9, about 90% should be correct. Raw LLM probabilities usually miss that mark, most often by being too sure. You measure the gap with expected calibration error (ECE) and the Brier score, and you close it by fitting a small calibrator on labels from your own data. Curva does this per question once it has 30 feedback labels, and applies the calibrator only when it improves on the raw probabilities on held-out labels. This post explains each step, with real before and after numbers, including runs where the right call was to change nothing.

A 0.9 that is right 70% of the time breaks every threshold

Say your classifier returns a confidence with every answer, and you automate everything at 0.9 or above. That rule is only as good as the 0.9. If the model says 0.9 on answers it gets right 70% of the time, three in ten of your "safe" automations are wrong. Nothing errors. The wrong items just move on with a high score attached.

That model is overconfident, and it is the common case. Language models are trained to produce likely text, not honest probabilities. A confidence the model writes in its reply is generated text, and tends to be high. Token log-probabilities are better, but they still reflect how the model reads your prompt and options, not how often it is right on your data. And the same question can be well calibrated on one model and badly overconfident on another.

Reliability bins: checking a model against the diagonal

To check calibration, group answers by confidence and compare each group's confidence with its accuracy. Take every answer given at about 0.8: roughly 80% of them should be right. The same holds low in the range and near the top.

Plot average confidence against accuracy, bin by bin, and you get the calibration curve, also called a reliability diagram. Curva is named after that curve. A calibrated model sits on the diagonal. An overconfident one sits below it: it says 0.9 and is right far less often. Curva's calibration report includes the `reliability` bins, and the built-in dashboard draws the diagram.

ECE, Brier and accuracy when automated in one table

Curva reports three numbers, before and after calibration:

A made-up example makes ECE concrete. Of 100 answers, 60 land in the 0.9 bin and 42 of those are right: accuracy 0.70, a gap of 0.20. The other 40 land in the 0.6 bin and 24 are right: accuracy 0.60, no gap. ECE is 0.6 × 0.20 + 0.4 × 0 = 0.12. The whole error comes from the confident bin, which is exactly the one you would automate. These numbers are invented for illustration.

ECE tells you whether the scores are honest. Accuracy when automated tells you what the scores are worth: if you automate at 0.9, how much traffic is that, and how often is it right.

How LLM calibration works in Curva

Calibration needs one thing from you: the true answer, whenever you learn it. Every decision has an `id`. When an agent closes the ticket or a reviewer fixes a label, send it back with `client.feedback(decision_id, "team", "billing")`. The label is the option key for a Choice, the level index for a Score, and `True` or `False` for a Noul.

The LLM calibration loop in Curva
  1. 1
    decide: raw probabilities
  2. 2
    you learn the true answer
  3. 3
    POST /v1/feedback
  4. 4
    30 labels for this exact question?
  5. 5
    fit calibrator, test on held-out folds
  6. 6
    clear gain in log-loss and Brier?
  7. 7
    answers return calibrated: true
  8. 8
    answers stay raw, calibrated: false

Labels flow back per question; from 30 of them a calibrator is fitted and kept only if it helps on held-out labels.

Temperature, bias and Platt scaling fitted after 30 labels

After 30 labels for the same exact question in a project, Curva fits one of three small calibrators:

These are deliberately small models, with one or two parameters, or one per answer for bias scaling. They fit well from a few dozen labels and can't overfit the way a large model would. With few labels they also move cautiously: about a third of the way to what the labels suggest at 30 labels, nearly all the way at 1,000.

Kept only when held-out log-loss and Brier both improve

A calibrator that makes things worse is worse than none. So Curva tests before it applies. It fits on four fifths of the labels and scores on the fifth it has not seen, five times over. It keeps the calibrator only if all three hold:

Otherwise answers stay raw, with `calibrated: false`. An already well-calibrated model is left alone. Curva's unit tests cover the mechanics: an overconfident model goes from ECE 0.25 to under 0.03 with 1,000 labels, a yes/no model that always says 1.0 moves to its true 70%, and a calibrated model keeps its raw probabilities.

Calibrators belong to a fingerprint of the question's type, wording and options, and to a project. Reword a question and its calibration starts over. Settle the wording before you collect labels.

Read the calibration report

python
report = client.calibration("team", project="support")
print(report["before"]["ece"], "→", report["after"]["ece"])
print(report["after"]["accuracy_when_automated"], report["after"]["automated"])

`before` is the raw model. `after` is held out: each half of the labels is calibrated by a fit on the other half, so the number isn't flattered by testing on the data the calibrator was fitted to. Both include `accuracy`, `ece`, `brier`, `automated` and `accuracy_when_automated`, plus `reliability` bins. `calibrator` shows the fitted parameters.

Real runs: AITA 0.404 to 0.149, Groq phishing 0.283 to 0.209, quiz sets left raw

These are held-out results from Curva's public benchmark runs on free-tier models, with order debiasing on, dated 2026-10-01. The benchmark fits each half on the other half's labels.

Three readings. Bias scaling fixed models that lean: Gemini's AITA error fell from 0.404 to 0.149 (0.230 with temperature alone), and on Yelp from 0.259 to 0.138, with accuracy up from 54% to 59%. Calibration improved an overconfident model without making it good: Groq on phishing went from 0.283 to 0.209, still above the 0.154 Jev publishes for that set. And the gate left good models alone. An earlier, looser gate let a harmful fit through on Gemini CommonsenseQA (0.118 to 0.176); the stricter gate keeps it raw.

These are samples of 69 to 127 rows. Treat them as first measurements on public data, not a promise for yours. Calibration can't create information that isn't in the labels.

Send feedback from the whole confidence range

Four rules follow from all this:

Next steps

The docs cover this in [probabilities and calibration](https://itsmohitrohilla.github.io/curva-docs/concepts/calibration/) and [calibrate with feedback](https://itsmohitrohilla.github.io/curva-docs/guides/feedback/). Go deeper on [expected calibration error](/blog/expected-calibration-error-llm/), the [held-out gate](/blog/llm-calibration-held-out-gate/) and [temperature scaling vs Platt scaling](/blog/temperature-scaling-vs-platt-scaling/). To act on calibrated scores, read [LLM confidence threshold and abstain](/blog/llm-confidence-threshold-abstain/), and for the full picture, [LLM classification confidence scores you can act on](/blog/llm-classification-confidence-scores/). Install with `pip install curva-ai`.