Expected calibration error (ECE) for LLM classifiers
How expected calibration error is computed, worked by hand on 100 answers, and why an LLM can match on accuracy yet lose on ECE, with measured numbers.
Expected calibration error (ECE) is the average gap between how confident a classifier says it is and how often it is actually right, weighted by how many answers fall at each confidence level. You sort answers into bins by confidence, take the absolute difference between each bin's average confidence and its accuracy, and average those gaps by bin size. 0 is perfect; lower is better. For an LLM classifier, ECE is the number that tells you whether a threshold like "automate above 0.9" means what it says. This post computes it by hand, shows why accuracy and ECE are separate verdicts, and shows where it appears in Curva's calibration report.
ECE in one sentence: the weighted gap between confidence and hit rate
Written out, for answers split into bins b:
ECE = sum over bins of (answers in b ÷ all answers) × |accuracy in b − average confidence in b|
Three things follow from the formula. The gap is absolute, so overconfidence and underconfidence both count. Bins with more answers weigh more, so a few odd answers at 0.5 barely move it. And it says nothing about accuracy itself: a model that is right half the time and always says 0.5 has an ECE near 0.
Curva's docs define it the same way: "the average gap between confidence and accuracy, weighted by how many answers fall in each bin". Lower is better, and 0 is perfect.
Worked example: expected calibration error by hand
Take a made-up run of 100 answers that fall into two confidence bins. The numbers are invented to show the arithmetic.
Made-up example: the high bin claims 0.9 and is right 70% of the time, the middle bin claims 0.6 and is right 60% of the time, so ECE is 0.12.
Step by step: the high bin holds 60 of 100 answers, so its weight is 0.6. Its answers claim 0.9 and 42 of 60 are right, an accuracy of 0.70 and a gap of 0.20. The middle bin's 40 answers claim 0.6 and 24 are right, so its gap is 0. ECE = 0.6 × 0.20 + 0.4 × 0 = 0.12.
Why the overconfident bin is the one you automate
Look at where the error sits. All of it comes from the high bin, the answers at 0.9. That is exactly the bin a rule like "automate at 0.9 or above" sends straight through. In the example, three in ten of those automated answers are wrong, and nothing in a log of confidences would tell you.
This is why ECE matters more for LLM classifiers than a single accuracy number does. Raw LLM probabilities tend to run high. A model that says 0.95 on answers it gets right 70% of the time is overconfident, and every threshold built on its numbers fails quietly. ECE measures that failure directly.
Pair it with the third number Curva reports, accuracy when automated: the accuracy on answers at or above 0.9, next to the share of answers that reach 0.9. ECE says whether the scores are honest. Accuracy when automated says what an honest 0.9 buys you.
Tie on accuracy, loss on ECE: BoolQ 88.9% with ECE 0.081 (n = 127)
Accuracy and calibration are separate verdicts, and Curva's own public runs show it.
All three rows are from 2026-10-01, with order debiasing on. On BoolQ, Gemini flash-lite matches Jev's published accuracy within the margin but has twice the calibration error: it is right as often, but its confidence is less honest. On PhishNChips, Groq beats Jev's published accuracy and still loses on ECE. Only Gemini on PhishNChips wins both. Jev's numbers are published by others on different samples, so compare with care.
The lesson for your own evaluation: report both. An accuracy number alone can hide a model whose high-confidence answers are right far less often than they claim.
Reading ece in the before and after blocks of GET /v1/calibration
Curva computes ECE for every question that has feedback labels. From Python:
report = client.calibration("team", project="support")
print(report["before"]["ece"], "→", report["after"]["ece"])
print(report["after"]["accuracy_when_automated"], report["after"]["automated"])Over HTTP it is `GET /v1/calibration?question=<key>&project=<name>`. `before` holds the raw probabilities. `after` is held out: each half of the labels is calibrated by a fit on the other half, so the ECE isn't flattered by testing on the data the calibrator was fitted to. Both blocks also carry `accuracy`, `brier`, `automated`, `accuracy_when_automated` and the `reliability` bins, which are the per-bin numbers ECE is built from.
A gap between `before` and `after` tells you what calibration would do. If the calibrator didn't clear Curva's held-out gate, answers stay raw and say `calibrated: false`, even when `after` looks a little better: a small gain that could be noise doesn't count.
The 0.03 target and where measured runs are today
How low is low enough? Curva's unit tests give one reference point: with 1,000 labels, an overconfident model goes from ECE 0.25 to under 0.03. That is what calibration can do with plenty of labels on a clean signal.
Measured runs on public sets are not there yet. On the quiz-style sets, raw ECE for the free models was 0.032 to 0.118 (n = 98 to 127, 2026-10-01), against published Jev numbers of 0.024 to 0.038. Calibration left those sets raw, because it couldn't show a clear held-out gain. Where a model leaned hard on one answer, calibration moved a lot: Gemini on AITA went from 0.404 to 0.149 held out (n = 100). Samples this size carry wide error bars, so treat them as first measurements.
The practical target is relative, not absolute: an ECE low enough that `accuracy_when_automated` at your threshold is the accuracy you need.
Next steps
The docs define the metrics in [probabilities and calibration](https://itsmohitrohilla.github.io/curva-docs/concepts/calibration/). For the full mechanism, read [LLM calibration explained](/blog/llm-calibration-explained/). For the other two lenses, see the [Brier score for LLMs](/blog/brier-score-llm/) and [reliability diagrams](/blog/reliability-diagram-llm/). For how it all drives automation, read [LLM classification confidence scores you can act on](/blog/llm-classification-confidence-scores/).