Abstain or a coverage set? Selective classification for LLMs
Abstain checks the top answer; a conformal set checks every option. A decision table by question type and the routing rule that combines both.
Selective classification vs conformal prediction is a choice between two ways of handling an unsure LLM answer. Selective classification lets the classifier abstain: if the top answer's confidence is below a threshold, it hands the item to a person. Conformal prediction keeps every answer but attaches a set of options that contains the true one at a promised rate. In Curva the first is `min_confidence`, which sets `abstain: true`, and the second is `coverage`, which adds a `set`. Abstain looks at the top answer; a coverage set looks at every option. They answer different questions, they work on different question types, and the strongest routing rule uses both.
Two questions: is the top answer good enough, which options must I keep
Every routing decision on an LLM answer comes down to one of two questions.
A word on vocabulary, because the two fields collide. In the selective classification literature, "coverage" usually means the share of inputs the classifier answers instead of abstaining. In Curva, `coverage` is the conformal target: the probability that the set contains the true answer. Curva's own name for the share answered is `automated`, in the calibration report.
Selective classification vs conformal prediction: support by question type
The two are not available on the same question types:
Two gaps are worth knowing. A Noul has no `min_confidence`: its answer is one number, P(yes), so you route it with your own thresholds, treating the middle band as unsure, or you set `coverage` and read the set. A set of `["true", "false"]` means the model can't tell. And a Multi accepts `coverage`, but feedback labels a Multi as a whole rather than option by option, so its sets carry no guarantee yet. Setting `coverage` on a type that doesn't support it gets 422.
What each needs: a calibrated threshold vs 30 representative labels
Both tools work from day one, but neither is fully meaningful until it has labels.
**Abstain needs calibration.** `min_confidence` compares the answer's confidence with your bar. Before the question is calibrated, that confidence is the model's own number, debiased across two option orders but not checked against your data. A bar of 0.8 is then a guess about the model. Once your feedback has fitted a calibrator (after 30 labels, and only when it beats the raw probabilities on held-out labels), the threshold means what it says.
**A coverage set needs 30 representative labels.** The set is built by split conformal prediction from the question's own feedback labels. Before 30 labels, the answer has `guaranteed: false` and no `set`. After 30, the guarantee holds whatever the model, as long as the labels are a fair sample of your traffic. Calibration helps by making sets smaller, but the guarantee doesn't depend on it.
Combined rule: a one-option set and abstain false
The two checks catch different failures, so combine them. Curva's docs give a common routing rule: automate when the set has one option and `abstain` is false, and send everything else to review.
q = Choice("Which team?", ["billing", "technical", "sales"], min_confidence=0.8, coverage=0.95)
d = curva.decide(ticket, {"team": q}, project="support")
a = d["team"]
if a.guaranteed and len(a.set) == 1 and not a.abstain:
route(ticket, a.choice)
else:
review(ticket, candidates=a.set or [a.choice])Why both? A one-option set says the model is sure enough to meet the coverage guarantee alone. `abstain: false` says the top answer also clears your own bar. Either check alone can let an answer through that the other would stop. And the `guaranteed` check guards the start: until the question has 30 labels, there is no set and everything goes to review.
Showing the set to a reviewer as candidates
The `else` branch above doesn't just send the item to a person. It sends the candidates. That is where a coverage set earns its keep even when it can't automate.
A reviewer choosing between two likely teams works faster than one starting from the full list. When the reviewer decides, send the answer back with `client.feedback(decision_id, "team", "technical")`. Each review adds a label that feeds both calibration and the conformal threshold.
Metrics to watch for each
Each tool has its own health signal.
**For abstain**, watch how much you automate and how well it goes. The calibration report gives `automated`, the share of answers at 0.9 or above, and `accuracy_when_automated`, how often those were right, before and after calibration. The server's `curva_abstains_total` counter, and the abstain rate on the dashboard's overview, show how much traffic reaches people.
**For a coverage set**, watch two things. First, the size of the sets: the share with one option is how much you can automate under the guarantee. Second, whether the sets still cover. Label a random sample of recent decisions and count how often the true answer was inside the set. At 0.95 coverage, expect about one miss in twenty; many more suggests the labels no longer match your traffic, and the weekly drift report is the place to check.
Neither replaces the other. Abstain is the simpler tool and works on Scores and extraction. Coverage gives a promise you can state in a contract with your own team. Use abstain everywhere it applies, and add coverage where a stated rate matters.
Next steps
The docs describe both in [guaranteed accuracy](https://itsmohitrohilla.github.io/curva-docs/guides/guaranteed-accuracy/) and [debiasing, escape and abstain](https://itsmohitrohilla.github.io/curva-docs/concepts/trust/). Go deeper on each side with [LLM confidence threshold and abstain](/blog/llm-confidence-threshold-abstain/) and [conformal prediction for LLM classification](/blog/conformal-prediction-llm/). For the accuracy you get at each threshold, read [selective accuracy for LLMs](/blog/llm-selective-accuracy/), and for the full picture, [LLM classification confidence scores you can act on](/blog/llm-classification-confidence-scores/).