Logprobs vs verbal confidence: which LLM number to trust
Token logprobs or a stated probability per label? How each mode works, which one auto picks, and the probe results that changed Curva's default model.
Logprobs vs verbal confidence comes down to where the number is read. With logprobs, the model answers with one label token per question and the probability of each label is read from the model's token log-probabilities. With verbal confidence, the model writes a probability for each label, and Curva constrains that reply to JSON over your declared labels so it is never a sentence to parse. Neither wins everywhere. Curva's default, `auto`, uses logprobs when the model returns them and verbal otherwise. And a small probe on 3 tickets showed that a logprobs model can fail a basic consistency check that a verbal model passes, which is why Curva's default model changed.
Logprobs vs verbal confidence: one label token, or JSON with a probability per label
Curva reads probabilities in one of two modes:
Both modes answer every question of a request in one call. Both map onto your declared labels only, so there is no free text in either case. The difference is what the probability means. A logprobs probability is the model's own distribution over the next token. A verbal probability is a number the model chose to write, which makes it more like a judgement than a measurement.
Here is a real response from the docs, answered in logprobs mode:
{
"id": "dec_19294a3c1f2000000",
"model": "inclusionai/ling-3.0-flash-fin:free",
"mode": "logprobs",
"latency_ms": 1144,
"cost_usd": 0.0,
"cached": false,
"answers": {
"department": { "choice": "billing", "probabilities": { "billing": 0.9999, "technical": 0.0, "sales": 0.0001 }, "confidence": 0.9999 },
"frustration": { "score": 0.65, "probabilities": [0.36, 0.62, 0.02], "confidence": 0.62 },
"refund_requested": { "noul": 0.999 }
}
}auto: logprobs when the model returns them, verbal otherwise
The default `mode` is `auto`. It tries logprobs first and falls back to verbal when the model doesn't return them. The result is remembered per full model id, so the fallback happens once, not on every call. The response's `mode` field always says which one actually answered.
You can force either one with `mode: "logprobs"` or `mode: "verbal"`. Forcing logprobs on a question that can't use it is an error, not a silent fallback: you get a 422.
To check a model before you rely on it, `curva spike` reports which models return usable label probabilities. Providers without logprobs, or without JSON-schema output, still work in verbal mode, and then results depend on how well the model follows the schema.
Probe results: Ling (logprobs) failed negation, Nemotron (verbal) passed
The probe that changed Curva's default asks a yes/no question and its negation about the same ticket: "is X?" and "is not X?". The two probabilities should add up to 1.
Ling's sums were far from 1 in both directions: the model ignored the word "NOT". Nemotron, reading verbally, summed to 1.000 each time. Because Ling failed the negation probe and Nemotron passed it, the default config `curva-1.1.0` (which `curva-latest` points to) uses Nemotron 3 Super. `curva-1.0.0` keeps Ling for callers who pinned it.
Read this carefully. It is a probe on 3 tickets, not a benchmark. It compares two models as much as two modes, so it doesn't prove that verbal beats logprobs in general. What it does show is that a token probability is not automatically a trustworthy one. Note too that a Noul in Curva is asked as a normalised two-option choice, which keeps P(yes) and P(no) for one question consistent; the probe tests something harder, two differently worded questions.
Built-in sets: 85% vs 40% of decisions at 0.9, both 100% right on those
The same two models ran Curva's built-in sets (routing, sentiment, policy, adversarial and a coin), 20 rows per model, uncalibrated, on 2026-09-27:
This is the useful nuance. The logprobs model was more decisive: 85% of its answers reached 0.9, and all of those were right. The verbal model was more accurate overall and better calibrated, but more cautious: only 40% reached 0.9, also all right. So on these 20 rows, the logprobs model would have automated more of the work at the same accuracy. With n = 20, none of this is firm. It shows why you measure your own model on your own data rather than choosing a mode by reputation.
Always verbal: over 20 options and any extraction question
Some questions can't use logprobs, whatever the model:
Where logprobs are available, they are the cheaper reply: one token per question, so every question in a request costs little more than one. Keep `mode` at `auto`, or `logprobs` where the model returns them, if cost matters.
Either way, calibrate before trusting thresholds
Both modes give raw probabilities. Neither is calibrated to your data out of the box. A logprobs model can be overconfident; a verbal model can write tidy numbers that don't match its hit rate. The fix is the same for both: send feedback, and after 30 labels per question Curva fits a calibrator and keeps it only when it beats the raw numbers on held-out labels.
One caution if you change modes or models. Calibrators belong to the exact question and project, and the `auto` decision is remembered per model id. Switching from a logprobs model to a verbal one changes the raw numbers the calibrator learned from. Re-check the calibration report after any switch.
Next steps
The docs explain both modes in [questions and answers](https://itsmohitrohilla.github.io/curva-docs/concepts/questions/) and show the probes on the [benchmarks page](https://itsmohitrohilla.github.io/curva-docs/benchmarks/). Test a model with [check LLM logprobs support](/blog/check-llm-logprobs-support/), see why the negation probe matters in [LLM negation consistency](/blog/llm-negation-consistency/), and read [LLM classification with many classes](/blog/llm-classification-many-classes/) for the 20-option limit. The full picture is in [LLM classification confidence scores you can act on](/blog/llm-classification-confidence-scores/).