LLM yes or no questions with a real probability of yes
A Noul question returns one number, P(yes), asked as a normalised two-option choice so P(x) and P(not x) agree. How to word it, threshold it and label it.
To get an LLM yes or no answer with a real probability, ask for the probability of yes instead of the word. In Curva that question type is called a Noul. You write a statement such as "The customer explicitly asks for a refund", and the answer is one number, `noul`, the probability that it is true. Curva asks it as a normalised two-option choice, so P(yes) and P(no) always add up and a question and its negation stay consistent. In Python the plain answer is `True` when P(yes) is at least 0.5, and the full answer keeps the number for your own threshold. This guide covers how to word a yes/no question, threshold it, and label it.
LLM yes no probability in one number: noul is P(yes)
Noul("The customer explicitly asks for a refund")A Noul is a yes/no question, and its answer is a single number: `noul`, the probability of yes. There is no separate confidence field, because the number already says how sure the model is. 0.97 is a confident yes, 0.03 a confident no, and 0.5 means the model can't tell.
Write the question as a statement, not a question with "or". "The customer explicitly asks for a refund" reads as a claim the model can judge true or false, and P(yes) is the probability that the claim holds for this state.
A Noul is answered in the same request as your other questions, whether they are Choices, Scores or extraction. All of a state's questions go to the model together.
Asked as a normalised two-option choice, so P(x) and P(not x) agree
Internally, Curva asks a Noul as a choice between two options, yes and no, and normalises the probabilities over those two. So P(yes) and P(no) for one question always sum to 1. With order debiasing on, it is also asked with the two options in both orders and averaged, so a model's lean towards whichever option comes first cancels out.
Curva's docs tie this mechanism to a trust probe: ask "is X?" and "is not X?" about the same input, and the two probabilities should add up to 1.
In a probe on 3 tickets (2026-09-27), Nemotron 3 Super in verbal mode summed to 1.000 each time. Ling 3.0 Flash in logprobs mode summed to 1.33, 0.14 and 0.56: it ignored the word "NOT". Curva's default config moved to Nemotron because of that result. A tiny probe, but a useful warning that a yes/no probability is only as good as the model behind it.
Plain answers: True when P(yes) is at least 0.5
The Python shorthand gives you a boolean when that is all you need:
import curva
d = curva.decide("I was charged twice, please refund me",
{"team": ["billing", "technical"], "refund": "Asks for a refund?", "total": float})
print(d.team, d.refund, d.total) # billing True None`d.refund` is `True` when P(yes) is at least 0.5. `d["refund"].noul` is the probability itself. The TypeScript client does the same: `d.values` holds `true` for the refund question when P(yes) is at least 0.5, and `d.answers.refund.noul` is the number.
The 0.5 cut-off is fine for a script. For anything that triggers an action, read the probability and choose your own cut-off, as below.
Spell out the no case in the instructions
The most useful line in a yes/no question is often the one that says what counts as no. Curva's shipped recipes do this on purpose:
Without the first line, an angry message about a double charge reads like a refund request, and P(yes) climbs on tickets that only complain. A sentence that names the near miss is cheaper and more reliable than any threshold you could tune afterwards.
Noul has no min_confidence: pick your own P(yes) cut-off
Choice and Score questions take `min_confidence` and come back with `abstain: true` below it. A Noul doesn't, because its answer is one number. Route it with two cut-offs and treat the middle as unsure:
p = d["refund"].noul
if p >= 0.9:
start_refund_flow(ticket)
elif p > 0.1:
send_to_human(ticket, decision_id=d.id)Below 0.1 is a confident no, above 0.9 a confident yes, and the band between goes to a person. The numbers are yours to set. Before calibration they are guesses about the model. After 30 labels, Curva fits a Platt scaling calibrator for the question (a logistic fit on P(yes)) and keeps it only when it improves on the raw numbers on held-out labels. Then your cut-offs mean what they say.
For a guarantee instead of hand-picked cut-offs, set `coverage` on the Noul. Once it has 30 labels, the answer carries a `set` that is a subset of `["true", "false"]` and contains the right value at least `coverage` of the time. A set with one value: act on it. A set with both: the model can't tell, so send it to a person.
Feedback with true or false, rules with answer true
When you learn the truth, send it back with `True` or `False`:
client.feedback(decision_id, "refund_requested", True)
Over HTTP the label is `true` or `false`. Sending feedback again for the same decision and question replaces the earlier label.
Some yes/no answers you already know without a model. A rule answers them on the spot:
"priority": {
"type": "noul",
"instructions": "This ticket needs a reply today",
"rules": [{"if": {"plan": ["enterprise", "premium"], "open_tickets": {"gte": 3}}, "answer": true}]
}A rule answer puts all its probability on the label, so `noul` is 1.0 here, and it carries `rule` with the index of the rule that matched. Rule answers are not model output, so they are not used for calibration, and feedback for them gets 404.
Shorthand: a string ending in ? or bool
For scripts and notebooks, two shorthands create a Noul:
In TypeScript, a string ending in `?` or `Boolean` does the same. Switch to an explicit `Noul` in production code, so the wording is explicit and stays fixed, because calibration belongs to the exact wording of a question.
Next steps
The docs describe Noul in [questions and answers](https://itsmohitrohilla.github.io/curva-docs/concepts/questions/). For all the types side by side, read [LLM classification question types](/blog/llm-classification-question-types/), and for wording, [how to word LLM classification questions](/blog/write-llm-classification-instructions/). See a Noul at work in [refund request detection](/blog/refund-request-detection/), and the probe in detail in [LLM negation consistency](/blog/llm-negation-consistency/). For the product overview, read [what is Curva](/blog/what-is-curva/).