Classify banking intents with 77 options
Build a 77-intent banking classifier with one Choice: why over 20 options Curva switches to verbal mode, and what Gemini scored on BANKING77 (n = 120).
Banking intent classification with an LLM can use one Choice question with all 77 intents as options. Curva accepts up to 255 options, adds a `none_of_these` escape option, and returns one intent with a probability for each. Over 20 options it reads those probabilities in verbal mode, because providers return at most 20 logprobs. On the public BANKING77 set, Gemini flash-lite scored 79.8% (95% range 73% to 87%, n = 120, 2026-10-01), which matches Jev's published 75.3% within the margin. This post covers the mechanics, the descriptions, the measured numbers and the cost.
Banking intent classification: 77 options in one Choice
A Choice picks exactly one option. Options map a key to a description, and the description may be empty. The order is kept. A Choice takes 2 to 255 options, so a full banking intent list fits in one question:
from curva import Curva, Choice
client = Curva()
INTENTS = {
"card_arrival": "asks when a new card will arrive",
"card_not_working": "a card is declined or not accepted",
"lost_or_stolen_card": "reports a lost or stolen card",
"pending_top_up": "a top-up shows as pending",
"top_up_failed": "a top-up was rejected or failed",
# ... one entry per intent, 77 in all
}
q = {"intent": Choice("What does the customer want?", INTENTS, min_confidence=0.8)}
d = client.decide({"message": "My card got declined at the shop twice today"}, q, project="banking")
print(d["intent"].choice, d["intent"].confidence)Curva also adds a `none_of_these` option by default, so a message that fits no intent isn't forced into one. In banking, that catches the complaint about branch opening hours that your list never planned for.
The question is asked as one request. If you also want an urgency Score or a "customer is upset" yes/no, add them to the same request; up to 64 questions are answered together.
Over 20 options: verbal mode
Curva reads probabilities in one of two modes. In `logprobs` mode the model answers with one label token per question, and Curva reads each label's probability from the token log-probabilities. In `verbal` mode the model returns a JSON object limited to the declared labels, with a probability for each.
Providers return at most 20 logprobs per position. A 77-option question can't be read that way, so Choices with more than 20 options are always answered in verbal mode. The changelog adds that these are answered with the top 5 labels. Setting `mode: logprobs` on such a question gets a 422.
Three things follow from that:
If the flat list is too much for your model, split it. A first Choice picks a category (cards, transfers, top-ups, account), and a second question that uses `depends_on` picks the intent within it, seeing the first answer. Each stage is one model call, at most 8 stages per request.
Writing descriptions for close intents
BANKING77 is hard because many intents are near twins: a declined card against a card that doesn't work at all, a pending top-up against a failed one, a transfer that hasn't arrived against one that was declined. The descriptions are what the model reads, so they carry the distinctions.
Models also tend to favour options by position, and a long list gives position bias more room. Order debiasing, on by default, asks the question twice with the options in original and reversed order and averages the two.
Measured on BANKING77
Curva's public benchmark runs use the BANKING77 set with 77 options, debiasing on, and samples balanced across intents. Jev's 75.3% is published by others (sanand0 llmevals, 77 items) on a different sample, so compare with care.
All rows are from 2026-10-01. "After calibration" means each half of the rows was calibrated by a fit on the other half, which is what you get once feedback flows in. Gemini's raw probabilities were already well calibrated, so calibration left them alone. Groq's were overconfident, and calibration brought ECE from 0.252 to 0.157.
The early paid runs, at n = 20 each and dated 2026-10-01, are too small to rank: gpt-4.1-mini 75.0%, Claude Haiku 4.5 65.0%, gpt-4o-mini 60.0%, and gpt-4.1-nano 50.0%, which is a loss against Jev's number. At n = 20 the range is about plus or minus 20 points. Treat these as a reason to test your own model, not as a verdict.
Gemini was accurate but slow here: a median of 1,362 ms per decision. Groq answered in 351 ms. If you need fast answers on a big intent list, measure a fast provider on your own labels.
Cost: the most expensive set we ran
Long option lists mean long prompts, and with debiasing on each decision is two calls. BANKING77 was the most expensive of the 8 public sets in the paid runs. Measured from `cost_usd`, 20 decisions per model on the set, 2026-09-30, debiasing on:
Figures are truncated, not rounded up. To bring the cost down, answer the obvious intents with `rules` (a message containing "stolen" can go straight to the lost card flow), let the decision cache answer exact repeats, and try `debias: "auto"`, which drops the second call once a question has shown no position bias.
Next steps
Read the [questions reference](https://itsmohitrohilla.github.io/curva-docs/concepts/questions/) in the docs and install with `pip install curva-ai`. For the full benchmark write-up, see [the BANKING77 LLM benchmark](/blog/banking77-llm-benchmark/). For the general problem of big label sets, read [LLM classification with many classes](/blog/llm-classification-many-classes/). A related fine-grained list is [chargeback reason code classification](/blog/chargeback-reason-classification/), and more ideas are in [LLM classification use cases](/blog/llm-classification-use-cases/).