One signed binary. Every feature compiled in. Free to run. Install Crowkis →
← back to the Roost
curva use casesOctober 3, 2026· 5 min read

Support ticket triage with calibrated confidence

The support-triage recipe end to end: route confident tickets, queue unsure ones, answer obvious ones with rules, and learn from agents' corrections.

Support ticket triage with AI works when the model returns a typed answer per ticket (which team, how urgent, does the customer want a refund) with a probability you can act on, and sends the unsure tickets to a person instead of guessing. Curva's built-in `support-triage` recipe asks five such questions in one call. Confident tickets route themselves, tickets below 0.8 confidence go to a review queue, rules settle the obvious cases with no model call, and every agent correction becomes a calibration label. This post builds that flow end to end.

Triage is a good first job for a language model. Tickets are text, the decisions are small, and a wrong call costs minutes, not money. It is also where naive triage fails quietly: the model sends a login problem to billing "with high confidence", nobody notices, and the customer waits a day in the wrong queue.

Support ticket triage AI: the five questions in the recipe

Get the recipe from the binary:

bash
pip install curva-ai
export OPENROUTER_API_KEY=sk-or-v1-...     # any OpenRouter key; free models work
curva recipe show support-triage > triage.json

It asks five questions per ticket, all in one request:

The wording carries lessons. `team` says "judge by what the customer needs done, not by the words they happen to use", so "I can't log in to see my invoice" goes to `account`, not `billing`. `refund_requested` says that complaining about a charge without asking for money back is no. The urgency levels are defined, not just named: High means "the customer cannot use the product, or money is being lost right now", and Critical means "an outage, a security or data-loss issue, or many users affected".

Triage one ticket

python
import json
import curva

TRIAGE = json.load(open("triage.json"))
client = curva.local()

ticket = {"subject": "Charged twice??",
          "body": "I was charged twice for order A-104. Please refund the duplicate. "
                  "If this keeps happening I'm cancelling."}
d = client.decide(ticket, TRIAGE, project="support")

d["team"].choice, d["team"].confidence, d["team"].abstain
d["urgency"].level            # most likely urgency level
d["frustration"].score        # expected level, 0 = Calm
d["refund_requested"].noul    # P(yes)
d["churn_risk"].noul          # P(yes)

Every answer is one of the labels in the file. `team` can also come back as `none_of_these`, the escape option every Choice gets, for a ticket that fits no team. The ticket itself is fenced as data in the prompt, so a customer who writes "ignore your instructions" is just a customer.

Route confident tickets, queue the rest

Support ticket triage AI flow
  1. 1
    ticket
  2. 2
    rule matches?
  3. 3
    team queue, no model call
  4. 4
    five questions, one request
  5. 5
    team abstains or none_of_these?
  6. 6
    review queue
  7. 7
    agent picks the team
  8. 8
    feedback: a calibration label

Rules settle known cases first; unsure or unmatched tickets go to people, whose answers flow back as labels.

The routing code checks abstain first, because an unsure answer still has a `choice`:

python
def route(ticket, d):
    team = d["team"]
    if team.abstain or team.choice == "none_of_these":
        return review_queue.add(ticket, decision_id=d.id)

    queue = queues[team.choice]
    priority = "p1" if d["urgency"].score >= 2.5 else "p2" if d["urgency"].score >= 1.5 else "p3"
    queue.add(ticket, priority=priority, decision_id=d.id)

    if d["churn_risk"].noul >= 0.8:
        notify_account_manager(ticket)

Store `d.id` with the ticket. It is how feedback finds the decision later.

Priority uses `score`, the expected level, rather than the single most likely level. A ticket whose score lands between High and Critical keeps that doubt visible in the number. You choose the cut-offs.

Rules for the obvious ones

Some routing is not a judgement call. An enterprise customer with three or more open tickets always gets a reply today. A ticket body that starts with an error code goes to technical. Rules answer those instantly, with no model call, and the model sees only the rest:

python
from curva import Choice, Noul

team = TRIAGE["team"]
questions = dict(TRIAGE)
questions["team"] = Choice(team["instructions"], team["options"], min_confidence=0.8) \
    .rule("technical", body={"starts_with": "Error"})
questions["reply_today"] = Noul("This ticket needs a reply today") \
    .rule(True, plan=["enterprise", "premium"], open_tickets={"gte": 3})

d = client.decide({**ticket, "plan": "enterprise", "open_tickets": 4}, questions, project="support")
d["reply_today"].rule      # 0: the first rule answered it

Rules read top-level fields of the state. Up to 32 per question, first match wins. A rule answer puts all its probability on that label, has `calibrated: false`, and says which rule gave it, so the audit log shows why. If every question is answered by rules, no model is called and the decision costs $0. Keep rules to what you would write in code anyway, and drop a rule once the model agrees with it.

Feedback from agents calibrates the team question

Your agents already correct triage: they reassign tickets. Send each correction back:

python
client.feedback(decision_id, "team", "account")        # the team that really handled it
client.feedback(decision_id, "refund_requested", True)
client.feedback(decision_id, "urgency", 2)              # level index: 0 = Low

After 30 labels per question in the `support` project, Curva fits a calibrator, and keeps it only when it improves the probabilities on held-out labels. A model that is already well calibrated is left alone. Once calibrated, the 0.8 cut-off on `team` means what it says on your tickets. Check it:

python
report = client.calibration("team", project="support")
print(report["before"]["ece"], "→", report["after"]["ece"])
print(report["after"]["accuracy_when_automated"], report["after"]["automated"])

`automated` is the share of answers at 0.9 confidence or above, and `accuracy_when_automated` is how often those were right. Together they tell you how much triage you can hand over. Send feedback for a sample of tickets that routed themselves too, not only the reviewed ones, or calibration sees only the hard cases.

Shadow-test before go-live

Test against what you do today before any ticket moves. Export a week of tickets with the team that handled each one, in feedback format:

json
{"state": {"ticket": "I was charged twice"}, "labels": {"team": "billing", "refund_requested": true}}
bash
curva shadow traffic.jsonl -q triage.json -o shadow.jsonl

The Markdown report shows agreement with your current routing per question, agreement in each confidence band, and the share of traffic you could automate at or above each band. It lists the most confident disagreements first. Read them: in each one, either Curva or your current routing is clearly wrong. Agreement is not accuracy, so label a sample of disagreements and send them as feedback.

Keep an eye on it after launch

The recipe is a starting point. Rename the teams, add your own, write the descriptions in your agents' words. Settle the wording before you collect feedback: calibration belongs to the exact question, and rewording starts it over.

Next steps

The docs cover [recipes](https://itsmohitrohilla.github.io/curva-docs/guides/recipes/), [rules](https://itsmohitrohilla.github.io/curva-docs/guides/rules/) and [shadow mode](https://itsmohitrohilla.github.io/curva-docs/guides/shadow/). Go deeper on one part with [LLM ticket routing in Python](/blog/llm-ticket-routing-python/), [refund request detection](/blog/refund-request-detection/) or [ticket urgency and SLA rules](/blog/ticket-urgency-sla-rules/). For other workloads, see [LLM classification use cases](/blog/llm-classification-use-cases/).