One signed binary. Every feature compiled in. Free to run. Install Crowkis →
← back to the Roost
curva guidesOctober 3, 2026· 5 min read

Batch LLM classification of a JSONL file with curva map

Classify thousands of records from the command line with any LLM: typed answers per line, rate limits respected, and a rerun resumes where it stopped.

Batch LLM classification means answering the same questions for thousands of records, such as tagging every ticket in an export by team and urgency. With Curva you do it from the command line: put one record per line in a JSONL file, write the questions once, and run `curva map`. Each output line holds typed answers with a probability for every option. The command stays within your provider's rate limits, and if it stops for any reason, rerunning the same command continues after the last line written, so nothing is scored or paid for twice.

The usual alternative is a script with a loop, a retry decorator, a progress file you forgot to write, and a quota error halfway through the file that sends you back to the first line. `curva map` is that script, done once. Here are the steps.

Install and set one provider key

bash
pip install curva-ai
curva --version        # curva 0.1.0
bash
export OPENROUTER_API_KEY=sk-or-v1-...

The wheel includes the `curva` binary for Linux, macOS and Windows. `curva map` calls the model provider directly, so you don't need a running Curva server. OpenRouter is the default provider and free models work. Other providers turn on when their key variable is set, such as `GEMINI_API_KEY` or `GROQ_API_KEY`, and their models are named with a prefix like `@groq/qwen/qwen3.8-27b`. Curva is free to use; you pay only your model provider.

The input: one JSON state per line

`tickets.jsonl` holds one JSON state per line. Blank lines are skipped.

json
{"ticket": "I was charged twice for order A-104"}
{"ticket": "The app crashes on launch"}

Keep your own id in each line if you have one. The model sees the whole line as the state, and the id makes joining the results back easy. The state is fenced as data in the prompt, never treated as instructions, so text inside a record can't redirect the model.

The questions file, exactly as in /v1/decide

`questions.json` is a JSON object of questions, the same shape as the `questions` field of a `/v1/decide` request:

json
{
  "team": {"type": "choice", "instructions": "Which team?",
           "options": {"billing": "payments, refunds", "technical": "bugs"}},
  "refund": {"type": "noul", "instructions": "The customer asks for a refund"}
}

All question types work: Choice, Score, Noul and Multi for labels, and Text, Number and Integer for extraction. All of a line's questions go in one request, so adding a question does not add a round trip.

You don't have to start from a blank file. Curva ships seven recipes, built into the binary: `support-triage`, `content-moderation`, `lead-qualification`, `phishing-check`, `llm-output-qa`, `rag-check` and `rag-rerank`.

bash
curva recipe list
curva recipe show support-triage > questions.json

Edit the wording for your domain before you run a large job.

Run batch LLM classification and read each output line

bash
curva map tickets.jsonl -q questions.json -o answers.jsonl

Each output line has the input line number, the model that answered and the typed answers:

json
{"line": 1, "model": "…", "answers": {"team": {"choice": "billing", "probabilities": {…}, "confidence": 0.99}, "refund": {"noul": 0.02}}}

Every answer is one of the labels you declared, with a probability for each. There is no reply text to parse. A Choice also gets a `none_of_these` option by default, so a record that fits no option says so instead of being forced into one. Sort by `confidence` and you have a review list: read the least confident lines first.

`line` counts from 1 and refers to the input line. Remove blank lines from the input first, and a few lines of Python put each answer next to its record:

python
import json

rows = [json.loads(l) for l in open("tickets.jsonl") if l.strip()]
for out in map(json.loads, open("answers.jsonl")):
    row = rows[out["line"] - 1]
    team = out["answers"]["team"]
    print(row.get("id"), team["choice"], round(team["confidence"], 2))
Batch LLM classification with curva map
  1. 1
    tickets.jsonl
  2. 2
    curva map
  3. 3
    questions.json
  4. 4
    model provider, within rate limits
  5. 5
    answers.jsonl, one line per input line

A rerun reads the output file, skips finished lines and continues with the next one.

Resume after a quota error or Ctrl-C

A daily quota runs out, the network drops, you press Ctrl-C. Run the same command again. It continues after the last line written to `answers.jsonl`, so nothing is scored or paid for twice. Within a run, repeated states come from the cache.

On a free tier, cap the model calls per day with `CURVA_DAILY_LIMIT`. When the limit is reached, requests get a 429 and the run stops. Tomorrow, the same command continues. That cap applies to OpenRouter calls; other providers enforce their own limits. Free tiers change often: on 2026-09-29 a new OpenRouter key got 50 free-model requests a day, and a Groq key about 180 debiased decisions a day.

Fallback chain, council, cascade or race from flags

The multi-model plans of the API work as flags:

Use only one of `--council`, `--cascade` or `--race` per run. A local model works too: `--model @ollama/qwen3:4b` needs no key and no network hop.

By default each question is asked twice, with the options in the original and the reversed order, and the two are averaged. That cancels position bias and costs two calls per line. For a large backfill, `--no-debias` halves the calls. Try it on a sample first and compare the answers before you run the whole file.

Question files can also use `when` (skip a question unless the state matches), `rules` (answer known cases with no model call) and `depends_on` (answer in stages). Since the batch-parity fix after 0.1.0, `curva map` runs them exactly like the API. A question skipped by its `when`, or answered by a rule, costs nothing.

curva map or the API: no server, so no calibration or audit

`curva map` talks to the provider directly. Its answers don't go through a Curva server, so there is no calibration from feedback, no audit log and no API keys.

Use `curva map` for one-off jobs and offline backfills. Use `decide_many` against a server when the answers should be calibrated by your feedback and show up in the audit log and dashboard.

Because a rerun resumes, a scheduler that retries a failed job gets the right behaviour for free. In Airflow:

python
from airflow.operators.bash import BashOperator

classify = BashOperator(
    task_id="classify_tickets",
    bash_command="curva map /data/{{ ds }}/tickets.jsonl -q /opt/curva/questions.json -o /data/{{ ds }}/answers.jsonl",
    env={"OPENROUTER_API_KEY": "{{ var.value.openrouter_key }}"},
    append_env=True,
    retries=3,
)

Next steps

The docs cover this in [batch files with curva map](https://itsmohitrohilla.github.io/curva-docs/guides/batch/) and the [CLI reference](https://itsmohitrohilla.github.io/curva-docs/reference/cli/). Start your questions from [Curva recipes](/blog/curva-recipes/), and before you replace an existing classifier, compare it on logged traffic with a [shadow test](/blog/shadow-test-llm-classifier/). To lower the bill on large files, read [reduce LLM classification cost](/blog/reduce-llm-classification-cost/). For the bigger picture, see [what is Curva](/blog/what-is-curva/).