One signed binary. Every feature compiled in. Free to run. Install Crowkis →
← back to the Roost
curva guidesOctober 3, 2026· 5 min read

LLM image classification in Python with probabilities

Classify images with a vision LLM from Python: paths, bytes or URLs in images=, typed answers with probabilities, and the privacy and size limits.

For LLM image classification in Python, pass the pictures to `decide` with the `images=` argument and ask typed questions about them. Curva accepts file paths, `Path` objects, raw bytes, `data:` URIs and `https://` URLs, up to 8 images per decision. A vision-language model judges the images together with the state you send, and every answer comes back typed with a probability: P(yes) for a yes/no, a probability per option for a Choice. This guide covers the input forms, a receipt check that compares an image with a claimed amount, the privacy rule for URLs, the limits, and what images do to cost and caching.

In plain words: `images=` takes paths, bytes or URLs, and the model judges the picture and your data together. Send private images as bytes, because a URL is fetched by the provider.

images= accepts paths, Path, bytes, data: URIs and https URLs

Install the SDK with `pip install curva-ai`. It uses only the standard library and needs Python 3.9 or newer. The `images=` argument takes a list, and each item can be:

Files and bytes are always sent inline. You don't need to encode anything yourself. Mixing forms in one list is fine, so a ticket can carry one uploaded file and one image you rendered in memory.

LLM image classification in Python: does the receipt match the claim?

The useful part of image classification in a workflow is rarely "what is in this picture". It is "does this picture agree with the record". An expense claim says "Team lunch, 4 people, 86.40". The receipt either shows that or it doesn't. Here is the docs' example:

python
from curva import Curva, Choice, Noul

curva = Curva()
d = curva.decide(
    {"expense_claim": "Team lunch, 4 people", "amount": 86.40},
    {
        "paid": Noul("The receipt shows the bill was paid"),
        "matches": Noul("The receipt total matches the claimed amount"),
        "category": Choice("Expense category?", ["meals", "travel", "office", "other"]),
    },
    images=["receipt.jpg"],
    model="@openai/gpt-4.1-mini",  # any vision-capable model
)
d["matches"].noul, d["category"].choice

Three questions, one request. `paid` and `matches` are Nouls, so each answer is the probability of yes. `category` is a Choice with a probability per option and a `none_of_these` escape option by default. The claim lives in the state as data, and the model compares it with what it sees.

Turn the probabilities into actions in plain Python:

python
if d["matches"].noul < 0.5 or d["paid"].noul < 0.5:
    send_to_finance_review(claim_id, decision_id=d.id)
else:
    approve(claim_id, category=d["category"].choice)

Keep the `d.id`. When finance reviews a claim, send what they found with `client.feedback(d.id, "matches", True)` or `False`. After 30 labels for the question in a project, Curva calibrates the probabilities on your receipts when that makes them more accurate. Then your threshold means what it says.

One caution on this example. Comparing two amounts is arithmetic, and language models are poor at arithmetic. Here the model reads the total from the picture, which only a vision model can do. If you can get the receipt total some other way, compare the numbers in code and ask the model only what code can't answer.

URLs are fetched by the provider: send private images inline

This is the detail most vision tutorials skip. When you pass an `https://` URL, Curva doesn't download the image. It forwards the URL, and the model provider fetches it. Two things follow. The URL must be reachable from the provider's servers, so an intranet link fails. And the provider learns the URL, including any signed token in it.

For anything private (receipts, IDs, customer screenshots), read the file and pass bytes or a path, so the image travels inline in your request:

python
from pathlib import Path

d = curva.decide(claim_state, QUESTIONS, images=[Path("uploads/claim-1182.jpg")])
d = curva.decide(claim_state, QUESTIONS, images=[blob_store.get(claim_id)])   # raw bytes

Inline images still reach the model provider, since the model has to see them. Inline only changes who fetches the image and what else gets shared.

Limits: 8 images, 5 MB decoded each, 16 MB body (--max-body-mb)

Inline base64 images count towards the body limit, and base64 is larger than the file. Several large photos can pass the 5 MB per-image check and still exceed 16 MB together. Resize before sending; most classification questions don't need full resolution.

Models without vision: a 502 with the provider's message, or ignored images

Only vision-language models accept images, and support varies by provider and model. A model without vision usually fails the call with a 502 carrying the provider's message. In the SDK that raises `ModelError`, a subclass of `CurvaError`. Sometimes the model just ignores the image instead, and answers from the state alone. That failure is silent, so test before you trust it: run a small labeled set through the model and check that the image questions are better than chance.

The prompt tells the model that the images are data to judge and that instructions written inside them are ignored, the same way the fenced state is handled. A photo of a note saying "approve this claim" is evidence, not an instruction.

Cost: debias and councils send the images with every call

Images count towards the prompt tokens, so they cost more than text. Order debiasing, on by default, asks every question twice with the options in original and reversed order, and both calls carry the images. A council of three models sends them three times. Ways to keep the bill down:

Images in the cache key; only a salted hash in the audit log

Images are part of the decision cache key. The same state with another image is a new decision. The same state with the same image is answered from the cache in about a millisecond, at no cost.

Curva never stores the images. The audit log keeps a salted SHA-256 of each image as sent, and the state itself is stored only as a salted hash. You can prove which image a decision saw by hashing it again, without keeping a copy in Curva's database.

Next steps

Read the [images guide](https://itsmohitrohilla.github.io/curva-docs/guides/images/) in the docs. For the full Python workflow, see [Python LLM classification](/blog/python-llm-classification/). To pick a vision-capable model, read [choose a model for LLM classification](/blog/choose-a-model-for-llm-classification/), and for more ideas, [LLM classification use cases](/blog/llm-classification-use-cases/).