Python LLM API error handling with CurvaError
Map each CurvaError subclass to the fix: 422 names the bad question, 429 carries retry_after, 502 splits model_error from model_unavailable.
Python LLM API error handling with Curva comes down to one base class and five subclasses. Every failure the client doesn't retry raises a `CurvaError` with `status`, `type` and `message`, matching the HTTP error the server sent. Catch `InvalidRequestError` to fix a bad question (a 422 names it), `RateLimitError` to wait (it carries `retry_after`), and `ModelError` for provider failures, where `model_error` means retry later and `model_unavailable` means fix the model id. A status of 0 means the server was never reached. And 429 and 5xx responses are retried for you first, so most transient failures never reach your `except` block.
Python LLM API error handling starts with one base class: CurvaError
from curva import CurvaError
try:
client.decide(state, questions)
except CurvaError as e:
print(e.status, e.type, e.message)The three attributes do different jobs. `status` is the HTTP status, or 0 when there was no response. `type` is a fixed word from the server, such as `invalid_request`, `rate_limited` or `model_unavailable`, and it is what your code should branch on. `message` is for people: it names the question or image at fault, and it is safe to log because provider error details stay in the server log and are never sent to callers.
Log `d.request_id` on success and keep it on failure where you can: it is the server's `x-request-id`, the same id that appears in the server's log line for the request.
The five subclasses: AuthError, InvalidRequestError, NotFoundError, RateLimitError, ModelError
Statuses without a subclass, such as 403 for a key bound to another project or 413 for a body over the server's limit, raise the base `CurvaError`. Catch subclasses first and the base class last:
from curva import CurvaError, RateLimitError, InvalidRequestError, ModelError
try:
d = client.decide(state, QUESTIONS, project="support")
except InvalidRequestError as e:
log.error("bad question: %s", e.message) # don't retry: fix the question
raise
except RateLimitError as e:
requeue(item, delay=e.retry_after)
except ModelError as e:
if e.type == "model_unavailable":
alert("model id or access is wrong: %s" % e.message)
requeue(item, delay=60)
except CurvaError as e:
log.error("curva %s %s: %s", e.status, e.type, e.message)
raiseStatus 0: curva.local() not started or a wrong CURVA_BASE_URL
A `CurvaError` with status 0 means the server was unreachable. No HTTP response came back at all. The usual causes:
One related error is raised before any request: with no provider key at all and no `CURVA_MODEL`, the module-level `curva.decide` raises an error naming the variables to set.
Dropped connections are retried along with 429 and 5xx, so a status 0 that reaches you has already failed `max_retries` times.
422 invalid_request names the question: option counts, levels, mode conflicts
A 422 means the request was valid JSON but something in it is not allowed. The message names the question or image, so log it in full. Common causes in Python code:
A 400 `invalid_json` is the other `InvalidRequestError`: the body was not JSON, or named an unknown `mode` or question `type`. Neither should be retried. The same request will fail the same way.
502 model_error vs model_unavailable: retry later or fix the model id
Both arrive as `ModelError` with status 502, and `type` tells them apart.
A 504 `timeout` is also a `ModelError`: the request took longer than the 120 s deadline.
404 on feedback: unknown decision or question key
`client.feedback(decision_id, key, label)` raises `NotFoundError` for an unknown decision or an unknown question key. Two cases surprise people because the decision and key both exist:
Check the answer's `rule` field and the skipped flag before sending feedback. A label that isn't a valid answer, such as a key not among the options, gets 422 and arrives as `InvalidRequestError`.
What never reaches your except block because the client retried it
The client retries before it raises:
from curva import Curva client = Curva(timeout=30.0, max_retries=3) # retries 429, 5xx and dropped connections
By default, statuses 429, 500, 502, 503 and 504 and dropped connections are retried up to `max_retries` times, honouring `Retry-After`. Only what is still failing after that raises. So a `RateLimitError` in your logs means a limit that lasted through every retry: a key's requests-per-minute limit, the provider's rate limit, or the server's daily budget, which resets at 00:00 UTC. Lower your concurrency or raise the limit rather than adding another retry loop on top.
For batches, `AsyncCurva.decide_many(..., return_exceptions=True)` puts each failure in its slot instead of raising the first one. Check each result with `isinstance(result, CurvaError)` and requeue only the failed rows.
Next steps
The docs cover the classes in the [Python SDK reference](https://itsmohitrohilla.github.io/curva-docs/reference/python-sdk/) and the statuses in the [HTTP API reference](https://itsmohitrohilla.github.io/curva-docs/reference/http-api/). For every status across all clients, read the [Curva troubleshooting guide](/blog/curva-troubleshooting-guide/). Tune retries in [LLM client retries and timeouts in Python](/blog/llm-client-retries-timeouts-python/), and handle batches in [async LLM classification in Python](/blog/async-llm-classification-python/). Start from the [Python LLM classification tutorial](/blog/python-llm-classification/).