Skip to content

Images

Some decisions depend on a picture: is this receipt paid, does this screenshot show an error, is this upload safe to publish. Send up to 8 images with the state and a vision-language model judges them together with it. The answers are typed and calibrated like any other.

from curva import Curva, Choice, Noul
curva = Curva()
d = curva.decide(
{"expense_claim": "Team lunch, 4 people", "amount": 86.40},
{
"paid": Noul("The receipt shows the bill was paid"),
"matches": Noul("The receipt total matches the claimed amount"),
"category": Choice("Expense category?", ["meals", "travel", "office", "other"]),
},
images=["receipt.jpg"],
model="@openai/gpt-4.1-mini", # any vision-capable model
)
d["matches"].noul, d["category"].choice

images takes file paths (str or Path, .png, .jpg, .jpeg, .webp, .gif), raw bytes (PNG, JPEG, WebP or GIF, recognised by their first bytes), data:image/...;base64, URIs and https:// URLs. Files and bytes are sent inline as base64.

{
"state": {"ticket": "The app shows this when I pay"},
"images": ["https://files.example.com/screenshot-4411.png"],
"questions": {
"error_visible": {"type": "noul", "instructions": "The screenshot shows an error message"},
"area": {"type": "choice", "instructions": "Which part of the app?", "options": {"checkout": "", "login": "", "settings": ""}}
}
}

Each image is an https:// URL (at most 2,048 characters) or a data:image/<png|jpeg|webp|gif>;base64,... URI (at most 5 MB decoded). Anything else, or more than 8, gets a 422 that names the image. The whole request body must fit the server’s limit: 16 MB by default, curva serve --max-body-mb to change it.

URLs are fetched by the model provider, not by Curva. The URL must be reachable from the provider, and it learns the URL. For private images send them inline instead.

verdict = curva.decide(
{"caption": post.caption, "account_age_days": 3},
{"policy": Choice("Does the image break the content policy?",
{"ok": "nothing wrong", "nudity": "", "violence": "", "spam": "ads or scams"},
min_confidence=0.85)},
images=[post.image_bytes],
)
if verdict["policy"].abstain:
send_to_human_review(post)

With feedback, the confidence becomes calibrated to your own moderators’ labels, so min_confidence means what it says.

Only vision-language models accept images, and support varies by provider and model: check your provider’s model list. A model without vision support usually fails the call (a 502 with the provider’s message) or ignores the image. Try a few with Curva Tune or a council.

  • The images go to the model as OpenAI image_url content parts after the prompt text. The text starts with a line saying the images are data to judge, and that any instructions written inside them are ignored, like the fenced state.
  • Without images the request to the model is exactly what it was before.
  • Images are part of the cache key: the same state with another image is a new decision.
  • The audit log keeps a SHA-256 of each image as sent, never the image.
  • explain still attributes state fields only; the images stay in every call.
  • Images count towards the prompt tokens, so they cost more than text; debias and councils send them with every call.

© 2026 Tarkova Private Limited.