Skip to content

Guaranteed accuracy

A calibrated probability tells you how often an answer is right on average. Sometimes you need a promise instead: “the right answer is in here at least 95% of the time”. Set coverage on a question and, once it has enough feedback, every answer comes with a prediction set that carries that guarantee.

from curva import Curva, Choice
curva = Curva()
d = curva.decide(ticket, {"team": Choice("Which team?", ["billing", "technical", "sales"], coverage=0.95)},
project="support")
d["team"].choice # "billing": the usual top answer, probabilities and confidence are all still there
d["team"].guaranteed # True once the question has 30 labels
d["team"].set # ["billing"], or ["billing", "technical"] when the model is torn
import { Curva, choice } from "curva-ai";
const d = await new Curva().decide(ticket, { team: choice("Which team?", ["billing", "technical", "sales"], { coverage: 0.95 }) });
d.answers.team.set; // string[] | undefined
Set Meaning What to do
One option The model is sure enough to meet the guarantee alone Automate
Several options The truth is one of these, at the promised rate Show the options, or route to a person
Every option The model can’t narrow it down for this input Route to a person

coverage works on choice, multi and noul questions. A noul’s set is a subset of ["true", "false"]. Any other type, or a value outside (0, 1), gets 422.

Curva uses split conformal prediction, with the question’s own feedback labels as the calibration set.

  1. For every labeled answer, the nonconformity score is 1 − p(true answer). The probabilities are the ones /v1/decide returns: calibrated once the question has a calibrator, raw before.
  2. With n scores and coverage c, the threshold q̂ is the ⌈(n+1)·c⌉-th smallest score. When that rank is past n (few labels and a high c), q̂ is 1 and the set is every option.
  3. A new answer’s set is every option with p ≥ 1 − q̂. The set is only empty when q̂ is 0; the top answer is returned then, so there is always something to act on.

As long as new inputs look like the labeled ones (the labels are a fair sample of your traffic), the set contains the true answer with probability at least c. This holds whatever the model is and however badly calibrated it is. A worse model gets bigger sets, not a broken promise.

The guarantee needs at least 30 labels for this exact question in this project, the same minimum as calibration. Before that, the answer has guaranteed: false and no set. It is still a normal answer, not an error. Send feedback to build up the labels.

Rewording the question starts over: its labels, calibrator and guarantee all belong to the exact wording. Changing only coverage does not. It isn’t part of the question’s fingerprint or of the decision cache key, because the set is worked out from the stored labels after the model has answered. Asking the same question at 0.8 and 0.95 costs one model call and returns two sets.

The two answer different questions and can be used together:

  • abstain looks at the top answer: is its probability above min_confidence?
  • set looks at all the options: which of them must you keep to be right coverage of the time?

A common routing rule is to automate when the set has one option and abstain is false, and to send everything else to review.

For a multi, the set is the options whose own P(selected) is at least 1 − q̂. Feedback currently labels a multi as a whole, not option by option, so there is no per-option calibration set yet. A multi with coverage returns guaranteed: false until per-option labels exist.

© 2026 Tarkova Private Limited.