Skip to content

Calibrate with feedback

Calibration is what turns a model’s raw confidence into a probability you can act on. It needs one thing from you: the true answer, whenever you learn it.

Every decision has an id (dec_…). Store it next to whatever you did with the answer.

d = client.decide(ticket, questions, project="support")
save(ticket_id, d.id)

When an agent closes the ticket, a reviewer corrects a label, or the outcome is otherwise known:

client.feedback(decision_id, "team", "billing")
Question type Label
Choice the option key, e.g. "billing"
Score the level index, e.g. 2
Noul True or False

Sending feedback again for the same decision and question replaces the earlier label. An unknown decision or question gets 404, and a label that isn’t a valid answer gets 422.

Over HTTP:

Terminal window
curl -s localhost:7777/v1/feedback -H 'content-type: application/json' \
-d '{"decision_id": "dec_…", "question": "team", "label": "billing"}'
# {"decision_id": "dec_…", "question": "team", "labels": 31, "min_labels": 30, "calibrated": true}

From 30 labels for the same exact question in a project, Curva fits a calibrator and refits it as labels arrive. When it makes the answers more accurate on labels it hasn’t seen, that question’s answers come back with calibrated=True; otherwise they stay raw (calibrated=False), because the model is already well calibrated. See Calibration.

report = client.calibration("team", project="support")
print(report["before"]["ece"], "→", report["after"]["ece"])
print(report["after"]["accuracy_when_automated"], report["after"]["automated"])

before is the raw model; after is held out (each half of the labels calibrated by a fit on the other half). Both include accuracy, ece, brier, automated and accuracy_when_automated (at 0.9 confidence), plus reliability bins for a calibration curve. calibrator shows the fitted parameters.

A common pattern is to automate confident answers and learn from the rest:

q = {"team": Choice("Which team?", {"billing": "", "technical": "", "sales": ""}, min_confidence=0.9)}
d = client.decide(ticket, q, project="support")
if d["team"].abstain:
team = ask_a_human(ticket) # a person decides...
client.feedback(d.id, "team", team) # ...and Curva learns from it
else:
route(ticket, d["team"].choice)

Also send feedback for a sample of automated answers, so the calibration covers the whole confidence range, not only the unsure cases.

© 2026 Tarkova Private Limited.