Skip to content

Drift

Models and inputs change quietly: a provider updates a model, a new product launches, a spam wave arrives. The drift report shows, week by week, how one question’s answers are distributed and how confident they are, and flags the latest week when it moved.

report = client.drift("team", project="support", weeks=8)
Terminal window
curl -s "localhost:7777/v1/drift?question=team&project=support&weeks=8" \
-H "Authorization: Bearer $CURVA_API_KEY"
{
"project": "support",
"question": "team",
"weeks": [
{"week": "2026-W38", "decisions": 412, "mix": {"billing": 0.52, "technical": 0.41, "sales": 0.07}, "confidence": 0.91},
{"week": "2026-W39", "decisions": 388, "mix": {"billing": 0.21, "technical": 0.72, "sales": 0.07}, "confidence": 0.84,
"mix_change": 0.31, "confidence_change": -0.07, "drift": true}
]
}

(Illustrative numbers.)

Field Meaning
week ISO week, oldest first. Weeks without decisions are left out
decisions Decisions that answered the question that week
mix Share of decisions per top answer (option key, level text, or yes/no)
confidence Average probability of the top answer
mix_change Latest week only: how far its mix moved from the average of the earlier weeks (total-variation distance, 0 to 1)
confidence_change Latest week only: change in average confidence against the earlier weeks
drift Latest week only: true when mix_change is over 0.2 or confidence moved by more than 0.1
  • weeks counts the current week and defaults to 8 (at most 53).
  • project defaults to "default".
  • Both mix and confidence use raw, uncalibrated probabilities, so a newly fitted calibrator doesn’t look like drift.
  • The report reads the audit log, so it covers every decision the server made for that question.

A flag is a prompt to look, not a verdict. A change in answer mix can be real (a billing outage really does bring more billing tickets). A drop in confidence more often means the inputs changed in a way the question doesn’t cover. Useful next steps:

  • Read recent decisions for the question in the audit log.
  • Label a sample and send it as feedback, then compare calibration before and after.
  • Poll the endpoint on a schedule and alert when the latest week has "drift": true.

© 2026 Tarkova Private Limited.