Curva Tune
Curva Tune answers: which model, mode and prompt setup is best on my labeled data? It searches a small grid of setups on one labeled set and writes the winner as a ready-to-send request.
Run it
Section titled “Run it”curva tune coin --per-set 40 # evals/coin.json, default modelcurva tune --dir my-evals tickets --models model-a,model-b --per-set 60 --out tuned.json --yesThe set is a JSON file in the eval format, <dir>/<set>.json:
{"question": {"type": "choice", "instructions": "Which team?", "options": {"billing": "", "technical": ""}}, "rows": [{"state": {"ticket": "Refund please"}, "label": "billing"}]}What it searches
Section titled “What it searches”| Knob | Values |
|---|---|
| Model plan | each --models model alone; with 2+ models also a council of all of them, and a cascade at escalate_below 0.7, 0.8 and 0.9 |
mode |
auto, verbal |
debias |
on, off |
| Few-shot | 0 or 4 examples (one per label first), drawn from the train rows |
That is 8 candidates per plan: 8 for one model, 48 for two. Before calling anything, it prints the
rows, the candidates and an upper bound on model calls, and refuses above 200 calls without
--yes. Lower --per-set or use fewer models to shrink it. Repeated calls (the same model,
mode and debias on the same row) are cached, so the real count is lower.
How it scores
Section titled “How it scores”- The rows (after
--per-set, spread evenly) are shuffled with a fixed seed and split once: a fifth (at least 4, at most half) is the few-shot pool, and the rest is the test set. Every candidate sees the same test rows, and reruns split the same way. - Each candidate answers every test row. The table shows accuracy, ECE, calibrated ECE (held out: each half of the test rows calibrated by a fit on the other half, when n ≥ 30), cost and failures.
- Candidates are ranked by accuracy, then ECE, then cost. The untuned default (first model,
auto, debias on, no examples) is marked, and the winner’s gain over it is printed.
Use the result
Section titled “Use the result”--out (default tuned.json) holds the winner as /v1/decide fields: model, mode, debias
and questions (the tuned examples live inside the question):
{"model": {"cascade": ["model-a", "model-b"], "escalate_below": 0.8}, "mode": "verbal", "debias": true, "questions": {"tickets": {"type": "choice", "instructions": "...", "options": {"...": ""}, "examples": []}}}Add a state and send it:
jq --arg t "The app crashes on launch" '. + {state: {ticket: $t}}' tuned.json \ | curl -s localhost:7777/v1/decide -H 'content-type: application/json' -d @-Or use the tuned questions for a batch:
jq .questions tuned.json > q.json && curva map in.jsonl -q q.json -o out.jsonl
