Shadow mode
Shadow mode answers one question before you switch: if Curva had made these decisions instead of my current system, how often would it agree, and where could it safely take over?
curva shadow replays logged traffic through Curva and compares its answers with those your
current system gave (an LLM prompt, rules, or people).
1. Log your traffic
Section titled “1. Log your traffic”One JSON line per past decision, with the current system’s answers in
feedback format (option key for Choice, level index for
Score, true/false for Noul):
{"state": {"ticket": "I was charged twice"}, "labels": {"team": "billing", "refund_requested": true}}A line may leave out a question; it is then counted as having no current answer. If the answers
live under another field, name it with --compare <field>.
2. Replay it
Section titled “2. Replay it”curva recipe show support-triage > questions.json # or your own questionscurva shadow traffic.jsonl -q questions.json -o shadow.jsonlcurva shadow traffic.jsonl -q questions.json -o shadow.jsonl --cascade free-model,strong-model --examples 10It takes the same model options as curva map (--model, --council, --cascade,
--race, --mode, --no-debias). It is resumable: every decision is appended to -o as it
is made, so after a quota or network failure you rerun the same command and it continues.
Repeated states are answered from the cache.
3. Read the report
Section titled “3. Read the report”The report is Markdown on stdout, built from the whole -o file.
- Header: decisions, the models that answered, Curva’s total and per-decision cost, and p50 and p95 latency.
- Per question:
- agreement with the current system;
- a table by Curva confidence band (
≥ 0.9,0.8–0.9, …< 0.6) with agreement per band, and the share of traffic and agreement if Curva decides everything at or above that band. This is the takeover view: if agreement at≥ 0.9is 99% on 70% of traffic, Curva can take that 70% and pass the rest to the current system; - the disagreements, most confident first (
--examples N, default 5), with the line number and the start of the state. A confident disagreement is where one of the two systems is clearly wrong, so read these first.
Multi questions are left out of the report, since their answers are sets.

