Skip to content

Shadow mode

Shadow mode answers one question before you switch: if Curva had made these decisions instead of my current system, how often would it agree, and where could it safely take over?

curva shadow replays logged traffic through Curva and compares its answers with those your current system gave (an LLM prompt, rules, or people).

One JSON line per past decision, with the current system’s answers in feedback format (option key for Choice, level index for Score, true/false for Noul):

{"state": {"ticket": "I was charged twice"}, "labels": {"team": "billing", "refund_requested": true}}

A line may leave out a question; it is then counted as having no current answer. If the answers live under another field, name it with --compare <field>.

Terminal window
curva recipe show support-triage > questions.json # or your own questions
curva shadow traffic.jsonl -q questions.json -o shadow.jsonl
curva shadow traffic.jsonl -q questions.json -o shadow.jsonl --cascade free-model,strong-model --examples 10

It takes the same model options as curva map (--model, --council, --cascade, --race, --mode, --no-debias). It is resumable: every decision is appended to -o as it is made, so after a quota or network failure you rerun the same command and it continues. Repeated states are answered from the cache.

The report is Markdown on stdout, built from the whole -o file.

  • Header: decisions, the models that answered, Curva’s total and per-decision cost, and p50 and p95 latency.
  • Per question:
    • agreement with the current system;
    • a table by Curva confidence band (≥ 0.9, 0.8–0.9, … < 0.6) with agreement per band, and the share of traffic and agreement if Curva decides everything at or above that band. This is the takeover view: if agreement at ≥ 0.9 is 99% on 70% of traffic, Curva can take that 70% and pass the rest to the current system;
    • the disagreements, most confident first (--examples N, default 5), with the line number and the start of the state. A confident disagreement is where one of the two systems is clearly wrong, so read these first.

Multi questions are left out of the report, since their answers are sets.

© 2026 Tarkova Private Limited.