Skip to content

Council, cascade and race

One request can use several models. Pick a plan by what you want to optimise.

Plan Models are asked Good for Answers get
Fallback chain one at a time, until one answers surviving outages and rate limits -
Council all at once, blended fewer confident mistakes agreement
Cascade cheapest first, escalating only unsure questions cost answered_by
Race all at once, first valid answer wins latency -

Four ways to use several models. A fallback chain asks one model at a time until one answers. A council asks all models at once and blends them by geometric mean. A cascade asks the cheapest model first and escalates only unsure questions. A race asks all models at once and the first valid answer wins.

client.decide(state, questions, model=["model-a", "model-b"])

If a model returns 429 or a server error, the next one takes over. The response’s model says which one answered.

d = client.decide(state, questions, council=["model-a", "model-b", "model-c"])
d["team"].agreement # share of members whose top answer matches the council's

2 to 5 models answer concurrently, and their probabilities are blended by geometric mean. Where members disagree, confidence drops, which is exactly what you want: disagreement is a signal to look closer. A member that fails is left out.

HTTP: "model": {"council": ["model-a", "model-b"]}.

d = client.decide(state, questions, cascade=["free-model", "strong-model"], escalate_below=0.8)
d["team"].answered_by # the model that gave the final answer

The cheapest model answers first. Only the questions whose confidence is below escalate_below (default 0.8) go to the next model; confident answers stand. If the stronger model fails, the cheaper answers are kept. When the cheap model is sure, the strong model is never called.

HTTP: "model": {"cascade": ["free-model", "strong-model"], "escalate_below": 0.8}.

client.decide(state, questions, race=["model-a", "model-b"])

All models are asked at once, the first valid answer wins, and the others are cancelled. This cuts slow outliers on shared providers, at the cost of one call per model.

HTTP: "model": {"race": ["model-a", "model-b"]}.

curva bench and curva map take the same plans:

Terminal window
curva bench --council model-a,model-b
curva map in.jsonl -q questions.json -o out.jsonl --cascade free-model,strong-model --escalate-below 0.8
curva bench --race model-a,model-b

Use only one of council, cascade or race per request.

© 2026 Tarkova Private Limited.