Skip to content

Speed and cost

Every decision is one model call per option order, per model. These settings cut calls, wait time or price, roughly from least to most effort.

Setting Saves Costs you
debias="auto" about half the calls, once a question is learned nothing, for questions without position bias
Decision cache every repeat of an identical decision memory on the server
config="curva-1.2.0" prompt tokens, where the provider caches prompts nothing
Cascade calls to the expensive model an extra call for unsure questions
Race waiting on slow providers one call per model
Local or fast providers network time and per-call price running or choosing the model

By default every question is asked twice, with the options in original and reversed order, and the answers are averaged (why). Many models answer the same either way for many questions, and then the second call buys nothing.

d = client.decide(state, questions, debias="auto")
d.debiased # True while learning or checking, False when only the original order was asked

With debias: "auto", the server keeps asking both orders for a (model, question) until it has seen at least 20 paired answers of which 95% agree (same top answer, top probability within 0.1). From then on it asks only the original order, and still asks both on every 10th request to keep checking. One disagreement puts the question back to full debiasing.

What it learned lives in the server’s memory (up to 10,000 model-question pairs) and starts over after a restart. A reworded question is a new question. debias: false skips the second call always, and keeps whatever position bias the model has.

HTTP: "debias": "auto".

Calls run at temperature 0, so an identical decision (same model, state, questions, config, project, privacy and think) is answered from memory in about a millisecond, with cached: true and no cost. Size it with curva serve --cache-size. Nothing to do on the client.

config="curva-1.2.0" puts the questions before the state. The questions repeat on every call, so providers and local servers that cache prompt prefixes (llama.cpp, vLLM, most hosted APIs) read them from cache and charge less for them. Answers can differ slightly from curva-1.1.0, so calibrate again after switching.

client.decide(state, questions, cascade=["cheap-model", "strong-model"], escalate_below=0.8)
client.decide(state, questions, race=["model-a", "model-b"])

A cascade asks the cheap model first and sends only the questions it is unsure of to the next one. A race asks every model at once and takes the first valid answer, which cuts the slow outliers of shared providers. See council, cascade and race.

A model on your own machine (@ollama/..., @llamacpp/...) has no network hop, no per-call price and no rate limit. Hosted providers built for low latency, such as @groq/..., are the other route. See model providers, and your own fast model to fine-tune a small model on your labels.

Keep mode at auto (or logprobs where the model returns them): a logprobs answer is one token per question, the cheapest reply there is.

© 2026 Tarkova Private Limited.