Your own fast model
A large model answers your questions well, but every decision pays its latency and price. Once Curva has seen enough of your traffic, you can train a small open model (1–8B parameters) to answer your questions, serve it yourself, and point Curva at it. Curva keeps doing what it does for any model: typed answers, debiasing, calibration on your feedback, abstain and audit.
The path:
- Collect answers you trust: feedback labels, and a strong model’s confident answers.
curva exportthem as chat training data, in exactly the prompt Curva sends.- Fine-tune a small model with LoRA.
- Serve it behind an OpenAI-compatible endpoint.
- Point Curva at it, compare with
curva bench, and calibrate with feedback.
1. Collect data
Section titled “1. Collect data”Training rows come from two sources.
- Feedback labels (ground truth). Every label you send with
/v1/feedbackis kept next to the decision it corrects. These are the best rows you have. - A strong model’s confident answers (distillation). Where there is no label,
curva exportcan use the answering model’s own top label when its raw probability is at least--min-confidence(default 0.9). This is distillation: the small model learns to copy the strong model where the strong model was sure. It inherits that model’s mistakes, so check the result against labels (step 5), and prefer a strong model for this step.
Curva never stores states, so you keep them yourself: one JSON state per line, exactly as you
sent it (same key order; whitespace doesn’t matter). A curva map input file already has this
shape.
Log each state you send to a JSONL file. curva export finds the matching decisions in
curva.db through the audit log’s salted state hash, then takes their question, raw
probabilities and any feedback label.
Run your states through a strong model with curva map, then export from its
output file:
curva map states.jsonl -q questions.json -o strong.jsonl --model <strong-model>A map run has no feedback labels, so this is distillation only.
How much? A few hundred labeled rows per question is a start; 1,000–5,000 per question is comfortable. Include the hard and rare cases, not only the easy majority.
2. Export
Section titled “2. Export”Hold out some states first, so the test is on states the model never saw:
shuf states.jsonl > shuffled.jsonlhead -n -500 shuffled.jsonl > train-states.jsonltail -n 500 shuffled.jsonl > test-states.jsonlThen export training rows:
# from the store: labels, plus the strong model's answers at ≥ 0.9curva export train-states.jsonl --db curva.db --project support -o train.jsonl
# or from a curva map runcurva export train-states.jsonl --map strong.jsonl -q questions.json -o train.jsonl| Option | Meaning |
|---|---|
--db <FILE> |
The server’s database (default curva.db) |
--map <FILE> -q <FILE> |
Read a curva map output file and its questions instead |
--project <P> |
Only this project (default: all) |
--question <KEY> |
Only this question (default: all) |
--labels-only |
Only rows with a feedback label: no distillation |
--min-confidence <P> |
Unlabeled rows need at least this raw top probability (default 0.9) |
--format chat |
(default) {"messages": [system, user, assistant]} |
--format curva |
{"state", "key", "question", "label", "source"}, label in feedback format |
The chat format is Curva’s own logprobs prompt for one question, and the assistant reply is the
answer’s label letter:
{"messages": [ {"role": "system", "content": "You are a decision function. … Reply with exactly one label letter per question …"}, {"role": "user", "content": "<state>{\"ticket\":\"I was charged twice\"}</state>\n\nQ1. Which team? Labels: A=billing, B=technical, C=none_of_these (none of the options fits)\n\nAnswer:"}, {"role": "assistant", "content": "A"}]}Each answer is written twice: in the question’s label order and reversed, because Curva’s order debiasing asks both ways. The fine-tuned model therefore answers Curva’s prompt with a single letter token, which is what Curva reads the probabilities from. Multi questions and questions with more than 20 labels are left out (they are not answered with one letter). Rows use Curva’s default prompt layout (the state first).
Also export the held-out states with labels only, in curva format, for step 5:
curva export test-states.jsonl --db curva.db --project support --labels-only --format curva -o test.jsonl3. Fine-tune with LoRA
Section titled “3. Fine-tune with LoRA”LoRA trains a small adapter on top of an open instruct model, so it fits on a laptop or one rented GPU. Good bases are 1–8B instruct models with a permissive license (for example Qwen2.5 1.5B/3B/7B Instruct, Llama 3.2 3B Instruct, Gemma 2 2B). Start small: a 1.5–3B model is often enough for a few fixed questions.
pip install mlx-lmmkdir -p datashuf train.jsonl > shuffled.jsonlhead -n -200 shuffled.jsonl > data/train.jsonltail -n 200 shuffled.jsonl > data/valid.jsonl
# train only on the answer letter, not the promptmlx_lm.lora --model Qwen/Qwen2.5-1.5B-Instruct --train --data data \ --mask-prompt --batch-size 4 --iters 1000 --num-layers 16
# merge the adapter into a standalone modelmlx_lm.fuse --model Qwen/Qwen2.5-1.5B-Instruct --adapter-path adapters \ --save-path curva-support-1.5b# pip install unslothfrom unsloth import FastLanguageModelfrom unsloth.chat_templates import train_on_responses_onlyfrom datasets import load_datasetfrom trl import SFTTrainer, SFTConfig
model, tokenizer = FastLanguageModel.from_pretrained("unsloth/Qwen2.5-3B-Instruct", max_seq_length=4096, load_in_4bit=True)model = FastLanguageModel.get_peft_model(model, r=16, lora_alpha=16, target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"])
data = load_dataset("json", data_files="train.jsonl", split="train")data = data.map(lambda row: {"text": tokenizer.apply_chat_template(row["messages"], tokenize=False)})
trainer = SFTTrainer(model=model, tokenizer=tokenizer, train_dataset=data, args=SFTConfig(dataset_text_field="text", per_device_train_batch_size=8, num_train_epochs=2, learning_rate=2e-4, output_dir="out"))# loss on the answer letter only (the markers depend on the base model's chat template)trainer = train_on_responses_only(trainer, instruction_part="<|im_start|>user\n", response_part="<|im_start|>assistant\n")trainer.train()model.save_pretrained_merged("curva-support-3b", tokenizer, save_method="merged_16bit")One or two epochs are usually enough; the answer is a single token, so the model overfits fast. Watch the validation loss and stop when it rises.
4. Serve it
Section titled “4. Serve it”Curva needs an OpenAI-compatible chat completions endpoint that returns logprobs with
top_logprobs, since that is where the probabilities come from.
mlx_lm.server --model curva-support-1.5b --port 8080vllm serve ./curva-support-3b --served-model-name curva-support-3b --port 8080 --max-logprobs 20# convert and quantize first (convert_hf_to_gguf.py, llama-quantize)llama-server -m curva-support-3b-Q4_K_M.gguf --port 8080# Modelfile: FROM ./curva-support-3b-Q4_K_M.ggufollama create curva-support-3b -f ModelfileCheck that your Ollama version returns logprobs on its OpenAI-compatible endpoint.
Check the endpoint returns usable label probabilities before going further:
curva spike @local/curva-support-3b5. Point Curva at it
Section titled “5. Point Curva at it”Custom endpoints are named with @<provider>/<model> ids, where the provider’s base URL comes
from CURVA_PROVIDER_<NAME>_URL. This provider support is being added alongside this guide; the
exact variables are on the providers page once it ships.
export CURVA_PROVIDER_LOCAL_URL=http://127.0.0.1:8080/v1curva serve --model @local/curva-support-3bCompare with the strong model on held-out labels. Turn test.jsonl into a
bench set per question:
mkdir -p own-evalsjq -s --arg k team '{question: (map(select(.key == $k))[0].question), rows: map(select(.key == $k) | {state, label})}' test.jsonl > own-evals/team.json
curva bench --dir own-evals team --model @local/curva-support-3bcurva bench --dir own-evals team --model <strong-model>Look at accuracy, ECE and latency side by side. A good result: accuracy within a point or two of the strong model, at a fraction of the latency and cost.
Calibrate it. Calibrators are kept per project and exact question, not per model. Give the
new model its own project (for example support-own), so its calibrator is fit on its own
answers, and keep sending feedback. From 30 labels per question, its answers come
back calibrated: true whenever calibration makes them more accurate.
Keep a safety net. A cascade sends only the questions your model is unsure about to the strong model:
{"model": {"cascade": ["@local/curva-support-3b", "<strong-model>"], "escalate_below": 0.8}}Ask one question per request if it scores better. Training rows hold one question each. Curva normally asks all of
a request’s questions in one call, which is a longer prompt than the model was trained on. If a
multi-question request scores worse in curva bench or shadow mode, send one question per
request, or export and train on the questions you ask together.
Cost and time
Section titled “Cost and time”Rough numbers for a few thousand rows of short states. Your times depend on state length, epochs and hardware.
| Setup | Model size | Training | Cost | Serving |
|---|---|---|---|---|
| MacBook, Apple silicon, 16 GB+ (MLX) | 1–3B | 30 min to a few hours | free | fine for development and low traffic |
| MacBook, Apple silicon, 32 GB+ (MLX) | 7–8B | several hours, can be slow | free | slow; better on a GPU |
| Rented GPU, 24 GB (Unsloth, 4-bit) | 1–8B | 15–60 min | about $1–3 per hour | vLLM on the same class of GPU |
| Rented GPU, 80 GB | 8B, 16-bit | 15–30 min | about $2–3 per hour | high throughput with vLLM |
Serving is a continuous cost: a rented GPU running all month is roughly $700–2,000. It is worth it when that is less than your current model bill, or when latency or data residency matters more than cost.
Checklist
Section titled “Checklist”- Questions are fixed and asked often
- States logged exactly as sent; a test set held out by state
-
curva exportwith labels, plus confident strong-model answers if labels are few - LoRA fine-tune, validation loss watched
- Served with
logprobs+top_logprobs;curva spikepasses -
curva benchon held-out labels against the strong model - New project for the new model, feedback flowing, calibration checked

