Skip to content

Your own fast model

A large model answers your questions well, but every decision pays its latency and price. Once Curva has seen enough of your traffic, you can train a small open model (1–8B parameters) to answer your questions, serve it yourself, and point Curva at it. Curva keeps doing what it does for any model: typed answers, debiasing, calibration on your feedback, abstain and audit.

The path:

  1. Collect answers you trust: feedback labels, and a strong model’s confident answers.
  2. curva export them as chat training data, in exactly the prompt Curva sends.
  3. Fine-tune a small model with LoRA.
  4. Serve it behind an OpenAI-compatible endpoint.
  5. Point Curva at it, compare with curva bench, and calibrate with feedback.

Training rows come from two sources.

  • Feedback labels (ground truth). Every label you send with /v1/feedback is kept next to the decision it corrects. These are the best rows you have.
  • A strong model’s confident answers (distillation). Where there is no label, curva export can use the answering model’s own top label when its raw probability is at least --min-confidence (default 0.9). This is distillation: the small model learns to copy the strong model where the strong model was sure. It inherits that model’s mistakes, so check the result against labels (step 5), and prefer a strong model for this step.

Curva never stores states, so you keep them yourself: one JSON state per line, exactly as you sent it (same key order; whitespace doesn’t matter). A curva map input file already has this shape.

Log each state you send to a JSONL file. curva export finds the matching decisions in curva.db through the audit log’s salted state hash, then takes their question, raw probabilities and any feedback label.

How much? A few hundred labeled rows per question is a start; 1,000–5,000 per question is comfortable. Include the hard and rare cases, not only the easy majority.

Hold out some states first, so the test is on states the model never saw:

Terminal window
shuf states.jsonl > shuffled.jsonl
head -n -500 shuffled.jsonl > train-states.jsonl
tail -n 500 shuffled.jsonl > test-states.jsonl

Then export training rows:

Terminal window
# from the store: labels, plus the strong model's answers at ≥ 0.9
curva export train-states.jsonl --db curva.db --project support -o train.jsonl
# or from a curva map run
curva export train-states.jsonl --map strong.jsonl -q questions.json -o train.jsonl
Option Meaning
--db <FILE> The server’s database (default curva.db)
--map <FILE> -q <FILE> Read a curva map output file and its questions instead
--project <P> Only this project (default: all)
--question <KEY> Only this question (default: all)
--labels-only Only rows with a feedback label: no distillation
--min-confidence <P> Unlabeled rows need at least this raw top probability (default 0.9)
--format chat (default) {"messages": [system, user, assistant]}
--format curva {"state", "key", "question", "label", "source"}, label in feedback format

The chat format is Curva’s own logprobs prompt for one question, and the assistant reply is the answer’s label letter:

{"messages": [
{"role": "system", "content": "You are a decision function. … Reply with exactly one label letter per question …"},
{"role": "user", "content": "<state>{\"ticket\":\"I was charged twice\"}</state>\n\nQ1. Which team? Labels: A=billing, B=technical, C=none_of_these (none of the options fits)\n\nAnswer:"},
{"role": "assistant", "content": "A"}
]}

Each answer is written twice: in the question’s label order and reversed, because Curva’s order debiasing asks both ways. The fine-tuned model therefore answers Curva’s prompt with a single letter token, which is what Curva reads the probabilities from. Multi questions and questions with more than 20 labels are left out (they are not answered with one letter). Rows use Curva’s default prompt layout (the state first).

Also export the held-out states with labels only, in curva format, for step 5:

Terminal window
curva export test-states.jsonl --db curva.db --project support --labels-only --format curva -o test.jsonl

LoRA trains a small adapter on top of an open instruct model, so it fits on a laptop or one rented GPU. Good bases are 1–8B instruct models with a permissive license (for example Qwen2.5 1.5B/3B/7B Instruct, Llama 3.2 3B Instruct, Gemma 2 2B). Start small: a 1.5–3B model is often enough for a few fixed questions.

Terminal window
pip install mlx-lm
mkdir -p data
shuf train.jsonl > shuffled.jsonl
head -n -200 shuffled.jsonl > data/train.jsonl
tail -n 200 shuffled.jsonl > data/valid.jsonl
# train only on the answer letter, not the prompt
mlx_lm.lora --model Qwen/Qwen2.5-1.5B-Instruct --train --data data \
--mask-prompt --batch-size 4 --iters 1000 --num-layers 16
# merge the adapter into a standalone model
mlx_lm.fuse --model Qwen/Qwen2.5-1.5B-Instruct --adapter-path adapters \
--save-path curva-support-1.5b

One or two epochs are usually enough; the answer is a single token, so the model overfits fast. Watch the validation loss and stop when it rises.

Curva needs an OpenAI-compatible chat completions endpoint that returns logprobs with top_logprobs, since that is where the probabilities come from.

Terminal window
mlx_lm.server --model curva-support-1.5b --port 8080

Check the endpoint returns usable label probabilities before going further:

Terminal window
curva spike @local/curva-support-3b

Custom endpoints are named with @<provider>/<model> ids, where the provider’s base URL comes from CURVA_PROVIDER_<NAME>_URL. This provider support is being added alongside this guide; the exact variables are on the providers page once it ships.

Terminal window
export CURVA_PROVIDER_LOCAL_URL=http://127.0.0.1:8080/v1
curva serve --model @local/curva-support-3b

Compare with the strong model on held-out labels. Turn test.jsonl into a bench set per question:

Terminal window
mkdir -p own-evals
jq -s --arg k team '{question: (map(select(.key == $k))[0].question),
rows: map(select(.key == $k) | {state, label})}' test.jsonl > own-evals/team.json
curva bench --dir own-evals team --model @local/curva-support-3b
curva bench --dir own-evals team --model <strong-model>

Look at accuracy, ECE and latency side by side. A good result: accuracy within a point or two of the strong model, at a fraction of the latency and cost.

Calibrate it. Calibrators are kept per project and exact question, not per model. Give the new model its own project (for example support-own), so its calibrator is fit on its own answers, and keep sending feedback. From 30 labels per question, its answers come back calibrated: true whenever calibration makes them more accurate.

Keep a safety net. A cascade sends only the questions your model is unsure about to the strong model:

{"model": {"cascade": ["@local/curva-support-3b", "<strong-model>"], "escalate_below": 0.8}}

Ask one question per request if it scores better. Training rows hold one question each. Curva normally asks all of a request’s questions in one call, which is a longer prompt than the model was trained on. If a multi-question request scores worse in curva bench or shadow mode, send one question per request, or export and train on the questions you ask together.

Rough numbers for a few thousand rows of short states. Your times depend on state length, epochs and hardware.

Setup Model size Training Cost Serving
MacBook, Apple silicon, 16 GB+ (MLX) 1–3B 30 min to a few hours free fine for development and low traffic
MacBook, Apple silicon, 32 GB+ (MLX) 7–8B several hours, can be slow free slow; better on a GPU
Rented GPU, 24 GB (Unsloth, 4-bit) 1–8B 15–60 min about $1–3 per hour vLLM on the same class of GPU
Rented GPU, 80 GB 8B, 16-bit 15–30 min about $2–3 per hour high throughput with vLLM

Serving is a continuous cost: a rented GPU running all month is roughly $700–2,000. It is worth it when that is less than your current model bill, or when latency or data residency matters more than cost.

  • Questions are fixed and asked often
  • States logged exactly as sent; a test set held out by state
  • curva export with labels, plus confident strong-model answers if labels are few
  • LoRA fine-tune, validation loss watched
  • Served with logprobs + top_logprobs; curva spike passes
  • curva bench on held-out labels against the strong model
  • New project for the new model, feedback flowing, calibration checked

© 2026 Tarkova Private Limited.