Skip to content

Benchmarks

Curva is measured on well-known public datasets and on probes that test whether its probabilities can be trusted. Every number here comes from curva bench --save, with the sample size and a 95% range next to it. The datasets are public, and curva bench ships in the curva program you install, so you can measure any model on your own labeled data (see Measure your own model).

8 public sets, 6,040 rows sampled with a fixed seed. Every run and every model sees the same rows, and each sample takes the labels in turn, so a small sample is balanced across labels.

Set Rows Question What it tests
phishnchips 2,000 Noul: phishing? Security triage, yes/no calibration
banking77 770 (10 per intent) Choice, 77 options Large option sets (intent routing)
boolq 500 Noul Reading comprehension
openbookqa 500 Choice A–D Calibration on unseen tasks
commonsenseqa 500 Choice A–E Calibration on unseen tasks
hellaswag 500 Choice A–D Calibration on unseen tasks
yelp (Yelp Review Full) 500 Score, 1–5 stars Ordinal scores
aita (Reddit AITA) 770 Choice YTA/NTA/ESH/NAH Subjective judgement, Brier score
Probe What must hold Curva mechanism
Negation consistency P(x) + P(not x) = 1 Noul is asked as a normalised two-option choice
Option-order stability Reordering the options doesn’t change the answer Order debiasing averages both orders
Fair coin P(heads) ≈ 0.5 Calibration on feedback

Nothing is claimed until curva bench --save has produced it.

Metric Target
Negation sum P(x) + P(not x) within 0.02 of 1.00
Option-order swap < 5% of answers change
Fair coin P(heads) 0.50 ± 0.05
ECE after calibration, each public set ≤ 0.03
Overall accuracy (built-in sets) ≥ 90%
Accuracy when automated at ≥ 0.9 confidence ≥ 98%, on ≥ 70% of decisions
Adversarial (injected) inputs ≥ 95%, never confidently wrong
Repeat decision 0 ms, $0 (cache)

Results (30 September 2026, calibration updated 1 October)

Section titled “Results (30 September 2026, calibration updated 1 October)”

Free models, verbal mode, order debiasing on. Accuracy is weighted to each dataset’s natural label mix (samples are balanced, published accuracies are not). ECE after calibration is held-out: each half of the rows is calibrated by a fit on the other half, which is what you get once you send feedback; it equals the raw ECE when calibration would not help.

Set n Accuracy (95% range) ECE ECE after calibration p50 latency
PhishNChips (phishing) 100 80.8% (73%–89%) 0.138 0.138 1146 ms
BANKING77 (77 intents) 120 79.8% (73%–87%) 0.097 0.097 1362 ms
BoolQ 127 88.9% (83%–94%) 0.081 0.081 1071 ms
OpenBookQA 100 92.2% (87%–97%) 0.056 0.056 1013 ms
CommonsenseQA 100 82.0% (74%–90%) 0.118 0.118 1263 ms
HellaSwag 98 78.6% (70%–87%) 0.032 0.032 1173 ms
Yelp (1–5 stars) 100 54.0% (44%–64%) 0.259 0.138 2832 ms
Reddit AITA (4 verdicts) 100 67.2% (58%–76%) 0.404 0.149 1318 ms
Set n Accuracy (95% range) ECE ECE after calibration p50 latency
PhishNChips (phishing) 100 72.8% (64%–82%) 0.283 0.209 178 ms
BANKING77 (77 intents) 69 74.7% (64%–85%) 0.252 0.157 351 ms
OpenBookQA 100 85.0% (78%–92%) 0.051 0.051 207 ms
CommonsenseQA 100 84.0% (77%–91%) 0.066 0.066 230 ms
HellaSwag 100 86.0% (79%–93%) 0.085 0.085 209 ms
Reddit AITA (4 verdicts) 100 55.7% (46%–65%) 0.475 0.221 265 ms

What this shows:

  • Calibration fixes models that lean or overclaim. Gemini’s AITA error falls from 0.404 to 0.149 and its Yelp error from 0.259 to 0.138 once bias scaling corrects its favourite answers (Yelp accuracy also rises from 54% to 59%); Groq’s overconfident phishing probabilities go from 0.283 to 0.209.
  • It never makes good answers worse. Where a model is already well calibrated (the quiz sets), Curva leaves its probabilities as they are.
  • Speed. Groq answers in about 0.2 s; Gemini Flash-Lite in about 1.1 s.
  • Both run free within the providers’ free tiers.

gpt-4.1-nano, gpt-4o-mini, gpt-4.1-mini and Claude Haiku 4.5 ran 20 balanced rows on each set, too few for firm per-set numbers. Measured cost per 1,000 decisions across the 8 sets: gpt-4.1-nano $0.078, gpt-4o-mini $0.118, gpt-4.1-mini $0.314, Claude Haiku 4.5 $1.64.

Trust probes (live server, auth and debiasing on)

Section titled “Trust probes (live server, auth and debiasing on)”
Probe Ling 3.0 Flash (logprobs) Nemotron 3 Super (verbal)
Negation: P(x) + P(not x), 3 tickets 1.33 / 0.14 / 0.56 (fails: the model ignores “NOT”) 1.000 / 1.000 / 1.000
Option-order swap, 4 tickets × 2 orders 0 of 4 change not run
Fair coin P(heads) 0.494 not run

Because Ling failed the negation probe and Nemotron passed it, the default config curva-1.1.0 (curva-latest) uses Nemotron 3 Super. curva-1.0.0 keeps Ling for callers who pinned it.

Put your labeled rows in a folder as <name>.json, with the question and one row per example:

{"question": {"type": "noul", "instructions": "The customer asks for a refund"},
"rows": [{"state": {"ticket": "Charged twice, refund please"}, "label": true}]}

Then run any model on it and re-analyse without new calls:

Terminal window
pip install curva-ai
curva bench --dir my-evals --model @gemini/gemini-flash-lite-latest --per-set 100 --save run.jsonl refunds
curva report run.jsonl --dir my-evals

A rerun with the same --save file skips rows already answered, so a small daily quota builds one growing sample.

  • DeepSeek, Z.ai GLM and Alibaba Qwen, which return real token probabilities.
  • Clearer verdict definitions for AITA.
  • More rows for the paid models.

© 2026 Tarkova Private Limited.