Benchmarks
Curva is measured on well-known public datasets and on probes that test whether its
probabilities can be trusted. Every number here comes from curva bench --save, with the sample
size and a 95% range next to it. The datasets are public, and curva bench ships in the curva
program you install, so you can measure any model on your own labeled data (see
Measure your own model).
Public datasets
Section titled “Public datasets”8 public sets, 6,040 rows sampled with a fixed seed. Every run and every model sees the same rows, and each sample takes the labels in turn, so a small sample is balanced across labels.
| Set | Rows | Question | What it tests |
|---|---|---|---|
phishnchips |
2,000 | Noul: phishing? | Security triage, yes/no calibration |
banking77 |
770 (10 per intent) | Choice, 77 options | Large option sets (intent routing) |
boolq |
500 | Noul | Reading comprehension |
openbookqa |
500 | Choice A–D | Calibration on unseen tasks |
commonsenseqa |
500 | Choice A–E | Calibration on unseen tasks |
hellaswag |
500 | Choice A–D | Calibration on unseen tasks |
yelp (Yelp Review Full) |
500 | Score, 1–5 stars | Ordinal scores |
aita (Reddit AITA) |
770 | Choice YTA/NTA/ESH/NAH | Subjective judgement, Brier score |
Trust probes
Section titled “Trust probes”| Probe | What must hold | Curva mechanism |
|---|---|---|
| Negation consistency | P(x) + P(not x) = 1 | Noul is asked as a normalised two-option choice |
| Option-order stability | Reordering the options doesn’t change the answer | Order debiasing averages both orders |
| Fair coin | P(heads) ≈ 0.5 | Calibration on feedback |
Targets
Section titled “Targets”Nothing is claimed until curva bench --save has produced it.
| Metric | Target |
|---|---|
| Negation sum P(x) + P(not x) | within 0.02 of 1.00 |
| Option-order swap | < 5% of answers change |
| Fair coin P(heads) | 0.50 ± 0.05 |
| ECE after calibration, each public set | ≤ 0.03 |
| Overall accuracy (built-in sets) | ≥ 90% |
| Accuracy when automated at ≥ 0.9 confidence | ≥ 98%, on ≥ 70% of decisions |
| Adversarial (injected) inputs | ≥ 95%, never confidently wrong |
| Repeat decision | 0 ms, $0 (cache) |
Results (30 September 2026, calibration updated 1 October)
Section titled “Results (30 September 2026, calibration updated 1 October)”Free models, verbal mode, order debiasing on. Accuracy is weighted to each dataset’s natural label mix (samples are balanced, published accuracies are not). ECE after calibration is held-out: each half of the rows is calibrated by a fit on the other half, which is what you get once you send feedback; it equals the raw ECE when calibration would not help.
Gemini Flash-Lite
Section titled “Gemini Flash-Lite”| Set | n | Accuracy (95% range) | ECE | ECE after calibration | p50 latency |
|---|---|---|---|---|---|
| PhishNChips (phishing) | 100 | 80.8% (73%–89%) | 0.138 | 0.138 | 1146 ms |
| BANKING77 (77 intents) | 120 | 79.8% (73%–87%) | 0.097 | 0.097 | 1362 ms |
| BoolQ | 127 | 88.9% (83%–94%) | 0.081 | 0.081 | 1071 ms |
| OpenBookQA | 100 | 92.2% (87%–97%) | 0.056 | 0.056 | 1013 ms |
| CommonsenseQA | 100 | 82.0% (74%–90%) | 0.118 | 0.118 | 1263 ms |
| HellaSwag | 98 | 78.6% (70%–87%) | 0.032 | 0.032 | 1173 ms |
| Yelp (1–5 stars) | 100 | 54.0% (44%–64%) | 0.259 | 0.138 | 2832 ms |
| Reddit AITA (4 verdicts) | 100 | 67.2% (58%–76%) | 0.404 | 0.149 | 1318 ms |
Qwen 3.8 27B on Groq
Section titled “Qwen 3.8 27B on Groq”| Set | n | Accuracy (95% range) | ECE | ECE after calibration | p50 latency |
|---|---|---|---|---|---|
| PhishNChips (phishing) | 100 | 72.8% (64%–82%) | 0.283 | 0.209 | 178 ms |
| BANKING77 (77 intents) | 69 | 74.7% (64%–85%) | 0.252 | 0.157 | 351 ms |
| OpenBookQA | 100 | 85.0% (78%–92%) | 0.051 | 0.051 | 207 ms |
| CommonsenseQA | 100 | 84.0% (77%–91%) | 0.066 | 0.066 | 230 ms |
| HellaSwag | 100 | 86.0% (79%–93%) | 0.085 | 0.085 | 209 ms |
| Reddit AITA (4 verdicts) | 100 | 55.7% (46%–65%) | 0.475 | 0.221 | 265 ms |
What this shows:
- Calibration fixes models that lean or overclaim. Gemini’s AITA error falls from 0.404 to 0.149 and its Yelp error from 0.259 to 0.138 once bias scaling corrects its favourite answers (Yelp accuracy also rises from 54% to 59%); Groq’s overconfident phishing probabilities go from 0.283 to 0.209.
- It never makes good answers worse. Where a model is already well calibrated (the quiz sets), Curva leaves its probabilities as they are.
- Speed. Groq answers in about 0.2 s; Gemini Flash-Lite in about 1.1 s.
- Both run free within the providers’ free tiers.
Paid models (early: n = 20 per set)
Section titled “Paid models (early: n = 20 per set)”gpt-4.1-nano, gpt-4o-mini, gpt-4.1-mini and Claude Haiku 4.5 ran 20 balanced rows on each set, too few for firm per-set numbers. Measured cost per 1,000 decisions across the 8 sets: gpt-4.1-nano $0.078, gpt-4o-mini $0.118, gpt-4.1-mini $0.314, Claude Haiku 4.5 $1.64.
Trust probes (live server, auth and debiasing on)
Section titled “Trust probes (live server, auth and debiasing on)”| Probe | Ling 3.0 Flash (logprobs) | Nemotron 3 Super (verbal) |
|---|---|---|
| Negation: P(x) + P(not x), 3 tickets | 1.33 / 0.14 / 0.56 (fails: the model ignores “NOT”) | 1.000 / 1.000 / 1.000 |
| Option-order swap, 4 tickets × 2 orders | 0 of 4 change | not run |
| Fair coin P(heads) | 0.494 | not run |
Because Ling failed the negation probe and Nemotron passed it, the default config
curva-1.1.0 (curva-latest) uses Nemotron 3 Super. curva-1.0.0 keeps Ling for callers who
pinned it.
Measure your own model
Section titled “Measure your own model”Put your labeled rows in a folder as <name>.json, with the question and one row per example:
{"question": {"type": "noul", "instructions": "The customer asks for a refund"}, "rows": [{"state": {"ticket": "Charged twice, refund please"}, "label": true}]}Then run any model on it and re-analyse without new calls:
pip install curva-aicurva bench --dir my-evals --model @gemini/gemini-flash-lite-latest --per-set 100 --save run.jsonl refundscurva report run.jsonl --dir my-evalsA rerun with the same --save file skips rows already answered, so a small daily quota builds
one growing sample.
Next runs
Section titled “Next runs”- DeepSeek, Z.ai GLM and Alibaba Qwen, which return real token probabilities.
- Clearer verdict definitions for AITA.
- More rows for the paid models.

