Scaling and capacity
One Curva server handles thousands of decisions a second on a laptop. In practice the model provider limits you long before Curva does: a decision takes as long as its model call, and Curva adds well under a millisecond to it. This page gives the numbers, what limits them, and how to grow past one server.
Single-node numbers
Section titled “Single-node numbers”Measured on an Apple M1 (8 cores, 16 GB) with a zero-latency fake model, a real on-disk SQLite
database and API-key auth on. The load generator runs in the same process, so every number
includes loopback HTTP and competes for the same cores. The tests are in the repository
(crates/curva/src/server/mod.rs); run any of them with
cargo test --release -p curva <name> -- --ignored --nocapture.
Throughput (scaling_test)
Section titled “Throughput (scaling_test)”Saturated POST /v1/decide (two questions, audit log and metrics written), 64 clients:
| Server worker threads | Decisions/s |
|---|---|
| 1 | 7,400 |
| 2 | 8,400 |
| 4 | 10,700 |
| 8 | 8,800 (8 server + 4 load threads on 8 cores) |
The same load generator reaches 88,000 req/s on GET /health, so the limit is in the decision
path. It is the database: the store alone, on one thread, does about 11,000 decisions/s (about
90 µs each, mostly its two commits: the decision rows and the audit row). More cores stop helping
at that point, because SQLite has one writer.
A big database (big_db_test)
Section titled “A big database (big_db_test)”1,000,000 decisions over 50 projects × 20 questions and a year, 105,000 feedback labels and 1,000,000 audit rows: a 470 MB file. One question has 50,000 decisions and 5,000 labels.
| Call | Median |
|---|---|
POST /v1/decide |
0.2 ms |
POST /v1/feedback, refitting on 5,000 labels |
16 ms (15 ms reading the labels, 1 ms fitting) |
POST /v1/feedback, between refits |
0.4 ms |
GET /v1/calibration (5,000 labels) |
17 ms |
GET /v1/audit, first page / a page 20,000 entries deep |
0.4 ms / 0.2 ms |
GET /v1/drift, 8 / 53 weeks, question with 50,000 decisions |
3 ms / 23 ms |
GET /v1/drift, 8 / 53 weeks, question with 1,000 decisions |
0.1 ms / 0.6 ms |
| Opening the store | 1 ms |
How these stay flat as data grows:
- A refit reads the question’s labels through an index on feedback and uses the most recent 5,000, so its cost no longer grows with the question’s decisions (reading 5,000 labels among 50,000 decisions took 150 ms warm and up to 3 s cold before). A question that already has a calibrator refits at most every 5 s; labels in between are stored and counted at once.
- Drift reads one covering index: a year of a busy question is a single range scan.
- Audit pages seek straight to
before, however deep. - Long reads (audit pages, drift, calibration reports, refits) run on four read-only connections, off the async workers. While four clients ran 53-week drift and calibration reports back to back, decisions kept a p99 of 5 ms (1.3 s with those reads on the writer’s connection).
- The WAL is checkpointed and truncated once a second on its own connection. Constant reads otherwise keep SQLite from restarting it, and every commit then pays a checkpoint.
The first start on a database from Curva 0.1.0 builds two new indexes and fills two new feedback columns: about 1 s per million decisions, once.
Long runs and many connections (soak_test, connections_test)
Section titled “Long runs and many connections (soak_test, connections_test)”-
10 minutes at 1,000 req/s of mixed traffic (per 100 requests: 88 decisions using rules,
when,depends_onand extraction, 5 feedback labels, 2 each of calibration, audit and 53-week drift reports, 1 metrics scrape): 600,001 requests, 0 errors, p50 0.30 ms, p99 39 ms, p99.9 0.4 s (0.4% over 100 ms). The database grew to about 750 MB, the WAL stayed at about 4 MB. RSS went from 38 MB after warm-up to about 70 MB at the end (lowest values; brief peaks to 122 MB while a backlog drained). -
10,000 concurrent keep-alive connections against a model that takes 2 s: all 10,000 requests answered in 3.1–3.4 s, twice in a row over the same connections. Peak RSS was 0.7–0.95 GB, for both ends of every connection (client and server share the process).
-
A slow provider: 2,000 requests always in flight against the 2 s model for 60 s: about 920 answered per second, none failed, RSS flat at about 350 MB.
Memory follows the requests in flight, not the time running. A slow provider means more requests in flight: cap concurrency at your reverse proxy if the provider can stall.
What limits one node
Section titled “What limits one node”- Your model provider. Rate limits, latency and cost bind first, nearly always. Use the decision cache, cascades and rules (Speed and cost) before anything on this page.
- SQLite’s single writer. About 10,000 decisions/s on a fast NVMe disk, fewer on network storage. Each decision writes one row per question and one audit row; reads never wait for writes (WAL mode).
- Disk. About 0.5 KB per question answered plus 0.3 KB of audit per decision. A million decisions with two questions each is roughly 1.3 GB.
- Memory. Tens of MB idle; add roughly 50–100 KB per request in flight.
Why SQLite
Section titled “Why SQLite”One file, no server to run, backups by copying, and faster than a network database for this
workload (every call is a few indexed statements). Its ceiling is one writer per database, which
is several times what one provider account allows. Durability: WAL with synchronous=NORMAL
survives a crash of Curva; after a power loss the last few commits may be missing.
Backups
Section titled “Backups”The database is the curva.db file (plus -wal and -shm while running). Do not copy it with
cp while Curva runs; use SQLite’s online backup, which is consistent and doesn’t stop the
server:
sqlite3 curva.db ".backup '/backups/curva-$(date +%F).db'"The Docker image has no shell, so run it on the host against the volume’s file, or from a small container that mounts the same volume. Run it from cron and ship the file off the machine. For a recovery point of seconds rather than a day, use a continuous SQLite replication tool that streams the WAL to object storage (S3, GCS, Azure Blob) and can restore to a point in time.
Several instances today
Section titled “Several instances today”Curva can run as several instances now, each with its own database:
- Calibrators are fit from the labels in their own database, so a project’s decisions and its feedback must reach the same instance.
- Put a load balancer in front that is sticky by project: route on the
projectfield (or a header your clients set from it, or a separate hostname per group of projects). A hash of the project name across instances works. - The audit log, drift and calibration reports of a project are on its instance.
- API keys are per database: create each key on every instance that should accept it.
This covers many independent projects (a platform or agency serving many customers). One very hot project stays on one instance: at ~10,000 decisions/s, that is rarely the limit.
When a shared database is worth it
Section titled “When a shared database is worth it”Add a shared database (Postgres, for instance) when one of these becomes true:
- One project needs more than one instance: more than ~10,000 decisions/s sustained, or instances that must be interchangeable (autoscaling, zero-downtime rolling deploys without sticky routing).
- You want one audit log and one set of calibrators across instances, or keys managed in one place.
- Your platform gives you managed Postgres with backups and failover, and a local disk with volume backups is the harder thing to run.
Until then one instance per database, with a backup, is simpler and faster. The store is one
module (crates/curva/src/store), so a Postgres backend replaces a file without touching the
API.

