Model providers
Curva talks to any OpenAI-compatible chat API, and to several at once. The model id picks the provider:
vendor/model(no prefix) goes to the default provider, OpenRouter unless configured otherwise.@<provider>/<model>goes to the named provider, e.g.@ollama/qwen3:4bor@vllm/Qwen/Qwen3-4B(only the first/splits, so vLLM is asked forQwen/Qwen3-4B).
Every place that takes a model takes either form, so a council can mix a local model with a hosted one:
curl -s localhost:7777/v1/decide -d '{ "state": "Refund request, order #123, arrived broken", "model": {"council": ["@ollama/qwen3:4b", "@groq/qwen/qwen3.8-27b"]}, "questions": {"intent": {"type": "choice", "options": ["refund", "exchange", "other"]}}}'Built-in providers
Section titled “Built-in providers”Hosted providers turn on as soon as their usual key variable is set. Local ones need nothing: start the server and use the prefix. Model ids below are examples; check the provider for current model names.
| Provider | Base URL | Key variable | Example model id |
|---|---|---|---|
openrouter (default) |
https://openrouter.ai/api/v1 |
OPENROUTER_API_KEY |
nvidia/nemotron-3-super-120b-a12b:free or @openrouter/... |
openai |
https://api.openai.com/v1 |
OPENAI_API_KEY |
@openai/gpt-4.1-mini |
anthropic |
https://api.anthropic.com/v1 |
ANTHROPIC_API_KEY |
@anthropic/claude-haiku-4-5 |
gemini |
https://generativelanguage.googleapis.com/v1beta/openai |
GEMINI_API_KEY or GOOGLE_API_KEY |
@gemini/gemini-flash-lite-latest |
groq |
https://api.groq.com/openai/v1 |
GROQ_API_KEY |
@groq/qwen/qwen3.8-27b |
mistral |
https://api.mistral.ai/v1 |
MISTRAL_API_KEY |
@mistral/mistral-small-latest |
deepseek |
https://api.deepseek.com/v1 |
DEEPSEEK_API_KEY |
@deepseek/deepseek-chat |
zai |
https://api.z.ai/api/paas/v4 |
ZAI_API_KEY |
@zai/glm-4.7-flash (free) |
alibaba |
https://dashscope-intl.aliyuncs.com/compatible-mode/v1 |
DASHSCOPE_API_KEY |
@alibaba/qwen-flash |
together |
https://api.together.xyz/v1 |
TOGETHER_API_KEY |
@together/meta-llama/Llama-3.3-70B-Instruct-Turbo |
fireworks |
https://api.fireworks.ai/inference/v1 |
FIREWORKS_API_KEY |
@fireworks/accounts/fireworks/models/llama-v3p3-70b-instruct |
xai |
https://api.x.ai/v1 |
XAI_API_KEY |
@xai/grok-3-mini |
ollama (local) |
http://127.0.0.1:11434/v1 |
none | @ollama/qwen3:4b |
lmstudio (local) |
http://127.0.0.1:1234/v1 |
none | @lmstudio/qwen3-4b |
vllm (local) |
http://127.0.0.1:8000/v1 |
none | @vllm/Qwen/Qwen3-4B |
llamacpp (local) |
http://127.0.0.1:8080/v1 |
none | @llamacpp/model |
Keys are sent as Authorization: Bearer <key>.
Any other provider
Section titled “Any other provider”Name it and point Curva at it; this also overrides any field of a built-in provider:
| Variable | Meaning |
|---|---|
CURVA_PROVIDER_<NAME>_URL |
Base URL (the part before /chat/completions). The provider is @<name>, <NAME> in lowercase |
CURVA_PROVIDER_<NAME>_KEY |
Optional key (wins over the built-in key variable) |
CURVA_PROVIDER_<NAME>_RPM |
Requests per minute. Default: unlimited for a URL on this machine, else 60 |
CURVA_PROVIDER_<NAME>_TIMEOUT |
Seconds one call may take, 1 to 600. Default 60. Raise it for a provider that queues (e.g. CURVA_PROVIDER_NVIDIA_TIMEOUT=180) |
CURVA_PROVIDER_<NAME>_PRICE |
<input>,<output> in dollars per million tokens, for cost_usd when the provider doesn’t report cost (e.g. CURVA_PROVIDER_MY_VLLM_PRICE=0.05,0.10 for your own GPU cost) |
export CURVA_PROVIDER_GPU_URL=http://gpu-box:8000/v1 # @gpu/<model>export CURVA_PROVIDER_OLLAMA_URL=http://127.0.0.1:9999/v1 # moves the built-in @ollamaexport CURVA_PROVIDER_GROQ_TIMEOUT=120 # a built-in provider, longer callsA value outside these ranges stops the server at startup with a message naming the variable.
Which provider gets plain model ids
Section titled “Which provider gets plain model ids”- OpenRouter, when
OPENROUTER_API_KEY(orCURVA_PROVIDER_URL) is set. - Otherwise the one hosted provider whose key is set, if there is exactly one.
- Otherwise none: plain ids fail with a message listing the configured providers, and every
model needs a
@provider/prefix.
curva serve logs the providers it found and where plain ids go (names only, never keys). It
starts without any key; a request fails only when it needs a provider that is not configured
(422).
Without --model and without OpenRouter, requests that name no model go to the first hosted
provider in the table above whose key is set, using the first of its usual models that its
/models endpoint lists (with only GEMINI_API_KEY: @gemini/gemini-flash-lite-latest). The
Python client picks the same way.
What differs per provider
Section titled “What differs per provider”- Request body. OpenRouter gets its routing fields (
provider.require_parameters, reasoning off,data_collection). Every other provider gets a plain OpenAI chat-completions body, since strict APIs reject fields they don’t know. - Probability mode.
autotries logprobs first and falls back to verbal, remembered per full model id. Providers without logprobs (or without JSON-schema output) work in verbal mode; results then depend on how well the model follows the schema. - Cost.
cost_usdis what OpenRouter reports; for OpenAI and Anthropic it is their list price times the tokens used; for any other provider it isCURVA_PROVIDER_<NAME>_PRICEtimes the tokens, or 0 when no price is set (always 0 on your own machine). privacy: strict. Enforced by OpenRouter’sdata_collection: deny. On a named provider it is accepted only when the URL is on this machine (prompts never leave it); otherwise the request gets 422.- Limits. Each provider has its own rate limiter.
CURVA_RPMandCURVA_DAILY_LIMITapply to OpenRouter only. Retries (429/5xx) apply everywhere. A call times out after 60 s, or afterCURVA_PROVIDER_<NAME>_TIMEOUTseconds on a named provider. - Models your account can’t use. When the provider says the model doesn’t exist, was retired,
or isn’t open to your account (“not found for account”, “no longer available to new users”,
“decommissioned”), the request gets 502
model_unavailable: check the model id and your access. The provider’s own words are in the server log only, never in the response. - Per-project keys (
curva project-key set) are OpenRouter keys, so they apply to OpenRouter only; named providers always use their own key. - Cache. The decision cache is keyed by the full model id, so
@ollama/mandmnever share answers.
Free tiers in practice
Section titled “Free tiers in practice”What we measured with new free keys on 2026-09-29. Providers change these limits often, so check yours before you plan around them.
| Provider | What a free key got |
|---|---|
| Gemini | gemini-flash-lite-latest: about 500 requests a day. gemini-2.5-flash is closed to new accounts |
| Groq | qwen/qwen3.8-27b: 200,000 tokens a day, about 180 debiased decisions |
| NVIDIA | Calls often queue for over 60 s, and many listed models aren’t enabled for free accounts (“Not found for account”). With CURVA_PROVIDER_NVIDIA_URL=https://integrate.api.nvidia.com/v1, raise CURVA_PROVIDER_NVIDIA_TIMEOUT |
| OpenRouter | Free models: 50 requests a day, 1,000 after a one-time $10 top-up |
With no CURVA_MODEL, curva.decide and curva.local() ask the provider which models your key
can use (GET /models) and take the first of a short list that is there (Gemini:
gemini-flash-lite-latest, then gemini-flash-latest; Groq: qwen/qwen3.8-27b, then
openai/gpt-oss-20b), so a model closed to new accounts doesn’t break a first run.
curva bench --save stops at the first daily-quota error and says so; the finished rows are in
the file, and running the same command the next day continues where it stopped. Other failures
(timeouts, unreadable replies) mark the row failed and the run goes on, unless 10 fail in a row.

