Skip to content

Model providers

Curva talks to any OpenAI-compatible chat API, and to several at once. The model id picks the provider:

  • vendor/model (no prefix) goes to the default provider, OpenRouter unless configured otherwise.
  • @<provider>/<model> goes to the named provider, e.g. @ollama/qwen3:4b or @vllm/Qwen/Qwen3-4B (only the first / splits, so vLLM is asked for Qwen/Qwen3-4B).

Every place that takes a model takes either form, so a council can mix a local model with a hosted one:

Terminal window
curl -s localhost:7777/v1/decide -d '{
"state": "Refund request, order #123, arrived broken",
"model": {"council": ["@ollama/qwen3:4b", "@groq/qwen/qwen3.8-27b"]},
"questions": {"intent": {"type": "choice", "options": ["refund", "exchange", "other"]}}
}'

Hosted providers turn on as soon as their usual key variable is set. Local ones need nothing: start the server and use the prefix. Model ids below are examples; check the provider for current model names.

Provider Base URL Key variable Example model id
openrouter (default) https://openrouter.ai/api/v1 OPENROUTER_API_KEY nvidia/nemotron-3-super-120b-a12b:free or @openrouter/...
openai https://api.openai.com/v1 OPENAI_API_KEY @openai/gpt-4.1-mini
anthropic https://api.anthropic.com/v1 ANTHROPIC_API_KEY @anthropic/claude-haiku-4-5
gemini https://generativelanguage.googleapis.com/v1beta/openai GEMINI_API_KEY or GOOGLE_API_KEY @gemini/gemini-flash-lite-latest
groq https://api.groq.com/openai/v1 GROQ_API_KEY @groq/qwen/qwen3.8-27b
mistral https://api.mistral.ai/v1 MISTRAL_API_KEY @mistral/mistral-small-latest
deepseek https://api.deepseek.com/v1 DEEPSEEK_API_KEY @deepseek/deepseek-chat
zai https://api.z.ai/api/paas/v4 ZAI_API_KEY @zai/glm-4.7-flash (free)
alibaba https://dashscope-intl.aliyuncs.com/compatible-mode/v1 DASHSCOPE_API_KEY @alibaba/qwen-flash
together https://api.together.xyz/v1 TOGETHER_API_KEY @together/meta-llama/Llama-3.3-70B-Instruct-Turbo
fireworks https://api.fireworks.ai/inference/v1 FIREWORKS_API_KEY @fireworks/accounts/fireworks/models/llama-v3p3-70b-instruct
xai https://api.x.ai/v1 XAI_API_KEY @xai/grok-3-mini
ollama (local) http://127.0.0.1:11434/v1 none @ollama/qwen3:4b
lmstudio (local) http://127.0.0.1:1234/v1 none @lmstudio/qwen3-4b
vllm (local) http://127.0.0.1:8000/v1 none @vllm/Qwen/Qwen3-4B
llamacpp (local) http://127.0.0.1:8080/v1 none @llamacpp/model

Keys are sent as Authorization: Bearer <key>.

Name it and point Curva at it; this also overrides any field of a built-in provider:

Variable Meaning
CURVA_PROVIDER_<NAME>_URL Base URL (the part before /chat/completions). The provider is @<name>, <NAME> in lowercase
CURVA_PROVIDER_<NAME>_KEY Optional key (wins over the built-in key variable)
CURVA_PROVIDER_<NAME>_RPM Requests per minute. Default: unlimited for a URL on this machine, else 60
CURVA_PROVIDER_<NAME>_TIMEOUT Seconds one call may take, 1 to 600. Default 60. Raise it for a provider that queues (e.g. CURVA_PROVIDER_NVIDIA_TIMEOUT=180)
CURVA_PROVIDER_<NAME>_PRICE <input>,<output> in dollars per million tokens, for cost_usd when the provider doesn’t report cost (e.g. CURVA_PROVIDER_MY_VLLM_PRICE=0.05,0.10 for your own GPU cost)
Terminal window
export CURVA_PROVIDER_GPU_URL=http://gpu-box:8000/v1 # @gpu/<model>
export CURVA_PROVIDER_OLLAMA_URL=http://127.0.0.1:9999/v1 # moves the built-in @ollama
export CURVA_PROVIDER_GROQ_TIMEOUT=120 # a built-in provider, longer calls

A value outside these ranges stops the server at startup with a message naming the variable.

  1. OpenRouter, when OPENROUTER_API_KEY (or CURVA_PROVIDER_URL) is set.
  2. Otherwise the one hosted provider whose key is set, if there is exactly one.
  3. Otherwise none: plain ids fail with a message listing the configured providers, and every model needs a @provider/ prefix.

curva serve logs the providers it found and where plain ids go (names only, never keys). It starts without any key; a request fails only when it needs a provider that is not configured (422).

Without --model and without OpenRouter, requests that name no model go to the first hosted provider in the table above whose key is set, using the first of its usual models that its /models endpoint lists (with only GEMINI_API_KEY: @gemini/gemini-flash-lite-latest). The Python client picks the same way.

  • Request body. OpenRouter gets its routing fields (provider.require_parameters, reasoning off, data_collection). Every other provider gets a plain OpenAI chat-completions body, since strict APIs reject fields they don’t know.
  • Probability mode. auto tries logprobs first and falls back to verbal, remembered per full model id. Providers without logprobs (or without JSON-schema output) work in verbal mode; results then depend on how well the model follows the schema.
  • Cost. cost_usd is what OpenRouter reports; for OpenAI and Anthropic it is their list price times the tokens used; for any other provider it is CURVA_PROVIDER_<NAME>_PRICE times the tokens, or 0 when no price is set (always 0 on your own machine).
  • privacy: strict. Enforced by OpenRouter’s data_collection: deny. On a named provider it is accepted only when the URL is on this machine (prompts never leave it); otherwise the request gets 422.
  • Limits. Each provider has its own rate limiter. CURVA_RPM and CURVA_DAILY_LIMIT apply to OpenRouter only. Retries (429/5xx) apply everywhere. A call times out after 60 s, or after CURVA_PROVIDER_<NAME>_TIMEOUT seconds on a named provider.
  • Models your account can’t use. When the provider says the model doesn’t exist, was retired, or isn’t open to your account (“not found for account”, “no longer available to new users”, “decommissioned”), the request gets 502 model_unavailable: check the model id and your access. The provider’s own words are in the server log only, never in the response.
  • Per-project keys (curva project-key set) are OpenRouter keys, so they apply to OpenRouter only; named providers always use their own key.
  • Cache. The decision cache is keyed by the full model id, so @ollama/m and m never share answers.

What we measured with new free keys on 2026-09-29. Providers change these limits often, so check yours before you plan around them.

Provider What a free key got
Gemini gemini-flash-lite-latest: about 500 requests a day. gemini-2.5-flash is closed to new accounts
Groq qwen/qwen3.8-27b: 200,000 tokens a day, about 180 debiased decisions
NVIDIA Calls often queue for over 60 s, and many listed models aren’t enabled for free accounts (“Not found for account”). With CURVA_PROVIDER_NVIDIA_URL=https://integrate.api.nvidia.com/v1, raise CURVA_PROVIDER_NVIDIA_TIMEOUT
OpenRouter Free models: 50 requests a day, 1,000 after a one-time $10 top-up

With no CURVA_MODEL, curva.decide and curva.local() ask the provider which models your key can use (GET /models) and take the first of a short list that is there (Gemini: gemini-flash-lite-latest, then gemini-flash-latest; Groq: qwen/qwen3.8-27b, then openai/gpt-oss-20b), so a model closed to new accounts doesn’t break a first run.

curva bench --save stops at the first daily-quota error and says so; the finished rows are in the file, and running the same command the next day continues where it stopped. Other failures (timeouts, unreadable replies) mark the row failed and the run goes on, unless 10 fail in a row.

© 2026 Tarkova Private Limited.