CLI Reference
Every command, flag, and config option for the dollama CLI.
Every command, flag, and config option for the dollama CLI.
This page is the reference — every command, flag, and config key. For the concept doc (how priority, privacy, and routing actually work), see Docs.
Start as a local proxy that forwards LLM requests to your own nodes only. Fails closed — if no own node is available, the request fails rather than falling through to a group or the public network. Your coding tools talk to localhost:11435 and the proxy handles routing, auth, and streaming.
| Flag | Type | Default | Description |
|---|---|---|---|
| -p, --port | int | 11435 | Local proxy listen port |
| --model | string | (relay) | Model to route requests to |
| --no-tray | bool | false | Run without system tray (headless) |
| --verbose | bool | false | Enable verbose debug logging |
Like private, but widens routing to your own nodes plus nodes shared by groups you belong to. Also fails closed — no fall-through to the public network.
| Flag | Type | Default | Description |
|---|---|---|---|
| -p, --port | int | 11435 | Local proxy listen port |
| --model | string | (relay) | Model to route requests to |
| --no-tray | bool | false | Run without system tray (headless) |
| --verbose | bool | false | Enable verbose debug logging |
Full network mode: use the network and share your idle compute. Runs the local proxy on localhost:11435 and, when your hardware qualifies, contributes your local Ollama instance to the network over a WebSocket.
| Flag | Type | Default | Description |
|---|---|---|---|
| -p, --port | int | 11435 | Local proxy listen port |
| --ollama-url | string | http://localhost:11434 | Local Ollama API URL |
| --model | string | (relay) | Model to advertise |
| --no-tray | bool | false | Run without system tray (headless) |
| --verbose | bool | false | Enable verbose debug logging |
Launch the system tray application with a web dashboard for configuration and control. This is the default when running dollama with no arguments.
Start the local proxy, set ANTHROPIC_BASE_URL, and launch Claude Code automatically. The fastest way to get started.
| Flag | Type | Default | Description |
|---|---|---|---|
| -p, --port | int | 11435 | Local proxy listen port |
| --model | string | (relay) | Model to use |
| --opus | bool | false | Map claude-opus requests to network model |
| --sonnet | bool | false | Map claude-sonnet requests to network model |
| --haiku | bool | false | Map claude-haiku requests to network model |
| --no-plan | bool | false | Skip the planning pre-pass for this session (lower latency). Sends X-Dollama-Planning: off via ANTHROPIC_CUSTOM_HEADERS; set planning = "off" in config to make it the default. |
Show current configuration with all fields, values, and defaults.
Print the config file path (typically ~/.dollama/config.toml).
Print a template with all config fields documented.
All settings live in ~/.dollama/config.toml. CLI flags override config values.
| Key | Type | Default | Description |
|---|---|---|---|
| relay_url | string | https://api.dollama.net | Relay server URL |
| ollama_url | string | http://localhost:11434 | Local Ollama API URL |
| model | string | qwen3.5:9b | Default model for inference and serving |
| listen_port | int | 11435 | Proxy listen port (private mode) |
| dashboard_port | int | 11436 | Web dashboard port |
| mode | string | "" | Active mode: private, group, network, or "" (idle) |
| Key | Type | Default | Description |
|---|---|---|---|
| relay_token | string | "" | API token (auto-generated on first run) |
| relay_node_id | string | "" | Node UUID from registration |
| relay_secret | string | "" | Node secret from registration |
| user_id | string | "" | Stable user ID from relay |
| github_login | string | "" | GitHub username (set via dollama login) |
| Key | Type | Default | Description |
|---|---|---|---|
| opus_model | string | "" | Model for claude-opus requests |
| sonnet_model | string | "" | Model for claude-sonnet requests |
| haiku_model | string | "" | Model for claude-haiku requests |
| Key | Type | Default | Description |
|---|---|---|---|
| auto_update | bool | true | Auto-update when new versions are available |
| battery_mode | string | pause_serve | On battery: pause_serve, keep_running, or stop_all |
| serve_num_ctx | int | 0 | Override Ollama num_ctx per request (0 = model default) |
| language | string | (system) | Response language (e.g. "English") |
| Key | Type | Default | Description |
|---|---|---|---|
| optimize_enabled | bool | true | Master switch for context optimization |
| optimize_tools_enabled | bool | true | Enable tool filtering |
| optimize_tools_strip_descriptions | bool | true | Strip tool descriptions |
| optimize_tools_simplify_schemas | bool | true | Simplify tool schemas |
| optimize_tools_max_tools | int | 0 | Max tools to keep (0 = unlimited) |
| optimize_context_enabled | bool | true | Enable context truncation |
| optimize_context_max_tokens | int | 6000 | Max context tokens |
| optimize_context_truncate_tool_results | int | 4000 | Truncate tool result tokens |
| optimize_context_prune_stale | bool | true | Prune stale messages |
| optimize_format_enabled | bool | true | Enable format optimization (markdown, whitespace, etc.) |
Direct, mutually authenticated tunnels between the computers on your own account (LAN first, relay fallback). Managed with dollama peer enable | disable | status, or from the dashboard under Settings → Experimental; every key below except peer_enabled only applies once that master switch is on. Changing any of them reconnects the serving connection (in-flight work drains first) — the peer identity is advertised in the connection handshake, so it cannot be attached to a live one.
| Key | Type | Default | Description |
|---|---|---|---|
| peer_enabled | bool | false | Master switch. On, this machine opens a UDP listener — the only non-loopback port dollama opens |
| peer_listen_port | int | 4256 | QUIC/UDP listen port (dollama peer enable --port N). A port already in use fails non-fatally: check dollama peer status for "listening" |
| peer_lan_discovery | bool | true | mDNS advertise + browse on the local network. A hint only — every connection still authenticates against pinned keys |
| peer_sync_enabled | bool | true | Replicate your hosted sites to your other peer-enabled machines |
| peer_content_enabled | bool | false | When one of your machines serves your own request, let it pull the conversation content over the direct link instead of through the relay. Needs to be on at both ends; falls back to the relay on any failure |
The relay operator can switch direct tunnels off for a whole account independently of these keys, which closes existing tunnels; dollama peer status and the dashboard both report it when that happens.
| Key | Type | Default | Description |
|---|---|---|---|
| terms_accepted_use | bool | false | Accepted terms for connect (use) mode |
| terms_accepted_give | bool | false | Accepted terms for serve (give) mode |
Authenticate with GitHub using the Device Authorization Flow. Opens a browser to github.com/login/device where you enter a one-time code. Links your GitHub account to your dollama identity.
Clear your authentication token. Optionally revokes the token server-side via POST /v1/auth/token/revoke.
Display the currently authenticated GitHub account and user ID.
dollama host publishes a static directory at <name>.host.dollama.net, served directly from your own node while it runs dollama network. For the full walk-through, scopes, and how it works end to end, see Personal Hosting.
Register a static directory as a hosted site and bind it to your current node. Provide exactly one scope. <dir> is the site root — put an index.html there.
| Flag | Type | Default | Description |
|---|---|---|---|
| --name | string | (required) | DNS label; becomes <name>.host.dollama.net |
| --private | bool | false | Owner-only site: only you can view it. Unrelated to Private mode for inference. Choosing no scope is an error |
| --group | string | "" | Share with this group by name (members only) |
| --public | bool | false | Anyone with the link, no login |
List your hosted sites with scope, bound node, and live online/offline status.
Revoke a hosted site: remove it from the relay and drop the local binding. Deletes the registration, not the files on disk.
Work the network does beyond LLM inference. See Capabilities for the full guide, including how to serve either capability from your own node.
Transcribe a local audio clip on a node advertising the audio_stt capability. Prints
{"text","language","words","duration_s"} as JSON, with per-word timestamps by default.
By default the audio goes only to machines on your own account, not to group members' machines or
public volunteers. If none of yours is available the command fails. Those machines must be running
dollama network, because only Open Network mode serves speech-to-text. Pass
--public to also allow public volunteers' machines.
| Flag | Type | Default | Description |
|---|---|---|---|
| --language | string | auto-detect | ISO-639-1 language hint (e.g. en) |
| --verbatim | bool | false | Retain fillers/disfluencies (um, uh) |
| --no-word-timestamps | bool | false | Omit per-word start/end times |
| --initial-prompt | string | "" | Vocabulary/glossary bias for domain terms |
| --model | string | (network) | STT model override |
| --public | bool | false | Also allow public volunteers' machines (default: your own machines only) |
| --output | string | stdout | Write the JSON result to this file |
| --relay-url | string | (config) | Relay URL override |
| --token | string | (config) | Auth token override |
Get embedding vectors from a node advertising the embeddings capability. Accepts
arguments, --file, or stdin. Prints
{"model","dims","embeddings","usage"} as JSON, one vector per input in input order.
Batch in one call rather than looping — up to 512 texts or 8 MiB per request.
| Flag | Type | Default | Description |
|---|---|---|---|
| --file | string | "" | Read texts from this file, one per non-empty line |
| --no-truncate | bool | false | Error instead of truncating over-long inputs |
| --model | string | (network) | Embedding model override |
| --output | string | stdout | Write the JSON result to this file |
| --relay-url | string | (config) | Relay URL override |
| --token | string | (config) | Auth token override |
Install and start the Speech-to-Text backend (Speaches), or check whether this machine can host
it. dollama setup offers this during onboarding; these are the standalone forms.
audio-probe exits 3 when the machine is not feasible.
Display the current state of the network including online nodes, capacity, and supported models. Calls GET /v1/status under the hood.
Send a minimal test request and trace its lifecycle through the network step-by-step. Useful for diagnosing why requests fail silently.
| Flag | Type | Default | Description |
|---|---|---|---|
| --relay-url | string | (config) | Relay URL |
| --token | string | (config) | Authentication token |
| -V, --verbose | bool | false | Show raw SSE frames |
Post-deploy smoke test: deep health check, node count, trace request, network health. Exits 0 on success, 1 on failure.
| Flag | Type | Default | Description |
|---|---|---|---|
| --relay-url | string | (config) | Relay URL |
| --token | string | (config) | Authentication token |
| --timeout | duration | 60s | Overall smoke test timeout |
Simulate fake compute nodes that connect via WebSocket and complete tasks with synthetic responses. No real Ollama needed.
| Flag | Type | Default | Description |
|---|---|---|---|
| --count | int | 1 | Number of simulated nodes |
| --delay | duration | 500ms | Simulated inference delay |
| --tokens | int | 20 | Simulated output token count |
| --fail-rate | float | 0 | Fraction of requests that error (0.0-1.0) |
| --timeout-rate | float | 0 | Fraction of requests that timeout (0.0-1.0) |
| --model | string | (relay) | Model to register as |
| --relay-url | string | (config) | Relay URL |
| --token | string | (config) | Authentication token |
Send concurrent test requests to measure end-to-end performance including routing, queue wait, and streaming latency.
| Flag | Type | Default | Description |
|---|---|---|---|
| --concurrent | int | 5 | Number of parallel requests |
| --count | int | 20 | Total requests to send |
| --timeout | duration | 120s | Per-request timeout |
| --relay-url | string | (config) | Relay URL |
| --token | string | (config) | Authentication token |
Run full-stack integration tests (relay terrarium + CLI E2E). Spins up an in-process test server and exercises the complete data path.
| Flag | Type | Default | Description |
|---|---|---|---|
| --scenarios | strings | all | Scenarios: happy-path, concurrent, node-failures, node-timeouts, node-disconnect, zero-nodes, burst-drain, sse-passthrough, idle-timeout, or all |
| --suite | string | all | Test suite: relay, e2e, or all |
| --replay-case | strings | Replay fixture case(s) for capture-replay tests | |
| --verbose | bool | false | Pass -v to go test |
| --short | bool | false | Skip slow tests (idle-timeout) |
| --timeout | duration | 5m | Overall test timeout |
Test whether a model supports function calling (tool use) by sending a minimal tool-calling request to local Ollama.
| Flag | Type | Default | Description |
|---|---|---|---|
| --ollama-url | string | http://localhost:11434 | Local Ollama API URL |
| --model | string | (config) | Model to test |
| -V, --verbose | bool | false | Show detailed response content |
Apply aggressive context compression to a captured Claude Code request JSON. Useful for analyzing compression effectiveness.
| Flag | Type | Default | Description |
|---|---|---|---|
| --stats | bool | false | Print size report only, no JSON output |
| -o, --output | string | "" | Write compressed JSON to file |
| --model | string | "" | Override model name |
Run a benchmark suite against your local Ollama instance to measure TPS, max context length, and concurrency limits. Results are saved under ~/.dollama/benchmarks/, one file per model.
| Flag | Type | Default | Description |
|---|---|---|---|
| --ollama-url | string | http://localhost:11434 | Local Ollama API URL |
| --model | string | (config) | Model to benchmark |
| --max-context | int | 0 | Max context length (0 = default 65536) |
| --conversation | bool | false | Run conversation degradation benchmark |
| --turns | int | 0 | Number of conversation turns (default 20) |
| --stt | bool | false | Benchmark the audio_stt capability instead — RTF, WER, and sustained concurrency. Saved to ~/.dollama/audio-benchmark.json. |
| --embeddings | bool | false | Benchmark the embeddings capability instead — throughput, p95 latency, and batch speedup. |
| --sweep | bool | false | With --stt or --embeddings, compare several models on your hardware sequentially instead of benchmarking one. |
Run agent workload evaluation against your local Ollama. Tests real-world coding scenarios to measure model capability.
| Flag | Type | Default | Description |
|---|---|---|---|
| --ollama-url | string | http://localhost:11434 | Local Ollama API URL |
| --model | string | (config) | Model to evaluate |
| --verbose | bool | false | Show detailed output |
| --timeout | duration | Per-task timeout | |
| --matrix | bool | false | Run comparison matrix |
| --strategy | string | Evaluation strategy | |
| --capture | bool | false | Capture request/response for replay |
| --session | string | Session name for grouping results |
Reset the stuck active request counter. Useful if the counter gets out of sync after a crash. Calls POST /v1/ledger/reset-active.
Install a launcher entry: macOS .app bundle (Spotlight/Launchpad + login agent), Linux .desktop file, or Windows Start Menu shortcut + login startup. Pass --no-login to skip auto-start at sign-in.
Remove dollama binary and config directory.
| Variable | Description |
|---|---|
| ANTHROPIC_BASE_URL | Override the Anthropic API base URL. Set to http://localhost:11435 to route through dollama. |
| ANTHROPIC_API_KEY | Anthropic API key. Set to dollama-proxy when using the local proxy. |
| ANTHROPIC_MODEL | Override the model name sent to the API. |
| DOLLAMA_DEBUG | Enable debug logging (set to any value). |
| DOLLAMA_DEBUG_UNSAFE | Enable unsafe debug features (set to "1"). |
| DOLLAMA_CAPTURE | Enable request capture mode for analysis. |
| DOLLAMA_MITIGATIONS | Enable mitigations (default: enabled; set to "0" to disable). |
| HOME | Home directory — config lives at $HOME/.dollama/config.toml. |
dollama benchmark, one file per model; --model runs go in benchmarks/adhoc/
dollama eval; every run is also appended to eval-history.jsonl