Docs
How the priority system works, what data flows where, and how to configure your tools.
How the priority system works, what data flows where, and how to configure your tools.
What leaves your machine and who can read it, how contributing raises your place in the queue when the network is busy, and how to point Claude Code at dollama. For endpoints, flags and setup steps, see the API reference, the CLI reference, Connect your app, Capabilities and Hosting.
Don't send secrets, credentials, or proprietary code you wouldn't share with a cloud API.
Your balance decides your place in the queue when the network is busy. It never decides whether you may use it.
balance = tokens served − tokens consumed
local: prefix, neither add to nor subtract from it.The short version of the Privacy Policy: who handles a request, and what each of them can read.
Runs your tools and keeps the saved copy of your conversation history. Whatever your tools put into a request (file contents, command output) is sent for inference.
Receives the whole conversation so far with every request, because it has to read it to run the model. It cannot reach your files and is not told who you are. dollama's software does not keep the content after the request; a contributor running modified software could.
| Data | Your machine | Relay | Machine that runs it |
|---|---|---|---|
| Files on disk | Stay here | No access | No access |
| File contents your tool read into the conversation | Stored here | Passes through, readable | Readable |
| Prompt, conversation history, response | Stored here | Readable; held up to 10 minutes | Readable, for the length of the request |
| Who you are | Known | Account ID; GitHub login if linked | Not sent |
Private mode runs requests only on your own machines, Group adds your groups' machines, and Open Network adds volunteers'. Private and Group never fall back to the open network: with no eligible machine, the request waits up to about two minutes and then fails. In all three the request still goes through the relay.
Run dollama launch claude — it starts the local proxy, configures the environment, and opens Claude Code automatically. That's it.
Run dollama network first, then dollama launch claude in another terminal — you'll use the network and share your idle compute.
dollama launch claude
{
"env": {
"ANTHROPIC_BASE_URL": "http://localhost:11435",
"ANTHROPIC_API_KEY": "dollama-proxy"
},
"model": "network:qwen3.5:9b"
}
ollama pull qwen3.5:9b
dollama network: dollama benchmarks your machine and the relay assigns it the role it can serve, from the 4B helper model on a CPU or small GPU up to the 9B worker on a GPU with about 8 GB of VRAM. A machine too slow to serve can still use the network as a client.localhost:11435 speaks the Anthropic Messages, OpenAI Chat Completions and OpenAI Responses APIs. Claude Code and Codex are tested, with dollama launch commands for each; OpenCode is lightly tested. Other tools that accept a custom base URL and use one of those formats should work but have not been tested. See Connect your app for a step-by-step guide (including Raspberry Pi and headless setups).