Back to dollama.net

Docs

How the priority system works, what data flows where, and how to configure your tools.

What this page covers

What leaves your machine and who can read it, how contributing raises your place in the queue when the network is busy, and how to point Claude Code at dollama. For endpoints, flags and setup steps, see the API reference, the CLI reference, Connect your app, Capabilities and Hosting.

Know exactly what you're sharing

On your machine

  • Your tools run here: file reads, edits and shell commands
  • Whatever they put into a request (file contents, command output) is sent for inference
  • The saved copy of your conversation history stays here, though the conversation so far is sent with each request

Readable outside your machine

  • Every request passes through the relay, in every mode
  • The relay can read it, and holds it for a few minutes while it is served
  • The machine that runs the request reads the whole conversation so far
  • Each hop is encrypted in transit (TLS), but requests are not yet encrypted client to client, so the relay can read them
  • The relay does not log content, and the request expires from its cache within minutes

Use it for

  • Open-source and hobby projects
  • Learning and experimentation
  • Non-sensitive coding tasks

Don't send secrets, credentials, or proprietary code you wouldn't share with a cloud API.

Contributing compute raises your priority when the network is busy

Your balance decides your place in the queue when the network is busy. It never decides whether you may use it.

Balance and queue

balance = tokens served − tokens consumed
  • Raises it — your machines serving other people's requests.
  • Spends it — your requests served by someone else's machine.
  • Busy network — one queue, ordered by balance bucket and how long each request has waited. A higher balance moves you up.
  • Idle network — a request goes out as soon as an eligible machine is free, whatever the balance. New accounts start at zero.

Waiting and groups

  • Waiting counts — the longer a request waits, the higher it ranks, so a low balance slows you down but does not lock you out. A request that finds no machine within about two minutes fails.
  • Groups — your group's machines are tried before the open network.
  • Balances cannot be bought or transferred.

Your own hardware

  • Free for your balance — requests served by your own machines, and models you run on your local Ollama with the local: prefix, neither add to nor subtract from it.
  • First in line — your requests go to your own machines ahead of other people's. A request already running there, or a model that has to load, can still delay yours.

How your data flows

The short version of the Privacy Policy: who handles a request, and what each of them can read.

Your machine

Runs your tools and keeps the saved copy of your conversation history. Whatever your tools put into a request (file contents, command output) is sent for inference.

The relay

  • Matches your request to a machine and carries the request and response between you, in every mode
  • It can read the content. Each hop is encrypted in transit (TLS), but requests are not yet encrypted client to client
  • Holds the request in its cache while it is served (up to 10 minutes) and for up to 5 minutes after it completes, then it expires; logs no content
  • Keeps usage records (who, which machine, token counts, timing) indefinitely

The machine that runs it

Receives the whole conversation so far with every request, because it has to read it to run the model. It cannot reach your files and is not told who you are. dollama's software does not keep the content after the request; a contributor running modified software could.

Data Your machine Relay Machine that runs it
Files on disk Stay here No access No access
File contents your tool read into the conversation Stored here Passes through, readable Readable
Prompt, conversation history, response Stored here Readable; held up to 10 minutes Readable, for the length of the request
Who you are Known Account ID; GitHub login if linked Not sent

Modes change which machine runs it, not the path

Private mode runs requests only on your own machines, Group adds your groups' machines, and Open Network adds volunteers'. Private and Group never fall back to the open network: with no eligible machine, the request waits up to about two minutes and then fails. In all three the request still goes through the relay.

Use with Claude Code

1

Launch with one command

Run dollama launch claude — it starts the local proxy, configures the environment, and opens Claude Code automatically. That's it.

tip

Want to contribute too?

Run dollama network first, then dollama launch claude in another terminal — you'll use the network and share your idle compute.

Launch Claude Code
dollama launch claude
Or configure manually in ~/.claude/settings.json
{
  "env": {
    "ANTHROPIC_BASE_URL": "http://localhost:11435",
    "ANTHROPIC_API_KEY": "dollama-proxy"
  },
  "model": "network:qwen3.5:9b"
}
Pull the model (contributors)
ollama pull qwen3.5:9b

Common questions

Yes. The machine that runs your request reads the prompt, including earlier turns and any file contents your tools put into it. It is not told who you are and cannot reach your files. In Private mode that machine is one of your own; in Open Network mode it can be a stranger's.
Not yet. The relay can read every request, in every mode, and so does the machine that runs it. Private mode removes the stranger but not the relay. Don't send credentials, secrets, or proprietary code you wouldn't share with a cloud API.
Everything in the request. Each hop is encrypted in transit (TLS), but requests are not yet encrypted client to client, and the relay carries them in between. It holds the request in its cache while it is served (up to 10 minutes) and for up to 5 minutes after it completes, after which it expires. It does not log content. What it keeps long term is usage records: which account, which machine, token counts and timing. See the Privacy Policy for the full table.
It is the current worker model: fast enough for consumer hardware (about 8 GB of VRAM) and capable enough for agentic coding. A smaller model, Qwen 3.5 4B, handles planning and helper tasks. See Models for the measurements behind those choices and their limits.
Ollama, and a machine that passes the benchmark. Run dollama network: dollama benchmarks your machine and the relay assigns it the role it can serve, from the 4B helper model on a CPU or small GPU up to the 9B worker on a GPU with about 8 GB of VRAM. A machine too slow to serve can still use the network as a client.
Yes. The local proxy at localhost:11435 speaks the Anthropic Messages, OpenAI Chat Completions and OpenAI Responses APIs. Claude Code and Codex are tested, with dollama launch commands for each; OpenCode is lightly tested. Other tools that accept a custom base URL and use one of those formats should work but have not been tested. See Connect your app for a step-by-step guide (including Raspberry Pi and headless setups).