The Distributed Open Llama Network · experimental

Your computers,
one API

Dollama turns the computers you already own into a personal AI cluster behind one local endpoint. Point your tools at it as you would a paid provider.

  • One endpoint for your coding tools and apps, in the Anthropic and OpenAI formats
  • Dollama picks which of your machines runs each request, and runs tasks in parallel across them
  • Opt in to a community network when you want more capacity
dollama mascot
– tokens processed
– tok/s capacity
– total users
– requests served

One cluster. All your machines.

Dollama connects the hardware you already own, so you can use your most capable computer’s compute when you are away from it. Each model runs on one machine; the cluster shares work out between machines. It does not pool their memory to run a larger model.

Use your desktop GPU from a lightweight laptop

Keep models and agents running on an always-on home server. Send work to whichever of your machines is best equipped for it.

Run several independent tasks at once

Run tasks across your machines in parallel. Continue working when one machine is busy or unavailable.

Keep inference on hardware you control

In Private mode your requests run only on your own machines. They still travel through the Dollama relay to get there, so read who can read your work before sending anything sensitive.

Machines on the network can also transcribe speech and produce embeddings, not only run chat models. Today those are served only by machines in Open Network mode. See capabilities.

Your machines first, the community only if you opt in

Each machine is benchmarked and given a role. A worker model does the main reasoning and tool calls; a small, fast model writes the plan and handles side jobs such as summarising and sorting context. Today those are Qwen 3.5 9B and Qwen 3.5 4B. Each request then runs on the first eligible machine, in this order. The mode you choose decides how far down the list it may go:

1

A machine you own

Private · Group · Open Network

Your requests prefer your own hardware first. A laptop on battery may use a desktop at home.

2

A machine shared by a group

Group · Open Network

Next, machines shared by groups you belong to. Private mode stops before this step.

3

A volunteer's machine

Open Network only

Last, and only in Open Network mode, an eligible contributor on the open network. Private and Group never reach this step: with no eligible machine, the request fails. The network is small and volunteer-run, so speed varies.

How a request is handled

A short plan, when it helps

Dollama checks each request. A long or multi-step one can get a short plan from the small, fast model before the worker starts; short follow-ups and mid-task tool results do not. You can turn planning off.

Worker and helper

The worker does the main reasoning and tool calls. Helpers take side jobs such as summarising and sorting context, and independent helper or sub-agent work can run at the same time on different machines of yours.

Recovery

If a worker stops early or an attempt comes back incomplete, Dollama can prompt it to continue, retry, or route the work elsewhere. Models still make mistakes.

Example sessionmodels and timings vary
# choose where work runs, then start your coding tool
$ dollama private
$ dollama launch claude

# notes Dollama adds to a turn inside the coding tool
[dollama] planned (guidance, qwen3.5:4b, 840ms)
[dollama] Recovered: retried incomplete response

Free to access

No subscription, per-token, network usage or membership fee. When the network is busy, contribution orders the queue: it affects priority, not permission.

Keep your tools

Point a compatible coding tool at the local endpoint. Your tools still run on your machine; Dollama only decides where the model runs. Claude Code and Codex are tested; OpenCode is lightly tested. Other tools that accept a custom base URL and speak one of the supported API formats should work but have not been tested.

What works today

Dollama is an experimental developer tool. This is what is available, what is still in testing, and what is only planned.

Available

  • Local endpoints for coding tools and applications
  • Private, Group and Open Network modes
  • Recovery: retries, continuation, tool-call repair

Available · experimental

  • Embeddings
  • Personal hosting of static sites
  • Direct links between your own machines

In testing · not yet available

  • A small model that decides when a request needs planning
  • Looking things up in offline reference collections

Planned

  • Client-to-client encryption
  • Sending less of each request to the relay
  • Public source code and licence

Where your work runs, and who can read it

The mode you choose decides which machines run your requests. In every mode the request travels through the Dollama relay, which can read it.

Private mode

Work runs only on machines signed in to your account, and your machine serves nobody else. If none of your machines can take a request, it waits up to about two minutes and then fails. It never falls through to a group or the open network.

dollama private

Group mode

Work may run on your own machines plus machines shared by groups you belong to. Like Private, Group fails closed: with no eligible machine, the request fails rather than reaching the open network.

dollama group

Open Network mode

Work may run on a volunteer's machine, and your idle hardware serves other people's requests. Use this mode only for tasks you are comfortable sending to a computer you don't control.

dollama network

What stays on your machine

  • Your tools run locally; what they put into a request is sent for inference
  • Your conversation history is stored locally
  • Your coding tool's permission prompts still decide what runs

What leaves, and who can read it

A request includes whatever your tool read into the conversation: file contents, command output, earlier turns. The relay carries it and holds it for a few minutes; the machine that runs it has to read it. Each hop is encrypted in transit (TLS), but requests are not yet encrypted client to client, so the relay can read them too. Keep secrets, credentials and regulated data out, in every mode. Privacy Policy

Install Dollama

Using other machines takes a few minutes. Running models on this machine also needs Ollama and several gigabytes of model downloads.

🦙
Ollama
Only to run models here
💾
8GB+ RAM
Recommended to run models
💻
macOS / Linux / Win
Cross-platform
🧠
qwen3.5:9b
Worker model, ~8 GB GPU
Terminal
curl -fsSL https://dollama.net/install.sh | sh

The installer downloads Dollama and opens the dashboard. To run models on this machine you also need Ollama; the setup wizard can install it and download the models. Using other machines needs neither.

PowerShell
irm https://dollama.net/install.ps1 | iex

Open Windows PowerShell (not Command Prompt) and paste the line above — irm and iex are PowerShell commands and won't work in cmd.exe. The installer downloads Dollama, adds a Start Menu shortcut, and opens the dashboard, where the setup wizard can install Ollama and download the models if you want to run models on this machine.

Aliases disabled? Use Invoke-RestMethod https://dollama.net/install.ps1 | Invoke-Expression instead.

Windows may show a SmartScreen warning, because the Dollama binary is not code-signed yet. The installer checks the download against the published SHA-256 checksums before installing, and you can check them yourself. If you are satisfied, choose "More info" then "Run anyway".

1

Choose where your work runs

Dollama starts idle. Pick a mode in the dashboard, the tray icon, or the terminal:

dollama private   # your own machines only
dollama group     # your own + your groups' machines
dollama network   # community compute; shares yours

Private and Group need at least one machine of yours (or your group's) that can run a model: install Dollama there too and sign in to the same account. Open Network works from any machine, including one with no GPU.

2

Launch your coding tool

dollama launch claude

No coding tool installed? dollama chat opens a chat in the terminal. In the current release it always uses the Open Network, whatever mode you chose.

3

Start working

Your requests now run wherever your mode allows. Dollama's local endpoint at localhost:11435 speaks the Anthropic Messages API and the OpenAI Chat Completions and Responses APIs, so other tools can use it too. See Connect your app.

What the installer does: downloads the latest dollama binary for your platform, verifies the checksum, and installs it to ~/.dollama/bin (macOS/Linux) or %LOCALAPPDATA%\dollama (Windows). On macOS it also creates ~/Applications/dollama.app for Spotlight/Launchpad; on Windows it adds a Start Menu shortcut. All platforms open the dashboard after install. Read the script before you run it: install.sh, install.ps1.

Direct binary download

After extracting on Windows: move dollama.exe to %LOCALAPPDATA%\dollama, add that folder to PATH, then run dollama install-app and launch via the Start Menu → Dollama shortcut (not by double-clicking the zip download).

Verify checksums

Download the archive for your platform and the checksum file into the same folder, then check one against the other. Linux:

curl -fsSLO https://dollama.net/dl/latest/dollama-linux-amd64.tar.gz
curl -fsSLO https://dollama.net/dl/latest/checksums.txt
sha256sum -c --ignore-missing checksums.txt

macOS: use the darwin archive and shasum -a 256 -c --ignore-missing checksums.txt. Windows PowerShell: run Get-FileHash dollama-windows-amd64.zip -Algorithm SHA256 and compare the result with that file's line in checksums.txt.

The herd at a glance

The network is small and early. Capacity is what is online right now; the other figures are all-time totals.

loading
Concurrent capacity
Messages handled simultaneously
loading
Processing capacity
Theoretical tok/s
loading
Requests served
All time
loading
Tokens processed
All time
loading
Contributors
All time unique node owners
loading
Users
All time

Tested with

Claude Code Codex OpenCode (lightly)

Endpoints

Anthropic Messages OpenAI Chat Completions OpenAI Responses

Runs models with

Ollama llama.cpp (experimental)

Common questions

Yes, on the Open Network. The machine that runs your request has to read it, including the earlier turns of the conversation and any file contents or command output your tool put into it. It is not told who you are and sees only what is in the request. Private and Group modes keep the work on machines you or your group control. In every mode the relay carries the request and can read it too; client-to-client encryption, which would stop the relay reading it, is on the roadmap and not in place today.
Not for anything you could not afford to expose. On the Open Network your code is read by a stranger's machine. Private mode keeps inference on your own machines, but the relay can still read the request: each hop is encrypted in transit, but not yet client to client. Keep secrets, credentials and regulated data out in every mode.
No. A machine that can't run a model can still use your other machines or the Open Network as a client. CPU-only machines can also serve speech-to-text, and small models if they pass the speed benchmark.
No. But when the network is busy, people who contribute useful compute get served first. When it's idle, everyone gets fast access.
No. One machine is enough to benefit: you still get planning, recovery and routing, plus the Open Network if you opt in. More machines simply give you more.
Ollama is the backbone of this project, and we love llamas. But a doe felt right for what we're building: gentle, graceful, part of a herd. Plus, dollama → doe.