Back to dollama.net

Connect your app

Point Claude Code, Codex, OpenCode, a home-assistant bot or your own script at dollama through the local proxy. A tool can use it if it accepts a custom base URL and speaks one of the proxy's API formats (Anthropic Messages, OpenAI Chat Completions or OpenAI Responses).

One local endpoint; your mode decides which machines it reaches

When you start dollama in any mode (dollama private, dollama group or dollama network), the CLI starts a small local proxy that listens on http://localhost:11435 and speaks the Anthropic Messages API, plus the OpenAI Chat Completions and Responses APIs. Your app talks to that local address exactly as if it were talking to the provider directly — the proxy handles authentication, routing to an available machine, context optimization, and SSE streaming for you.

The mode you choose decides which machines the endpoint can use. Private uses only your own machines. Group uses yours and your groups'. Open Network also uses volunteers' machines.

The examples below use dollama private, which runs requests only on your own machines. That needs at least one machine on your account that can run a model; with none, requests wait about two minutes and fail. Use dollama network instead to run on the open network.

A tool needs two things: a way to override its API base URL, and support for one of those three formats. A custom base URL alone is not enough. Claude Code and Codex are tested end to end, and OpenCode more lightly. Other tools that meet both requirements should work with no code changes, but they have not been tested. Tools still run on your machine. What they read into the conversation is sent with the request; see what a request contains.

Three steps

1 — Install the CLI

macOS / Linux (including Raspberry Pi on a 64-bit OS):

install
curl -fsSL https://dollama.net/install.sh | sh
2 — Start the proxy

This runs in the foreground and prints the address it's listening on. On first run it auto-provisions an API token. Add --no-tray on a headless box (see Raspberry Pi & headless).

connect
dollama private
3 — Point your app at it

Set your app's Anthropic base URL to the local proxy and use the network model. Details below.

environment
export ANTHROPIC_BASE_URL="http://localhost:11435"
export ANTHROPIC_API_KEY="dollama-proxy"
export ANTHROPIC_MODEL="network:qwen3.5:9b"

What to point where

Most apps expose these as environment variables or settings. The values are always the same:

SettingValueNotes
Base URLhttp://localhost:11435Where the proxy listens. Override the port with dollama private -p <port>.
API keydollama-proxyA placeholder — the proxy supplies the real token. Any non-empty value works.
Modelnetwork:qwen3.5:9bThe network worker model. See Docs for the model tiers.

Remote device? If your app runs on a different machine than the proxy, bind/point at that machine's IP instead of localhost and make sure the proxy port is reachable on your LAN. Treat the proxy as trusted-LAN only — don't expose it to the public internet.

Common tools

Claude Code

One command does everything — starts the proxy, sets the env vars, and launches Claude Code:

launch
dollama launch claude
Any app via environment

Start dollama private in one terminal, then run your app with the env vars set. This is the generic pattern for Continue, Aider, home-assistant bots, and your own scripts.

shell
ANTHROPIC_BASE_URL="http://localhost:11435" \
ANTHROPIC_API_KEY="dollama-proxy" \
ANTHROPIC_MODEL="network:qwen3.5:9b" \
  your-app
Your own script (Python, Anthropic SDK)

Point the official SDK at the proxy by setting base_url:

python
from anthropic import Anthropic

client = Anthropic(
    base_url="http://localhost:11435",
    api_key="dollama-proxy",
)

msg = client.messages.create(
    model="network:qwen3.5:9b",
    max_tokens=512,
    messages=[{"role": "user", "content": "Say hello from the dollama network."}],
)
print(msg.content[0].text)
Your own script (curl)

The endpoint is plain Anthropic Messages — anything that can POST JSON works:

curl
curl --no-buffer -X POST http://localhost:11435/v1/messages \
  -H "Content-Type: application/json" \
  -H "x-api-key: dollama-proxy" \
  -d '{
    "model": "network:qwen3.5:9b",
    "max_tokens": 512,
    "stream": true,
    "messages": [{"role": "user", "content": "Hello from my app"}]
  }'

Running on small or headless devices

The CLI ships native Linux arm64 builds, so it runs on a Raspberry Pi — perfect for always-on assistants and bots. A couple of things to know first:

  • 64-bit OS required Use 64-bit Raspberry Pi OS (Pi 3/4/5). We don't publish 32-bit (armv7) binaries — check with uname -m (you want aarch64).
  • Which command If the Pi will use the open network, run dollama network. A Pi usually can't serve LLM inference, so it will use the network and serve nothing: with no Ollama, or no model that qualifies, the CLI starts the local proxy as a client and does not contribute compute. If Ollama later becomes reachable and the machine qualifies, it starts serving other people's requests. If the Pi should only reach your own other machines, run dollama private instead. That needs one of them to be running a model, or requests fail after about two minutes.
  • Run headless Add --no-tray so it never tries to start a system-tray UI.
Headless connect
shell
dollama network --no-tray    # use the open network
dollama private --no-tray    # use only your own machines
Keep it running with systemd

Drop this in ~/.config/systemd/user/dollama-connect.service, then systemctl --user enable --now dollama-connect:

dollama-connect.service
[Unit]
Description=dollama private proxy
After=network-online.target

[Service]
ExecStart=%h/.local/bin/dollama private --no-tray
Restart=on-failure

[Install]
WantedBy=default.target

Adjust ExecStart to wherever the installer placed the binary (which dollama). The unit uses private; change it to network if the Pi should use the open network.

Speech-to-text and embeddings

The local proxy does not serve these. It has no /v1/audio/transcriptions or /v1/embeddings endpoint; a request to /v1/embeddings returns a 404. There is no OpenAI-compatible route for either.

Today the supported way is to call the CLI from your application and read its JSON output. Both commands print JSON to standard output, or to a file with --output.

# Speech-to-text: prints {"text","language","words","duration_s"}
dollama transcribe recording.wav --output transcript.json

# Embeddings: prints {"model","dims","embeddings","usage"}
dollama embed "first chunk" "second chunk" > vectors.json

You cannot call the relay for these with a plain HTTP client either. Every request needs a live content tunnel from your machine, and audio is fetched over it, so the CLI is the route to use.

Limits. Speech-to-text and embeddings are served only by machines in Open Network mode. By default, dollama transcribe sends audio only to machines on your own account, not to group members' machines, so one of yours must be running dollama network. Pass --public to also allow public volunteers' machines. In CLI v0.72.0, dollama embed always uses the Open Network, whatever mode the app is in. See known gaps. Flags, output and routing are in Capabilities.

Talk to the relay without running the CLI on-device

dollama gateway is a small persistent process (not itself serverless) that holds the content tunnel open and exposes a plain HTTP endpoint. Use it when a client can't run the CLI itself, for example a serverless function or a constrained device. The pattern is: your client → HTTP → a self-hosted gateway (holds the tunnel) → the relay.

You can't call the relay's /v1/messages directly with a plain HTTP client. The relay requires a persistent content tunnel (X-Dollama-CLI-Tunnel-ID) that a stateless request can't provide. A deployment guide for the gateway will be published with the source.

If you'd rather implement the tunnel yourself instead of running a gateway, the full mechanism (token types, tunnel handshake, content offload) is in the reverse tunnel reference. Full request/response shapes, streaming events, and more language examples are on the API reference.

If something's off

  • Connection refused The proxy isn't running, or your app is pointed at the wrong port. Confirm dollama private is up and the base URL port matches (default 11435).
  • Empty or 0ms responses Usually no nodes are online to serve your request. Check network health with dollama status, or trace a single request with dollama debug trace.
  • Wrong architecture on install On a Pi, verify uname -m reports aarch64. A 32-bit OS won't have a matching binary.
  • App hardcodes api.anthropic.com It must support a custom base URL. If it only reads ANTHROPIC_BASE_URL, set that env var; otherwise look for a base-URL setting in its config.

Every flag and config key lives in the CLI reference. For privacy specifics (what nodes can see), see the Docs FAQ.