Capabilities
The network does more than chat completion. Transcribe recordings with dollama transcribe and turn text into embedding vectors with dollama embed. Each request goes to a machine that has chosen to offer that service.
The network does more than chat completion. Transcribe recordings with dollama transcribe and turn text into embedding vectors with dollama embed. Each request goes to a machine that has chosen to offer that service.
Two commands do work other than chat. Both take a local input and print JSON:
dollama transcribe recording.wav
dollama embed "what is the capital of France?"
dollama transcribe turns a recording into text with per-word timestamps.
dollama embed turns text into vectors for search or retrieval. Details for each follow below.
Where these requests run. A machine can offer transcription
(audio_stt) or embeddings in addition to whatever LLM role it was assigned, and the
relay keeps a separate pool of machines for each. Your CLI sends the request, the relay picks a
machine from that pool, and the machine streams the result back. The serving machine pulls audio
from your computer over the content tunnel, and the relay does not store it. Text to embed travels
in the request itself, which the relay holds for a few minutes while it is served. The relay can
read what passes through it; see the Privacy Policy.
Routing & privacy explains which machines are eligible by default.
You can see what the live network currently offers — pool depth, benchmark medians, and each
capability's wire contract — under the capabilities key of
GET /v1/network/models.
Sends a local audio file to a node advertising audio_stt and prints the transcript
as JSON. Per-word timestamps are on by default. Progress streams back as the transcription runs,
so a long clip reports partial text rather than sitting silent.
dollama transcribe recording.wav
dollama transcribe --language en --verbatim voicemail.ogg
dollama transcribe clip.wav --output transcript.json
| Flag | Type | Default | Description |
|---|---|---|---|
| --language | string | auto-detect | ISO-639-1 hint (e.g. en). Omit to let the model detect the language. |
| --verbatim | bool | false | Retain fillers and disfluencies (“um”, “uh”). Routes to a filler-preserving model with more precise timestamps. |
| --no-word-timestamps | bool | false | Omit per-word start/end times. Word timestamps are on unless you turn them off. |
| --initial-prompt | string | "" | Vocabulary or glossary bias for domain terms — names, jargon, product words the model would otherwise mishear. |
| --model | string | (network default) | Explicit STT model. Omit unless you have a reason. |
| --temperature | float | (unset) | Decoding temperature, 0 to 1. 0 is greedy and deterministic; pin it for reproducible transcripts. |
| --hotwords | string | "" | Bias the decoder toward these terms. Separate from --initial-prompt; you can set both. |
| --public | bool | false | Also allow public volunteers' machines. Off by default, so only your own machines are used — see Routing & privacy. |
| --output | string | stdout | Write the JSON result to this file instead of stdout. |
| --relay-url | string | (config) | Relay URL override. |
| --token | string | (config) | Auth token override. |
{"text", "language", "words", "duration_s"}. Each entry in words carries
word, start, and end in seconds — enough to build subtitles,
seek to a phrase, or align a transcript against the source audio.
{
"text": "the quick brown fox",
"language": "en",
"duration_s": 3.0,
"words": [
{ "word": "the", "start": 0.00, "end": 0.20 },
{ "word": "quick", "start": 0.20, "end": 0.55 }
]
}
Word-level probability is part of the schema but the backend does not populate it in
practice, so treat it as absent rather than as a confidence score.
Use --initial-prompt when the recording contains names or jargon. It biases the decoder
toward words it would otherwise guess wrong:
dollama transcribe standup.wav \
--initial-prompt "dollama, Ollama, Qwen, relay, OET, Fly.io, Neon"
Pass --language when you know it. Auto-detection costs accuracy on
short or noisy clips. Use --verbatim only when hesitations matter —
for meeting notes or captions, leave it off so fillers are dropped.
STT runs on Speaches,
an OpenAI-compatible faster-whisper server that dollama installs and manages alongside Ollama.
dollama setup offers to install it; you can also add it later:
dollama audio-install # install and start Speaches
dollama audio-probe # check whether this machine can host it
dollama benchmark --stt # measure RTF and WER, cache the result
dollama benchmark --stt --sweep # compare whisper model sizes on your hardware
The install is CPU-capable and gated only on disk space, not on VRAM — a machine too small to serve LLM inference can still contribute transcription, which is why setup recommends it more strongly on client-only hardware. You can also enable, benchmark, and compare models from the Settings → Speech-to-Text card in the local dashboard.
Your benchmark decides whether the node joins the audio_stt pool at all:
| Axis | Meaning | Gate |
|---|---|---|
| rtf | Real-time factor — processing time divided by audio duration. Lower is better; below 1.0 means faster than real time. | ≤ 1.0 |
| wer | Word error rate against reference transcripts, as a percentage. | ≤ 25% |
WER needs a real speech corpus with reference transcripts to measure. Set stt_corpus_dir
in your config to a folder of clip.wav / clip.txt pairs to calibrate it;
without one, the benchmark measures RTF against a synthetic probe and reports WER as uncalibrated.
Sends one or more texts to a node advertising embeddings and prints the vectors as
JSON. Built for scripts, RAG indexing pipelines, and one-off lookups — not an interactive session.
It lets you build retrieval on dollama without a local GPU budget for an embedding model.
dollama embed "what is the capital of France?"
# A batch in one round trip — measurably faster per item than looping.
dollama embed "first chunk" "second chunk" "third chunk"
dollama embed --file chunks.txt # one text per non-empty line
cat chunks.txt | dollama embed # same, from stdin
| Flag | Type | Default | Description |
|---|---|---|---|
| --file | string | "" | Read texts from this file, one per non-empty line, instead of args or stdin. |
| --no-truncate | bool | false | Error instead of truncating inputs that exceed the model's context. |
| --model | string | (network default) | Embedding model override. The network supports one embedding model so every vector is comparable — see below. |
| --output | string | stdout | Write the JSON result to this file instead of stdout. |
| --relay-url | string | (config) | Relay URL override. |
| --token | string | (config) | Auth token override. |
{"model", "dims", "embeddings", "usage"}. embeddings is one float array
per input text, in the same order as the inputs, and dims is the
vector width.
{
"model": "qwen3-embedding:0.6b",
"dims": 1024,
"embeddings": [[0.0123, -0.0456, "…"]],
"usage": { "prompt_tokens": 8 }
}
Batch rather than loop. Sending N texts in one call is one network round trip and one backend call; sending them one at a time is N of each. A single request accepts up to 512 texts or 8 MiB of text, whichever comes first.
The network supports a single embedding model rather than a hot-swappable set. Vectors from different models are not comparable, so one model network-wide is what makes an index built today still searchable tomorrow — and it keeps a small embedding model resident alongside a node's LLM role instead of thrashing them against each other.
Embeddings need no extra host process — Ollama already serves embedding models
natively, so your node's existing Ollama is the backend. dollama setup offers to pull
the network's embedding model and benchmark it; the Settings → Embeddings card in the
dashboard does the same at any time.
dollama benchmark --embeddings # measure throughput and latency
dollama benchmark --embeddings --sweep # compare embedding models on your hardware
The model is small (0.6B parameters), but it still needs memory. On one 12 GB card, Ollama reported its runner at about 4 GB in total, with 1.5 GB of that on the GPU. Dollama keeps it loaded next to your LLM only if free GPU memory allows; otherwise it loads the model when a request arrives. Whether a machine joins the pool depends on the benchmark thresholds below.
| Axis | Meaning | Gate |
|---|---|---|
| embed_tok_per_sec | Input tokens embedded per second on a single-item call. | ≥ 20 |
| p95_latency_ms_single | 95th-percentile latency for a single-item request. | ≤ 2000 ms |
| batch_speedup | Batch throughput divided by single-item throughput. | informational |
batch_speedup never gates membership — it exists so the fleet-wide benchmark surface
can show how much batching actually helps. A node below either real gate simply doesn't join the
pool; nothing else about it changes.
Three separate questions apply here. Where does the work run? By default,
dollama transcribe sends your audio only to machines on your own account. It does not use
group members' machines, and if none of yours is available the request fails instead of going elsewhere.
Pass --public to also allow public volunteers' machines.
Whose requests does a machine serve? Only machines running in Open Network mode
(dollama network) serve speech-to-text and embeddings, so one of your machines must be
running that mode, even if it is the machine you are typing on.
Who can read the audio? The serving machine, and the relay while the audio passes
through it.
The relay enforces the own-machines rule twice: when it picks a machine, and again when that machine claims the request.
In CLI v0.72.0 and earlier, dollama embed always uses the Open Network, whatever mode
the app is in. A fix that makes it follow your chosen mode
(dollama private / group / network) is merged and will ship in
the next release. Until then, do not use dollama embed for sensitive material. See
Privacy Policy §12.
Audio clips travel as a content reference the serving node pulls over the reverse content tunnel; the relay passes the audio along in memory and does not store it. Embedding inputs ride inline with the request for latency reasons, so they sit in the relay's cache like any other request: minutes, then they expire. No content is written to the relay's database or logs. The relay can read what passes through it. Each hop is encrypted in transit (TLS), but requests are not yet encrypted from your machine to the serving machine, so the relay can read them. See the Privacy Policy for the retention table and Docs for the full data-flow picture.
| Not available | What to do instead |
|---|---|
OpenAI-compatible /v1/audio/transcriptions or /v1/embeddings REST endpoints |
Use the CLI. A request to POST /v1/messages naming the capability needs the reverse content tunnel, so a plain HTTP client cannot send it; see the API Reference and Connect your app. |
| Microphone capture and voice-activity detection | Record with any tool you like and pass the file. dollama has no client-side capture pipeline. |
| Streaming audio in — live dictation from an open mic | Send whole clips. Transcription progress streams back out, but the input is a complete file; clips are capped at 5 minutes. |
| Text-to-speech / read-aloud | Not implemented. Designed, not built. |
| Speaker diarization (“who said what”) | Not implemented. |
| Vector storage or search | dollama embed returns vectors; storing and searching them is yours. Any vector store works. |
Both capabilities depend on nodes choosing to advertise them. If a request fails with no node
available, check the capabilities key of
GET /v1/network/models for the current pool depth — and consider serving it yourself.