Back to dollama.net

Capabilities

The network does more than chat completion. Transcribe recordings with dollama transcribe and turn text into embedding vectors with dollama embed. Each request goes to a machine that has chosen to offer that service.

Speech-to-text and embeddings

Two commands do work other than chat. Both take a local input and print JSON:

dollama transcribe recording.wav
dollama embed "what is the capital of France?"

dollama transcribe turns a recording into text with per-word timestamps. dollama embed turns text into vectors for search or retrieval. Details for each follow below.

Where these requests run. A machine can offer transcription (audio_stt) or embeddings in addition to whatever LLM role it was assigned, and the relay keeps a separate pool of machines for each. Your CLI sends the request, the relay picks a machine from that pool, and the machine streams the result back. The serving machine pulls audio from your computer over the content tunnel, and the relay does not store it. Text to embed travels in the request itself, which the relay holds for a few minutes while it is served. The relay can read what passes through it; see the Privacy Policy. Routing & privacy explains which machines are eligible by default.

You can see what the live network currently offers — pool depth, benchmark medians, and each capability's wire contract — under the capabilities key of GET /v1/network/models.

Transcribe an audio clip

dollama transcribe <audio-file>

Sends a local audio file to a node advertising audio_stt and prints the transcript as JSON. Per-word timestamps are on by default. Progress streams back as the transcription runs, so a long clip reports partial text rather than sitting silent.

dollama transcribe recording.wav
dollama transcribe --language en --verbatim voicemail.ogg
dollama transcribe clip.wav --output transcript.json
FlagTypeDefaultDescription
--languagestringauto-detectISO-639-1 hint (e.g. en). Omit to let the model detect the language.
--verbatimboolfalseRetain fillers and disfluencies (“um”, “uh”). Routes to a filler-preserving model with more precise timestamps.
--no-word-timestampsboolfalseOmit per-word start/end times. Word timestamps are on unless you turn them off.
--initial-promptstring""Vocabulary or glossary bias for domain terms — names, jargon, product words the model would otherwise mishear.
--modelstring(network default)Explicit STT model. Omit unless you have a reason.
--temperaturefloat(unset)Decoding temperature, 0 to 1. 0 is greedy and deterministic; pin it for reproducible transcripts.
--hotwordsstring""Bias the decoder toward these terms. Separate from --initial-prompt; you can set both.
--publicboolfalseAlso allow public volunteers' machines. Off by default, so only your own machines are used — see Routing & privacy.
--outputstringstdoutWrite the JSON result to this file instead of stdout.
--relay-urlstring(config)Relay URL override.
--tokenstring(config)Auth token override.

Output

{"text", "language", "words", "duration_s"}. Each entry in words carries word, start, and end in seconds — enough to build subtitles, seek to a phrase, or align a transcript against the source audio.

{
  "text": "the quick brown fox",
  "language": "en",
  "duration_s": 3.0,
  "words": [
    { "word": "the",   "start": 0.00, "end": 0.20 },
    { "word": "quick", "start": 0.20, "end": 0.55 }
  ]
}

Word-level probability is part of the schema but the backend does not populate it in practice, so treat it as absent rather than as a confidence score.

Getting better transcripts

Use --initial-prompt when the recording contains names or jargon. It biases the decoder toward words it would otherwise guess wrong:

dollama transcribe standup.wav \
  --initial-prompt "dollama, Ollama, Qwen, relay, OET, Fly.io, Neon"

Pass --language when you know it. Auto-detection costs accuracy on short or noisy clips. Use --verbatim only when hesitations matter — for meeting notes or captions, leave it off so fillers are dropped.

Serving speech-to-text

STT runs on Speaches, an OpenAI-compatible faster-whisper server that dollama installs and manages alongside Ollama. dollama setup offers to install it; you can also add it later:

dollama audio-install          # install and start Speaches
dollama audio-probe            # check whether this machine can host it
dollama benchmark --stt        # measure RTF and WER, cache the result
dollama benchmark --stt --sweep  # compare whisper model sizes on your hardware

The install is CPU-capable and gated only on disk space, not on VRAM — a machine too small to serve LLM inference can still contribute transcription, which is why setup recommends it more strongly on client-only hardware. You can also enable, benchmark, and compare models from the Settings → Speech-to-Text card in the local dashboard.

Pool thresholds

Your benchmark decides whether the node joins the audio_stt pool at all:

AxisMeaningGate
rtfReal-time factor — processing time divided by audio duration. Lower is better; below 1.0 means faster than real time.≤ 1.0
werWord error rate against reference transcripts, as a percentage.≤ 25%

WER needs a real speech corpus with reference transcripts to measure. Set stt_corpus_dir in your config to a folder of clip.wav / clip.txt pairs to calibrate it; without one, the benchmark measures RTF against a synthetic probe and reports WER as uncalibrated.

Turn text into vectors

dollama embed [text...]

Sends one or more texts to a node advertising embeddings and prints the vectors as JSON. Built for scripts, RAG indexing pipelines, and one-off lookups — not an interactive session. It lets you build retrieval on dollama without a local GPU budget for an embedding model.

dollama embed "what is the capital of France?"

# A batch in one round trip — measurably faster per item than looping.
dollama embed "first chunk" "second chunk" "third chunk"
dollama embed --file chunks.txt          # one text per non-empty line
cat chunks.txt | dollama embed           # same, from stdin
FlagTypeDefaultDescription
--filestring""Read texts from this file, one per non-empty line, instead of args or stdin.
--no-truncateboolfalseError instead of truncating inputs that exceed the model's context.
--modelstring(network default)Embedding model override. The network supports one embedding model so every vector is comparable — see below.
--outputstringstdoutWrite the JSON result to this file instead of stdout.
--relay-urlstring(config)Relay URL override.
--tokenstring(config)Auth token override.

Output

{"model", "dims", "embeddings", "usage"}. embeddings is one float array per input text, in the same order as the inputs, and dims is the vector width.

{
  "model": "qwen3-embedding:0.6b",
  "dims": 1024,
  "embeddings": [[0.0123, -0.0456, "…"]],
  "usage": { "prompt_tokens": 8 }
}

Batch rather than loop. Sending N texts in one call is one network round trip and one backend call; sending them one at a time is N of each. A single request accepts up to 512 texts or 8 MiB of text, whichever comes first.

The network supports a single embedding model rather than a hot-swappable set. Vectors from different models are not comparable, so one model network-wide is what makes an index built today still searchable tomorrow — and it keeps a small embedding model resident alongside a node's LLM role instead of thrashing them against each other.

Serving embeddings

Embeddings need no extra host process — Ollama already serves embedding models natively, so your node's existing Ollama is the backend. dollama setup offers to pull the network's embedding model and benchmark it; the Settings → Embeddings card in the dashboard does the same at any time.

dollama benchmark --embeddings          # measure throughput and latency
dollama benchmark --embeddings --sweep  # compare embedding models on your hardware

The model is small (0.6B parameters), but it still needs memory. On one 12 GB card, Ollama reported its runner at about 4 GB in total, with 1.5 GB of that on the GPU. Dollama keeps it loaded next to your LLM only if free GPU memory allows; otherwise it loads the model when a request arrives. Whether a machine joins the pool depends on the benchmark thresholds below.

Pool thresholds

AxisMeaningGate
embed_tok_per_secInput tokens embedded per second on a single-item call.≥ 20
p95_latency_ms_single95th-percentile latency for a single-item request.≤ 2000 ms
batch_speedupBatch throughput divided by single-item throughput.informational

batch_speedup never gates membership — it exists so the fleet-wide benchmark surface can show how much batching actually helps. A node below either real gate simply doesn't join the pool; nothing else about it changes.

Where your audio and text actually go

Speech-to-text: your own machines by default

Three separate questions apply here. Where does the work run? By default, dollama transcribe sends your audio only to machines on your own account. It does not use group members' machines, and if none of yours is available the request fails instead of going elsewhere. Pass --public to also allow public volunteers' machines. Whose requests does a machine serve? Only machines running in Open Network mode (dollama network) serve speech-to-text and embeddings, so one of your machines must be running that mode, even if it is the machine you are typing on. Who can read the audio? The serving machine, and the relay while the audio passes through it.

The relay enforces the own-machines rule twice: when it picks a machine, and again when that machine claims the request.

Embeddings and your mode

In CLI v0.72.0 and earlier, dollama embed always uses the Open Network, whatever mode the app is in. A fix that makes it follow your chosen mode (dollama private / group / network) is merged and will ship in the next release. Until then, do not use dollama embed for sensitive material. See Privacy Policy §12.

What the relay holds

Audio clips travel as a content reference the serving node pulls over the reverse content tunnel; the relay passes the audio along in memory and does not store it. Embedding inputs ride inline with the request for latency reasons, so they sit in the relay's cache like any other request: minutes, then they expire. No content is written to the relay's database or logs. The relay can read what passes through it. Each hop is encrypted in transit (TLS), but requests are not yet encrypted from your machine to the serving machine, so the relay can read them. See the Privacy Policy for the retention table and Docs for the full data-flow picture.

What these don't do yet

Not availableWhat to do instead
OpenAI-compatible /v1/audio/transcriptions or /v1/embeddings REST endpoints Use the CLI. A request to POST /v1/messages naming the capability needs the reverse content tunnel, so a plain HTTP client cannot send it; see the API Reference and Connect your app.
Microphone capture and voice-activity detection Record with any tool you like and pass the file. dollama has no client-side capture pipeline.
Streaming audio in — live dictation from an open mic Send whole clips. Transcription progress streams back out, but the input is a complete file; clips are capped at 5 minutes.
Text-to-speech / read-aloud Not implemented. Designed, not built.
Speaker diarization (“who said what”) Not implemented.
Vector storage or search dollama embed returns vectors; storing and searching them is yours. Any vector store works.

Both capabilities depend on nodes choosing to advertise them. If a request fails with no node available, check the capabilities key of GET /v1/network/models for the current pool depth — and consider serving it yourself.