dollama · models, benchmarking & evaluation
What can your machine run, and how do you find out?
dollama runs a small set of supported open-weight models. After benchmarking a contributing machine, the relay assigns it one language-model serving role. This page says what runs today, how to test a model yourself, and what the published results show.
What runs today
| Model | Role | Who can use it | Status |
|---|---|---|---|
qwen3.5:9b | Worker: the main reasoning, tool calls and streamed answer. | Every mode. Private: your own machines. Group: also machines of members of your groups. Open Network: also volunteers’ machines. | In use |
qwen3.5:4b | Helper (summaries, trimming context, classifier calls, sub-agents) and planner (writes a plan when one is called for; see below). One model, two roles. | Every mode, on the same machines as above. | In use |
qwen3.6:35b-a3b | Candidate worker. | Not offered on the network. | Pending evaluationOne of three needed submissions; see section 05. |
bonsai:27b | Your own machine only. Runs on PrismML’s llama.cpp, declared in config.toml. | You and your other machines. Never offered to anyone else, and it takes no network role. | ExperimentalLinux and macOS only. Ollama cannot load it. |
After you benchmark, the relay gives your machine a role: worker, planner or helper. The machine then answers that kind of request. The relay compares the machine’s measured speed and context window with a minimum for each role; at the time of writing the worker needs 15 tokens a second and a 32,768-token context window, and the helper 20 tokens a second. The current minimums are at /v1/network/models. A machine that clears none still uses the network as a client.
Role assignment in detail
- A machine has one language-model role. Other capabilities (speech-to-text, embeddings, personal hosting) can be hosted on the same machine and are not roles; see Capabilities.
- A model declared for your own machines only, such as
bonsai:27b, takes no network role. - A machine with memory to spare can also keep the other tier’s model loaded and be picked for that tier too. Its assigned role does not change, and it still serves one request at a time.
When dollama plans
By default (planning = "auto" in config.toml) dollama scores each turn and writes a plan only for turns that score as complex. On the serving machine the plan is normally written by the model already loaded there, unless the relay supplied one. Planning can be turned off; the CLI reference lists the options.
Benchmarking versus evaluation
| Benchmark | Evaluation | |
|---|---|---|
| Question | How well does this model run on my machine? Node performance. | How well does this model do defined tasks through dollama? Model capability. |
| Measures | Generation speed, prompt-processing speed, the largest context window the model can load, the window it holds at a comfortable speed, and GPU memory used. This decides the role your machine gets. | A fixed set of coding-agent scenarios (finding files, choosing tools, editing code, knowing when to stop) run through dollama’s pipeline and graded by checks. It reports how many runs passed and how many passed first-shot (with no retry). It runs at the context window your benchmark measured. |
Whether the network accepts a model for a role is a separate decision, shown in the table in section 01.
Run a benchmark or evaluation
You need Ollama running and room on disk for the model. Stop other programs that hold GPU memory first: otherwise part of the model can be offloaded to the CPU, and the benchmark measures that slower setup. All flags are in the CLI reference.
-
Benchmark the network’s models
dollama benchmarkBenchmarks the network’s models your hardware can plausibly run (helper, then worker, then planner) and updates your machine’s role. Asks before downloading a missing model, or add
--pull. Time depends on your hardware.Saves
~/.dollama/benchmarks/<model>.jsonand~/.dollama/classification.json. -
Benchmark any one model
dollama benchmark --model <name>Does not change your role.
Saves
~/.dollama/benchmarks/adhoc/<model>.json. -
Evaluate a model on one role
dollama eval --model qwen3.5:9b --role workerRuns one role’s scenarios, offering to download the model and to run a context benchmark if needed.
--tierchoosesquick,standard(the default) orfull; the evaluation spec estimates about 8, 15 and 45 minutes a role, but your hardware sets the real time. Nothing is uploaded unless you add--submit.Saves
~/.dollama/eval-network.json,~/.dollama/eval-history.jsonland transcripts under~/.dollama/eval-transcripts/.
Benchmarking also happens by itself. dollama setup and starting dollama network benchmark a machine that has no results. A contributing machine then re-benchmarks, in a quiet moment (nothing served for five minutes), when a new model is published, the method changes or the relay asks; it checks every six hours. DOLLAMA_AUTO_REBENCHMARK=0 turns that off. A result that is merely 30 days old only prompts a notice. Evaluation never runs by itself unless you opt in (section 04).
What happens to the results
| Benchmark | Evaluation you run | |
|---|---|---|
| What is sent | Performance numbers for each model, and a description of your machine. No prompts or transcripts. | A performance summary: pass and first-shot counts, timings and the suite version. No prompts or transcripts; those stay on your machine. |
| Optional? | No. Sent automatically each time your machine connects to the relay as a contributor. | Yes. Sent only with --submit (you must be logged in with dollama login). |
| How it is used | Decides your machine’s role and routing to it. | Informs which models the network accepts. |
Exactly what is sent
Benchmark. For each model: its role, generation and prompt speed, context sizes, the context-versus-speed curve, when it was measured, and the runtime that served it. About your machine: CPU and GPU names, memory sizes, a hashed hardware fingerprint and the hostname. Benchmarks of other models you ran by hand are reported too, flagged so they are never used for routing.
Evaluation. Model, role, suite version, pass and first-shot counts as fractions, a tool-use basics score, the same per scenario (scenario name, class and the names of failed checks), counts per capability, timing figures (time to first token, tokens a second, tokens and time per solved task), the profile and scenario-set hashes, hardware lane, CLI version and the runtime identity. Quick-tier and diagnostic runs are refused. The transcripts stay in ~/.dollama/eval-transcripts/.
How evaluation results count. Your submission is stored as an ordinary result. It counts toward quorum: three accounts, each at least seven days old, running the identical test set and configuration, with first-shot scores within ten points of each other, and the model must be runnable by enough of the fleet. Nothing is then promoted automatically; the relay operator sets which model fills each role. Results recorded by a dollama maintainer account are the baselines in section 05.
Opt-in: letting the network evaluate a model on your idle machine
The relay operator can nominate a candidate model and ask opted-in machines to evaluate it. This is off by default: set eval_assist_enabled = true or use the dashboard switch. A machine that has not opted in is not affected.
- When it runs: overnight only (22:00 to 06:00 local) unless you set
eval_assist_window = "idle". It waits while you are serving requests or the machine is under load, and until the disk has room for the model plus 2 GB. It refuses a model too big for your video memory plus 75% of your RAM, or outside youreval_assist_allowlist. - What it does: downloads the model if you lack it, checks its digest when the relay gave one, runs
dollama evalas a separate process, and uploads the score and the transcripts (at most 8 MiB). The transcripts are of dollama’s built-in test scenarios, not your own conversations or files. They also record your machine’s hardware, operating system version and resource readings during the run, and the relay keeps them so the result can be checked. - Afterwards: it deletes a model only if it downloaded it for this job. A model you already had is never deleted.
Published evaluation results
All three rows are maintainer-recorded: a dollama maintainer account recorded the result and nobody independent has validated it. Passed: runs that ended correctly, recovery steps allowed. First-shot: runs that passed with no retry; its 95% interval is the range the true rate could plausibly fall in, and some rows give only the lower end.
These results are from an earlier suite
The runs were made on 2026-09-01 and 2026-09-02 with suite version 7; the suite is now version 12. On 2026-09-29 dollama began keeping the model’s own reasoning between tool calls by default. The current default pipeline has not been re-measured. The rows differ in date and run count, so do not compare one with another.
| Model · role | Suite | Measured | Runs | Hardware | Passed | First-shot · 95% interval | Verdict at test date |
|---|---|---|---|---|---|---|---|
| qwen3.5:9bWorker | v7 | 2026-09-02 | 84 | RTX 3080 | 96% | 90% · 82–95 | Recommended |
| qwen3.5:4bPlanner | v7 | 2026-09-01 | 66 | RTX 3080 | 86% | 79% · lower bound 67 | Recommended |
| qwen3.5:4bHelper | v7 | 2026-09-01 | 36 | RTX 3080 | 100% | 89% · lower bound 75 | Recommended |
Pending evaluation record: qwen3.6:35b-a3b
| Submissions | First-shot | Runnable by the fleet | On the network |
|---|---|---|---|
| 1 of the 3 independent ones needed | 36 of 42 | No machine can currently run it | Not offered |
Where the records live. The relay holds them, and two public endpoints return JSON for one role at a time: /v1/eval/baseline?role=worker (the maintainer-recorded baseline) and /v1/eval/candidates?role=worker (submissions, quorum and feasibility per candidate). There is no browsable page yet.
Method
Passed, first-shot and the verdict
Every graded run ends one of three ways: it passed first-shot, it passed after dollama’s recovery loop re-prompted the model, or it failed. Passed includes both kinds of pass. First-shot counts only the first, so passes that needed a retry do not raise it.
A model is Recommended for a role when its pass rate clears the role’s bar and the lower end of its first-shot 95% (Wilson) interval clears a second bar. The pass rate itself is judged on its point value, not its lower end.
| Role | Pass rate at least | First-shot lower bound at least |
|---|---|---|
| Worker | 85% | 60% |
| Planner | 80% | 50% |
| Helper | 95% | 70% |
The worker’s 84 runs are 7 scenarios by 12 repeats. A 12-repeat floor for maintainer-recorded runs was decided on 2026-09-02, after the planner and helper rows (fewer runs) were measured. Since suite v7, repairs that need no second model call do not count against first-shot, so it is not a measure of the unaided model.
Measured on
- The three maintainer-recorded rows in section 05 were run on suite v7, on one remote machine with an RTX 3080, through Ollama and dollama’s full pipeline.
- All three maintainer-recorded records are for the default hardware build (4-bit GGUF). None exists for the Apple (MLX) or Blackwell (FP4) builds.
- A harness fault, found and fixed in late August 2026, had counted dollama’s own status notices as model answers. Results from before the fix, including the July evaluation, overstate first-shot scores.
When results are re-run
A score belongs to the model, the pipeline and the suite together, so results only compare within one suite version. Each suite bump needs the current models re-baselined, and a change to default pipeline behaviour can move a score too. The suite has moved from version 7 to 12 and no maintainer-recorded re-run exists yet; that re-run is the next measurement due. New models are screened with the same harness and need quorum before they can be considered.
Archive. The July 2026 evaluation of ten candidate models (throughput and memory charts, full-tier tables) is kept on the archive page. It predates the late-August harness fix, so its first-shot scores are upper bounds. It is a record, not a recommendation.