dollama · model field report · archive · measured jul 2026
Archived evaluations: the July 2026 board.
The ten candidates we graded in July 2026, how they behaved under load, what contributor machines could run, and the full-tier results behind each role. This is a record, not a recommendation: for today’s models and results, see the Models page.
How the July board was run
July 2026 · supersededEvery model gets onto — or off — the network the same way: a broad screen, a full-tier stress test for any role that's up for a decision, and a hand-read of the transcripts. The same gate applies whether a model is new to the board or defending a slot it already holds.
Quick tier, every candidate, all three roles — 6 scenarios × 1 repeat on a reference RTX 3060. Cheap and deliberately coarse: scores land on a 1/6 grid, so a tie at 100 means "in contention", not "proven equal".
Full tier: 3 repeats, harder scenarios (including a 338-tool catalogue), real confidence intervals. The quick-tier ceilings break apart and the contenders separate.
Numbers can't see a hollow pass or a format fumble. Every shortlisted run is read by hand; where scores tie, substance and cost per request decide.
Capability, native, and the gate between them
Every graded run ends one of three ways: native (solved cleanly, first shot), recovered (the model stumbled, dollama's retry/recovery loop re-prompted it to a good answer), or failed. Capability counts success with the loop allowed; native counts first-shot passes only.
The recovery loop is real dollama value — qwen3.5:4b as a worker passes 62% first-shot yet grades 100 capability with the loop behind it. To stop that from flattering fragile models, Recommended requires the 95% (Wilson) lower bound of the native score to clear the role's bar. A model the pipeline merely rescues can't qualify.
One caveat on the native figures below. The gate described here is sound, but until 2026-08-25 it was fed by a harness that credited the proxy’s own status notices as the model’s answer — so runs that ran out of turns without answering counted as first-shot passes. Native scores on this board are therefore upper bounds; see the note under the table.
The candidates
July 2026 · supersededTen models graded across dollama's three roles — planner, worker, helper. Each cell is the quick-tier capability score (0–100) with the first-shot native score beneath, coloured by verdict. Speed and memory are measured on a 12 GB RTX 3060: VRAM runs base (weights, at 4k) → peak (at the model's max context, mostly KV); usable context is the inflection point where throughput starts to fall (models survive further, slower).
July 2026 board — superseded. These scores predate a harness fix; read native as an upper bound
Every number on this board was measured before 2026-08-25. Until then the eval harness counted its own [dollama] status notices as the model’s final answer, so a run that used up its turn budget while still calling tools — never actually answering — was scored as a pass. Measured across 24 runs on the day of the fix, 14 of them (58%) ended with no model text at all.
The effect is one-directional: it flatters, never penalises. Treat every native score below as a ceiling. The three models the network actually runs have been measured again since; those results are on the Models page. The other models on this board have not.
The numbers are left as they were measured rather than quietly edited.
| Model | qualityPlanner | qualityWorker | qualityHelper | toolsBasics | measuredSpeed t/s | measuredVRAM base→peak | measuredUsable ctx |
|---|
Where the scores lie — in both directions
A binary tool-call grader misleads both ways, which is why step 3 of the method exists. False negatives: the phi4 family (the † rows) reasons well but narrates <note> tags or declines the call instead of emitting Read() — those scores measure format adherence, not capability. False positives: correct tool plumbing can pass a task on hollow reasoning — string checks matching the model's own tool output rather than any answer it wrote — and we've caught that in high scorers, not just stragglers (a grader fix is on the list). No grade on this page stands alone: the "read by hand" columns in §04 carry as much weight as the numbers.
Throughput and memory as the context window grows, on the reference RTX 3060 (July 2026). The dot on each line marks the inflection point — where speed starts dropping off. Tap a model to isolate it.
What contributor machines can run
July 2026 · supersededQuality is only half the decision — a network model has to run well on the machines people actually contribute. Today's fleet is a handful of hand-built machines, and it sets the second constraint on every pick.
The honest headline: on the three GPU machines in the fleet, almost every model in the board should run — the differentiator isn't whether it fits but how much context and how fast. qwen3.5:9b and gemma4:12b both peak near 7.6–7.8 GB and hold a 64k usable window (128k survivable) on the 12 GB and 16 GB machines; on the 10 GB 3080 they run comfortably at a trimmed context. Only the 12 GB RTX 3060 was benchmarked for this; the figures for the other two are estimates from its measurements.
The lone exception is phi4:14b: it climbs to ~10 GB and caps out near a 16k window, so it fits the 3060 and the M1 but crowds the 3080. That memory ceiling — not raw capability — is exactly the kind of thing that keeps a model off the standard fleet.
So we lean on models that are both strong and comfortable on mid-range contributor cards. As bigger GPUs join, the heavier tiers open up on their own.
Each machine plotted by its VRAM and how fast it runs the universal helper model (qwen3.5:4b). Filled dot = measured; hollow = estimated from GPU class — only the RTX 3060 is directly benchmarked. A role needs both enough VRAM for its model (helper ~3 GB, worker ~6 GB, planner ~8 GB) and enough speed to serve (~12 t/s) — the bands are rough. All three GPU machines clear every role (two of them by estimate); the CPU/iGPU stragglers can't hold a model at all (they help with speech-to-text instead).
Full-tier results by role
July 2026 · supersededThe July full-tier tables that decided each role’s slot, with the reading notes written at the time. Their native scores are upper bounds. Rows for the three results that have since been re-measured carry a link to the current figure.
| Planner candidate | Capability | Native · 95% CI | Verdict | Plan quality — read by hand | Tokens | Time |
|---|---|---|---|---|---|---|
| qwen3.5:4bRe-measured Sep 2026 as planner: 86% passed, 79% first-shot → | 85 | 79% · 62–89 | Recommended | Substantive — concrete plans with correct line citations | 841 | 44s |
| ornith:9b | 78 | 82% · 66–91 | Recommended | Inflated — 2 of 5 "passes" are empty answers that cleared string checks on tool output alone | 739 | 75s |
| gemma4:12b | 76 | 79% · 62–89 | Marginal | Correct but verbose — 1.8× the tokens, and weak on the anti-stall scenario | 1528 | 113s |
| qwen3.5:9b | 74 | 79% · 62–89 | Recommended | Mostly real, but misses the profile-planning task entirely (0/3) | 962 | 62s |
The planner: qwen3.5:4b July 2026 · superseded
It wins on all three axes at once. Highest capability, genuinely substantive plans (we read them — precise and correctly cited), and the most efficient by a wide margin: a good plan in 841 tokens / 44s against gemma4:12b's 1528 / 113s, for a step that runs on every request. ornith's near-identical score is a mirage — two of its "passes" are empty answers that satisfied the string checks from tool output alone (the false-positive mode from §02), so its real plan quality sits below the number. gemma4:12b is thorough but Marginal here, and pays for its depth in tokens and seconds. With turn counts essentially equal, the honest separators are substance and cost — and qwen3.5:4b takes both.
| Worker candidate | Capability | Native · 95% CI | Verdict | Behaviour — read by hand | Tokens |
|---|---|---|---|---|---|
| qwen3.5:9bRe-measured Sep 2026 as worker: 96% passed, 90% first-shot → | 100 | 95% · 77–99 | Recommended | Fast, native-clean, and finds the task-relevant tool 3/3 first-shot in a 338-tool catalogue. The leanest of the five. Our worker pick. | 103 |
| gemma4:12b | 100 | 95% · 77–99 | Recommended | Matches qwen3.5:9b on quality, needle included 3/3. The cost is verbosity — ~2.5× the tokens — and it's the slowest of the shortlist. | 264 |
| ornith:9b | 90 | 81% · 60–92 | Marginal | Zero-recovery on everyday work and 3/3 on the needle — its weak-tool-calls reputation is a planner trait, not a worker one. Held back by the rest of the suite, not the catalogue. | 185 |
| qwen3.5:4b | 100 | 62% · 41–79 | Marginal | Fastest and nearly the leanest, but the most recovery-dependent: it emits a tool call then falls silent, and dollama's retry loop lands it. Its 62% first-shot is exactly the gap the native gate exists to catch. | 113 |
| gemma4:e4b | 76 | 62% · 41–79 | Marginal | The noisiest model we grade: 95 → 86 → 86 → 76 across four runs on byte-identical inputs, so its point score means little. Verbose. | 332 |
The worker: qwen3.5:9b July 2026 · superseded
Two models clear Recommended at identical quality — 100 capability, 95% first-shot, CI 77–99 — so the slot is decided on cost-to-serve: qwen3.5:9b runs ~50 t/s to gemma4:12b's ~34 (about 47% faster) and emits ~2.5× fewer tokens per response, at a similar ~8 GB footprint. On a volunteer network, throughput is served requests. One footnote the transcripts earned: the needle scenario originally failed across the board because of a tool-presentation bug in our pipeline — the one relevant tool was buried in a recovery menu — and the numbers above are against the corrected pipeline. Reading failures before blaming models is standing policy here; more than once the failure has been ours.
The helper: qwen3.5:4b July 2026 · superseded
Re-measured Sep 2026 as helper: 100% passed, 89% first-shot →All three finalists clear the helper bar at full tier (qwen3.5:9b 100, qwen3.5:4b 98, gemma4:12b 90), so the helper is chosen for the fleet, not the leaderboard: qwen3.5:4b is already the planner model — one pulled artifact serves two roles — it fits 2 GB-class machines and CPU boxes that could never hold the 9B, and at ~76 t/s it's the fastest of the three. Helper work (curation, summarisation, classification) also tolerates the recovery loop far better than worker-grade code editing does, so its retry-dependence costs little here.
Measurement notes
July 2026 · superseded- Quality:
dollama eval --role {planner|worker|helper} --tier quickfor the §02 board — 6 scenarios × 1 repeat, single RTX 3060, no error bars. Shortlists re-run at--tier full(3 repeats, Wilson confidence intervals) for the §04 tables. - Speed / memory: throughput and the VRAM/context curve from
dollama benchmarkon an RTX 3060 (12 GB); VRAM in binary GiB (MB / 1024). Base = weights at 4k; peak = at the model's max context. Your numbers scale with your card. - Usable context is the benchmark's inflection point (where t/s falls off); models survive to a higher context before running out of memory.
- More context isn't always better. A bigger window relaxes the pipeline's budget pressure, so it stops trimming the tool catalog — and small models handle a lean, relevant toolset better than a bloated one. We measured a worker score higher on ordinary edits with a trimmed catalog than with a roomy one; dollama now keeps the catalog lean regardless of window size. One more reason usable, not survivable, is the number that matters.
- † phi4 family: scores reflect text-tool-format adherence, not reasoning — see the note in §02.
- phi4-reasoning:14b is excluded: at ~16 t/s its thinking traces peg the GPU long enough to hard-lock the test box.