dollama · model field report · archive · measured jul 2026

Archived evaluations: the July 2026 board.

The ten candidates we graded in July 2026, how they behaved under load, what contributor machines could run, and the full-tier results behind each role. This is a record, not a recommendation: for today’s models and results, see the Models page.

01

How the July board was run

July 2026 · superseded

Every model gets onto — or off — the network the same way: a broad screen, a full-tier stress test for any role that's up for a decision, and a hand-read of the transcripts. The same gate applies whether a model is new to the board or defending a slot it already holds.

1Screen broadly

Quick tier, every candidate, all three roles — 6 scenarios × 1 repeat on a reference RTX 3060. Cheap and deliberately coarse: scores land on a 1/6 grid, so a tie at 100 means "in contention", not "proven equal".

2Stress-test the shortlist

Full tier: 3 repeats, harder scenarios (including a 338-tool catalogue), real confidence intervals. The quick-tier ceilings break apart and the contenders separate.

3Read the transcripts

Numbers can't see a hollow pass or a format fumble. Every shortlisted run is read by hand; where scores tie, substance and cost per request decide.

Capability, native, and the gate between them

Every graded run ends one of three ways: native (solved cleanly, first shot), recovered (the model stumbled, dollama's retry/recovery loop re-prompted it to a good answer), or failed. Capability counts success with the loop allowed; native counts first-shot passes only.

The recovery loop is real dollama value — qwen3.5:4b as a worker passes 62% first-shot yet grades 100 capability with the loop behind it. To stop that from flattering fragile models, Recommended requires the 95% (Wilson) lower bound of the native score to clear the role's bar. A model the pipeline merely rescues can't qualify.

One caveat on the native figures below. The gate described here is sound, but until 2026-08-25 it was fed by a harness that credited the proxy’s own status notices as the model’s answer — so runs that ran out of turns without answering counted as first-shot passes. Native scores on this board are therefore upper bounds; see the note under the table.

02

The candidates

July 2026 · superseded

Ten models graded across dollama's three roles — planner, worker, helper. Each cell is the quick-tier capability score (0–100) with the first-shot native score beneath, coloured by verdict. Speed and memory are measured on a 12 GB RTX 3060: VRAM runs base (weights, at 4k) → peak (at the model's max context, mostly KV); usable context is the inflection point where throughput starts to fall (models survive further, slower).

July 2026 board — superseded. These scores predate a harness fix; read native as an upper bound

Every number on this board was measured before 2026-08-25. Until then the eval harness counted its own [dollama] status notices as the model’s final answer, so a run that used up its turn budget while still calling tools — never actually answering — was scored as a pass. Measured across 24 runs on the day of the fix, 14 of them (58%) ended with no model text at all.

The effect is one-directional: it flatters, never penalises. Treat every native score below as a ceiling. The three models the network actually runs have been measured again since; those results are on the Models page. The other models on this board have not.

The numbers are left as they were measured rather than quietly edited.

July 2026 · supersededCandidates board · quick tier · verdicts are July’s, not today’s
Model qualityPlanner qualityWorker qualityHelper toolsBasics measuredSpeed t/s measuredVRAM base→peak measuredUsable ctx
Jul: Recommended Jul: Marginal Jul: Below bar † score held down by tool-format handicap, not capability — see the note below

Where the scores lie — in both directions

A binary tool-call grader misleads both ways, which is why step 3 of the method exists. False negatives: the phi4 family (the † rows) reasons well but narrates <note> tags or declines the call instead of emitting Read() — those scores measure format adherence, not capability. False positives: correct tool plumbing can pass a task on hollow reasoning — string checks matching the model's own tool output rather than any answer it wrote — and we've caught that in high scorers, not just stragglers (a grader fix is on the list). No grade on this page stands alone: the "read by hand" columns in §04 carry as much weight as the numbers.

How they behave under load July 2026 · superseded
context length · log scale

Throughput and memory as the context window grows, on the reference RTX 3060 (July 2026). The dot on each line marks the inflection point — where speed starts dropping off. Tap a model to isolate it.

03

What contributor machines can run

July 2026 · superseded

Quality is only half the decision — a network model has to run well on the machines people actually contribute. Today's fleet is a handful of hand-built machines, and it sets the second constraint on every pick.

The honest headline: on the three GPU machines in the fleet, almost every model in the board should run — the differentiator isn't whether it fits but how much context and how fast. qwen3.5:9b and gemma4:12b both peak near 7.6–7.8 GB and hold a 64k usable window (128k survivable) on the 12 GB and 16 GB machines; on the 10 GB 3080 they run comfortably at a trimmed context. Only the 12 GB RTX 3060 was benchmarked for this; the figures for the other two are estimates from its measurements.

The lone exception is phi4:14b: it climbs to ~10 GB and caps out near a 16k window, so it fits the 3060 and the M1 but crowds the 3080. That memory ceiling — not raw capability — is exactly the kind of thing that keeps a model off the standard fleet.

So we lean on models that are both strong and comfortable on mid-range contributor cards. As bigger GPUs join, the heavier tiers open up on their own.

Which machines can run which roles July 2026 · superseded
speed = helper model (qwen3.5:4b) t/s

Each machine plotted by its VRAM and how fast it runs the universal helper model (qwen3.5:4b). Filled dot = measured; hollow = estimated from GPU class — only the RTX 3060 is directly benchmarked. A role needs both enough VRAM for its model (helper ~3 GB, worker ~6 GB, planner ~8 GB) and enough speed to serve (~12 t/s) — the bands are rough. All three GPU machines clear every role (two of them by estimate); the CPU/iGPU stragglers can't hold a model at all (they help with speech-to-text instead).

04

Full-tier results by role

July 2026 · superseded

The July full-tier tables that decided each role’s slot, with the reading notes written at the time. Their native scores are upper bounds. Rows for the three results that have since been re-measured carry a link to the current figure.

The planner slot July 2026 · superseded
July 2026 · supersededPlanner candidates · full tier · verdicts are July’s, not today’s
Planner candidate CapabilityNative · 95% CIVerdict Plan quality — read by hand TokensTime
qwen3.5:4bRe-measured Sep 2026 as planner: 86% passed, 79% first-shot →8579% · 62–89Recommended Substantive — concrete plans with correct line citations84144s
ornith:9b7882% · 66–91Recommended Inflated — 2 of 5 "passes" are empty answers that cleared string checks on tool output alone73975s
gemma4:12b7679% · 62–89Marginal Correct but verbose — 1.8× the tokens, and weak on the anti-stall scenario1528113s
qwen3.5:9b7479% · 62–89Recommended Mostly real, but misses the profile-planning task entirely (0/3)96262s
Full tier · 3 repeats · curated planner scenarios · RTX 3060. Tokens & time are per plan over passing runs; turn counts were ~5 for all four.

The planner: qwen3.5:4b July 2026 · superseded

It wins on all three axes at once. Highest capability, genuinely substantive plans (we read them — precise and correctly cited), and the most efficient by a wide margin: a good plan in 841 tokens / 44s against gemma4:12b's 1528 / 113s, for a step that runs on every request. ornith's near-identical score is a mirage — two of its "passes" are empty answers that satisfied the string checks from tool output alone (the false-positive mode from §02), so its real plan quality sits below the number. gemma4:12b is thorough but Marginal here, and pays for its depth in tokens and seconds. With turn counts essentially equal, the honest separators are substance and cost — and qwen3.5:4b takes both.

The worker slot July 2026 · superseded
July 2026 · supersededWorker candidates · full tier · verdicts are July’s, not today’s
Worker candidate CapabilityNative · 95% CIVerdict Behaviour — read by hand Tokens
qwen3.5:9bRe-measured Sep 2026 as worker: 96% passed, 90% first-shot →10095% · 77–99Recommended Fast, native-clean, and finds the task-relevant tool 3/3 first-shot in a 338-tool catalogue. The leanest of the five. Our worker pick.103
gemma4:12b10095% · 77–99Recommended Matches qwen3.5:9b on quality, needle included 3/3. The cost is verbosity — ~2.5× the tokens — and it's the slowest of the shortlist.264
ornith:9b9081% · 60–92Marginal Zero-recovery on everyday work and 3/3 on the needle — its weak-tool-calls reputation is a planner trait, not a worker one. Held back by the rest of the suite, not the catalogue.185
qwen3.5:4b10062% · 41–79Marginal Fastest and nearly the leanest, but the most recovery-dependent: it emits a tool call then falls silent, and dollama's retry loop lands it. Its 62% first-shot is exactly the gap the native gate exists to catch.113
gemma4:e4b7662% · 41–79Marginal The noisiest model we grade: 95 → 86 → 86 → 76 across four runs on byte-identical inputs, so its point score means little. Verbose.332
Full tier · 3 repeats · curated worker scenarios + a 338-tool "needle" retrieval · RTX 3060. Tokens are mean output per response.

The worker: qwen3.5:9b July 2026 · superseded

Two models clear Recommended at identical quality — 100 capability, 95% first-shot, CI 77–99 — so the slot is decided on cost-to-serve: qwen3.5:9b runs ~50 t/s to gemma4:12b's ~34 (about 47% faster) and emits ~2.5× fewer tokens per response, at a similar ~8 GB footprint. On a volunteer network, throughput is served requests. One footnote the transcripts earned: the needle scenario originally failed across the board because of a tool-presentation bug in our pipeline — the one relevant tool was buried in a recovery menu — and the numbers above are against the corrected pipeline. Reading failures before blaming models is standing policy here; more than once the failure has been ours.

The helper: qwen3.5:4b July 2026 · superseded

Re-measured Sep 2026 as helper: 100% passed, 89% first-shot →All three finalists clear the helper bar at full tier (qwen3.5:9b 100, qwen3.5:4b 98, gemma4:12b 90), so the helper is chosen for the fleet, not the leaderboard: qwen3.5:4b is already the planner model — one pulled artifact serves two roles — it fits 2 GB-class machines and CPU boxes that could never hold the 9B, and at ~76 t/s it's the fastest of the three. Helper work (curation, summarisation, classification) also tolerates the recovery loop far better than worker-grade code editing does, so its retry-dependence costs little here.

05

Measurement notes

July 2026 · superseded
  • Quality: dollama eval --role {planner|worker|helper} --tier quick for the §02 board — 6 scenarios × 1 repeat, single RTX 3060, no error bars. Shortlists re-run at --tier full (3 repeats, Wilson confidence intervals) for the §04 tables.
  • Speed / memory: throughput and the VRAM/context curve from dollama benchmark on an RTX 3060 (12 GB); VRAM in binary GiB (MB / 1024). Base = weights at 4k; peak = at the model's max context. Your numbers scale with your card.
  • Usable context is the benchmark's inflection point (where t/s falls off); models survive to a higher context before running out of memory.
  • More context isn't always better. A bigger window relaxes the pipeline's budget pressure, so it stops trimming the tool catalog — and small models handle a lean, relevant toolset better than a bloated one. We measured a worker score higher on ordinary edits with a trimmed catalog than with a roomy one; dollama now keeps the catalog lean regardless of window size. One more reason usable, not survivable, is the number that matters.
  • † phi4 family: scores reflect text-tool-format adherence, not reasoning — see the note in §02.
  • phi4-reasoning:14b is excluded: at ~16 t/s its thinking traces peg the GPU long enough to hard-lock the test box.