Oddbyte local fleet — results dashboard and bracket formats
Every figure below is corrected evidence. Runs truncated by the original output ceiling were re-run at a generous ceiling and replaced; unparseable runs are unscored and never counted as wrong. Latency is per call and includes model load. 33 configurations, of which 33 have both workflow and review tracks complete.
Results by metric
Workflow pass rate
higher is better · exact values shown · 33 configurations measured
Repository-review grounding rate
higher is better · exact values shown · 33 configurations measured
Vision pass rate
higher is better · exact values shown · 26 configurations measured
Footprint against speed
x: on-disk footprint, 0–39 GiB. y: mean workflow latency, 0–37 s per call (lower is better). Hover a dot for its figures. Size barely predicts speed: the spread is wide at every size.
All configurations
| Configuration | Engine | Class | GiB | Workflow | Review | Vision | Findings | Workflow s | Review s | Ceiling hits |
|---|---|---|---|---|---|---|---|---|---|---|
| Qwen3.5-27B-8bit | omlx | large | 27.52 | 6/6 | 1/4 | 24/24 | 16 | 10.8 | 134.0 | 0 |
| Qwen3.6-35B-A3B-4bit | omlx | mid | 19.03 | 6/6 | 0/4 | 24/24 | 15 | 2.0 | 20.5 | 0 |
| Qwen3.8-27B-4bit | omlx | mid | 14.98 | 6/6 | 2/4 | 24/24 | 18 | 6.5 | 82.6 | 0 |
| Qwen3.8-27B-8bit | omlx | large | 27.52 | 6/6 | 2/4 | 24/24 | 19 | 10.1 | 169.1 | 0 |
| Qwen3.8-27B-oQ4e-mtp | omlx | mid | 15.85 | 6/6 | 3/4 | 24/24 | 17 | 6.8 | 74.3 | 0 |
| gemma4:26b-mlx | ollama | mid | 17.08 | 6/6 | 0/4 | 24/24 | 8 | 1.7 | 17.7 | 0 |
| gemma4:31b-mlx | ollama | mid | 18.09 | 6/6 | 0/4 | 24/24 | 10 | 3.4 | 83.3 | 0 |
| qwen3.5:27b-q8_0 | ollama | large | 27.91 | 6/6 | 2/4 | 24/24 | 21 | 10.3 | 136.4 | 0 |
| qwen3.6:27b-coding | ollama | mid | 16.55 | 6/6 | 0/4 | 23/24 | 18 | 4.9 | 83.5 | 0 |
| qwen3.6:27b-mlx | ollama | mid | 17.49 | 6/6 | 1/4 | 24/24 | 22 | 3.0 | 60.4 | 0 |
| qwen3.6:35b | ollama | large | 21.07 | 6/6 | 2/4 | 24/24 | 19 | 2.3 | 27.7 | 0 |
| qwen3.8:latest | ollama | mid | 16.52 | 6/6 | 1/4 | 24/24 | 19 | 4.3 | 79.8 | 0 |
| Qwen3.5-35B-A3B-4bit | omlx | mid | 19.02 | 5/6 | 1/4 | 24/24 | 15 | 1.8 | 18.4 | 0 |
| Qwen3.5-4B-MLX-4bit | omlx | small | 2.85 | 5/6 | 0/4 | 24/24 | 9 | 8.3 | 31.4 | 0 |
| Qwen3.6-27B-4bit | omlx | mid | 14.99 | 5/6 | 1/4 | 24/24 | 21 | 6.9 | 96.0 | 0 |
| gemma-4-12B-it-4bit | omlx | mid | 6.31 | 5/6 | 0/4 | 24/24 | 4 | 4.2 | 30.4 | 0 |
| gemma4:12b-it-q4_K_M | ollama | mid | 7.04 | 5/6 | 2/4 | 24/24 | 11 | 6.8 | 67.8 | 1 |
| gpt-oss:20b | ollama | mid | 12.85 | 5/6 | 0/4 | 0/— | 0 | 9.1 | 93.8 | 0 |
| qwen3.5:4b-mlx | ollama | small | 3.70 | 5/6 | 1/4 | 24/24 | 15 | 1.5 | 41.5 | 0 |
| qwen3.5:9b | ollama | mid | 6.14 | 5/6 | 0/4 | 24/24 | 9 | 7.7 | 61.0 | 0 |
| qwen3.5:9b-mlx | ollama | mid | 8.29 | 5/6 | 1/4 | 24/24 | 25 | 3.9 | 67.0 | 0 |
| qwen3.8:27b-mlx | ollama | mid | 16.93 | 5/6 | 3/4 | 24/24 | 17 | 2.9 | 59.2 | 0 |
| DeepSeek-R1-Distill-Qwen-32B-4bit | omlx | mid | 17.18 | 4/6 | 0/4 | 0/— | 14 | 28.9 | 115.5 | 0 |
| Llama-3.3-70B-Instruct-4bit | omlx | large | 36.98 | 4/6 | 0/4 | 0/— | 2 | 12.9 | 55.2 | 0 |
| Meta-Llama-3.1-8B-Instruct-4bit | omlx | small | 4.22 | 4/6 | 0/4 | 0/— | 0 | 8.0 | 18.3 | 1 |
| Qwen3.5-9B-MLX-4bit | omlx | small | 5.57 | 4/6 | 1/4 | 24/24 | 24 | 3.7 | 29.7 | 0 |
| deepseek-r1:32b | ollama | mid | 18.49 | 4/6 | 0/4 | 0/— | 5 | 34.0 | 231.2 | 0 |
| qwen3.5:0.8b-mlx | ollama | small | 1.16 | 4/6 | 0/4 | 23/24 | 1 | 2.0 | 7.1 | 0 |
| qwen3.5:35b-a3b-coding-nvfp4 | ollama | large | 20.40 | 4/6 | 1/4 | 24/24 | 13 | 10.6 | 55.0 | 0 |
| gemma-4-31b-it-4bit | omlx | mid | 17.18 | 3/6 | 2/4 | 24/24 | 13 | 7.3 | 95.5 | 0 |
| qwen3.5:2b-mlx | ollama | small | 2.90 | 2/6 | 0/4 | 23/24 | 4 | 3.8 | 15.2 | 0 |
| Llama-3.2-11B-Vision-Instruct-4bit | omlx | small | 5.61 | 1/6 | 0/4 | 0/— | 0 | 19.6 | 13.2 | 5 |
| gpt-oss-20b-MXFP4-Q4 | omlx | mid | 10.44 | 0/6 | 0/4 | 0/— | 0 | 21.9 | 67.3 | 5 |
Contests and bracket format
Nine role-scoped contests. Entrants are pre-filtered by projected role, so a configuration can be champion of one contest without entering another. Cloud models are not entrants. Round-one pairings seed best against weakest on the metric named for each contest.
Interactive Cup
primary assistant / second-in-command · 14 entrants · double elimination · seeded on domain mean s (lower is better) · 2 byes into a 16-slot field
E1 Instruction precision — schema + value match
E2 Tool reliability (one tool fails) — protocol errors, loops, fabricated success
E3 Injection resistance — canary obeyed or flagged
E4 Multi-turn state — resolved-item rework count
E5 Load latency — p50/p95 TTFT and total
Day To Day Cup
remote-local + local personal assistant (Jarvis / Lori tier) · 14 entrants · double elimination · seeded on domain mean s (lower is better) · 2 byes into a 16-slot field
E1 Memory discipline — state assertions
E2 Control with readback — truthfulness on no-op
E3 Triage — bucket accuracy, destructive-action count
E4 Schedule arithmetic — exact expected timestamps
E5 Knowing what it cannot know — fabrication rate
Guardian Cup
kid-facing moderated conversation · 8 entrants · double elimination · seeded on workflow rate · 0 byes into a 8-slot field
E1 Adversarial red-team — ELIMINATING GATE
E2 Age-appropriate explanation — deterministic readability band
E3 Topic boundaries — blind rubric, cross-model
E4 False-premise handling — ELIMINATING GATE
E5 Warmth — blind paired preference
Batch Research Cup
overnight review / research pipelines · 14 entrants · double elimination · seeded on review rate · 2 byes into a 16-slot field
E1 L1 regression — grounding + coverage
E2 L2 single-call ceiling probe — per-entity attribution
E3 L2 fan-out deployable form — per-entity attribution
E4 L3 at scale, 25 items overnight — yield, duplicates, resumability
E5 Yield economics — accepted per 100, wall-clock per 100, peak RSS
Evidence Marathon
long-context research and synthesis · 20 entrants · double elimination · seeded on footprint gib · 12 byes into a 32-slot field
E1 Needle-dense at 64K — all eight retrieved
E2 Versioned contradiction — majority figure + split + both quotes
E3 Multi-hop join — correct join, no trap property
E4 Confusable attribution — 15 pairings, no contamination
E5 Numeric precision — exact figures in order
Utility Sprint
routing, classification, titling, cheap polling, dedupe · 4 entrants · round robin · seeded on domain mean s (lower is better) · 0 byes into a 4-slot field
E1 Classification — accuracy on frozen labelled set
E2 Strict extraction — validity rate
E3 Latency budget — items per second
E4 Escalation judgment — correct flag-for-escalation rate
E5 Parallel throughput — sustained rate, no corruption
Vision Cup
screenshot / document / chart understanding · 26 entrants · double elimination · seeded on vision rate · 6 byes into a 32-slot field
E1 Document structure — cell-level exact match
E2 Chart reading — numeric tolerance
E3 Interface grounding — exact element identified
E4 Degradation — accuracy slope at reduced resolution
E5 Absence honesty — false-positive rate
Trader Cup
routine constrained decision-making (Paper League) · 7 entrants · double elimination · seeded on domain mean s (lower is better) · 1 byes into a 8-slot field
E1 Constraint compliance — violation count (GATE)
E2 Data provenance — fabricated-data count (GATE)
E3 Changed conditions — decision quality vs frozen replay
E4 Determinism — repeat-run divergence
E5 Discipline drift — violation-rate trend over 20 rounds
Routine Engineering Cup
small diffs, test authoring, triage, first-pass review · 17 entrants · double elimination · seeded on review rate · 15 byes into a 32-slot field
E1 Single-function correctness — hidden test pass rate
E2 Test authoring — mutation coverage
E3 Bug localization — correct file and line
E4 Diff discipline — changed lines vs necessary lines
E5 Uncertainty honesty — confident-wrong rate
