Oddbyte local fleet — results dashboard and bracket formats

Every figure below is corrected evidence. Runs truncated by the original output ceiling were re-run at a generous ceiling and replaced; unparseable runs are unscored and never counted as wrong. Latency is per call and includes model load. 33 configurations, of which 33 have both workflow and review tracks complete.

Results by metric

Workflow pass rate

higher is better · exact values shown · 33 configurations measured

Qwen3.5-27B-8bit · omlx
6/6
Qwen3.6-35B-A3B-4bit · omlx
6/6
Qwen3.8-27B-4bit · omlx
6/6
Qwen3.8-27B-8bit · omlx
6/6
Qwen3.8-27B-oQ4e-mtp · omlx
6/6
gemma4:26b-mlx · ollama
6/6
gemma4:31b-mlx · ollama
6/6
qwen3.5:27b-q8_0 · ollama
6/6
qwen3.6:27b-coding · ollama
6/6
qwen3.6:27b-mlx · ollama
6/6
qwen3.6:35b · ollama
6/6
qwen3.8:latest · ollama
6/6
Qwen3.5-35B-A3B-4bit · omlx
5/6
Qwen3.5-4B-MLX-4bit · omlx
5/6
Qwen3.6-27B-4bit · omlx
5/6
gemma-4-12B-it-4bit · omlx
5/6
gemma4:12b-it-q4_K_M · ollama
5/6
gpt-oss:20b · ollama
5/6
qwen3.5:4b-mlx · ollama
5/6
qwen3.5:9b · ollama
5/6
qwen3.5:9b-mlx · ollama
5/6
qwen3.8:27b-mlx · ollama
5/6
DeepSeek-R1-Distill-Qwen-32B-4bit · omlx
4/6
Llama-3.3-70B-Instruct-4bit · omlx
4/6
Meta-Llama-3.1-8B-Instruct-4bit · omlx
4/6
Qwen3.5-9B-MLX-4bit · omlx
4/6
deepseek-r1:32b · ollama
4/6
qwen3.5:0.8b-mlx · ollama
4/6
qwen3.5:35b-a3b-coding-nvfp4 · ollama
4/6
gemma-4-31b-it-4bit · omlx
3/6
qwen3.5:2b-mlx · ollama
2/6
Llama-3.2-11B-Vision-Instruct-4bit · omlx
1/6
gpt-oss-20b-MXFP4-Q4 · omlx
0/6

Repository-review grounding rate

higher is better · exact values shown · 33 configurations measured

Qwen3.8-27B-oQ4e-mtp · omlx
3/4
qwen3.8:27b-mlx · ollama
3/4
Qwen3.8-27B-4bit · omlx
2/4
Qwen3.8-27B-8bit · omlx
2/4
qwen3.5:27b-q8_0 · ollama
2/4
qwen3.6:35b · ollama
2/4
gemma4:12b-it-q4_K_M · ollama
2/4
gemma-4-31b-it-4bit · omlx
2/4
Qwen3.5-27B-8bit · omlx
1/4
qwen3.6:27b-mlx · ollama
1/4
qwen3.8:latest · ollama
1/4
Qwen3.5-35B-A3B-4bit · omlx
1/4
Qwen3.6-27B-4bit · omlx
1/4
qwen3.5:4b-mlx · ollama
1/4
qwen3.5:9b-mlx · ollama
1/4
Qwen3.5-9B-MLX-4bit · omlx
1/4
qwen3.5:35b-a3b-coding-nvfp4 · ollama
1/4
Qwen3.6-35B-A3B-4bit · omlx
0/4
gemma4:26b-mlx · ollama
0/4
gemma4:31b-mlx · ollama
0/4
qwen3.6:27b-coding · ollama
0/4
Qwen3.5-4B-MLX-4bit · omlx
0/4
gemma-4-12B-it-4bit · omlx
0/4
gpt-oss:20b · ollama
0/4
qwen3.5:9b · ollama
0/4
DeepSeek-R1-Distill-Qwen-32B-4bit · omlx
0/4
Llama-3.3-70B-Instruct-4bit · omlx
0/4
Meta-Llama-3.1-8B-Instruct-4bit · omlx
0/4
deepseek-r1:32b · ollama
0/4
qwen3.5:0.8b-mlx · ollama
0/4
qwen3.5:2b-mlx · ollama
0/4
Llama-3.2-11B-Vision-Instruct-4bit · omlx
0/4
gpt-oss-20b-MXFP4-Q4 · omlx
0/4

Vision pass rate

higher is better · exact values shown · 26 configurations measured

Qwen3.5-27B-8bit · omlx
24/24
Qwen3.6-35B-A3B-4bit · omlx
24/24
Qwen3.8-27B-4bit · omlx
24/24
Qwen3.8-27B-8bit · omlx
24/24
Qwen3.8-27B-oQ4e-mtp · omlx
24/24
gemma4:26b-mlx · ollama
24/24
gemma4:31b-mlx · ollama
24/24
qwen3.5:27b-q8_0 · ollama
24/24
qwen3.6:27b-mlx · ollama
24/24
qwen3.6:35b · ollama
24/24
qwen3.8:latest · ollama
24/24
Qwen3.5-35B-A3B-4bit · omlx
24/24
Qwen3.5-4B-MLX-4bit · omlx
24/24
Qwen3.6-27B-4bit · omlx
24/24
gemma-4-12B-it-4bit · omlx
24/24
gemma4:12b-it-q4_K_M · ollama
24/24
qwen3.5:4b-mlx · ollama
24/24
qwen3.5:9b · ollama
24/24
qwen3.5:9b-mlx · ollama
24/24
qwen3.8:27b-mlx · ollama
24/24
Qwen3.5-9B-MLX-4bit · omlx
24/24
qwen3.5:35b-a3b-coding-nvfp4 · ollama
24/24
gemma-4-31b-it-4bit · omlx
24/24
qwen3.6:27b-coding · ollama
23/24
qwen3.5:0.8b-mlx · ollama
23/24
qwen3.5:2b-mlx · ollama
23/24

Footprint against speed

x: on-disk footprint, 0–39 GiB. y: mean workflow latency, 0–37 s per call (lower is better). Hover a dot for its figures. Size barely predicts speed: the spread is wide at every size.

All configurations

Configuration Engine Class GiB Workflow Review Vision Findings Workflow s Review s Ceiling hits
Qwen3.5-27B-8bit omlx large 27.52 6/6 1/4 24/24 16 10.8 134.0 0
Qwen3.6-35B-A3B-4bit omlx mid 19.03 6/6 0/4 24/24 15 2.0 20.5 0
Qwen3.8-27B-4bit omlx mid 14.98 6/6 2/4 24/24 18 6.5 82.6 0
Qwen3.8-27B-8bit omlx large 27.52 6/6 2/4 24/24 19 10.1 169.1 0
Qwen3.8-27B-oQ4e-mtp omlx mid 15.85 6/6 3/4 24/24 17 6.8 74.3 0
gemma4:26b-mlx ollama mid 17.08 6/6 0/4 24/24 8 1.7 17.7 0
gemma4:31b-mlx ollama mid 18.09 6/6 0/4 24/24 10 3.4 83.3 0
qwen3.5:27b-q8_0 ollama large 27.91 6/6 2/4 24/24 21 10.3 136.4 0
qwen3.6:27b-coding ollama mid 16.55 6/6 0/4 23/24 18 4.9 83.5 0
qwen3.6:27b-mlx ollama mid 17.49 6/6 1/4 24/24 22 3.0 60.4 0
qwen3.6:35b ollama large 21.07 6/6 2/4 24/24 19 2.3 27.7 0
qwen3.8:latest ollama mid 16.52 6/6 1/4 24/24 19 4.3 79.8 0
Qwen3.5-35B-A3B-4bit omlx mid 19.02 5/6 1/4 24/24 15 1.8 18.4 0
Qwen3.5-4B-MLX-4bit omlx small 2.85 5/6 0/4 24/24 9 8.3 31.4 0
Qwen3.6-27B-4bit omlx mid 14.99 5/6 1/4 24/24 21 6.9 96.0 0
gemma-4-12B-it-4bit omlx mid 6.31 5/6 0/4 24/24 4 4.2 30.4 0
gemma4:12b-it-q4_K_M ollama mid 7.04 5/6 2/4 24/24 11 6.8 67.8 1
gpt-oss:20b ollama mid 12.85 5/6 0/4 0/— 0 9.1 93.8 0
qwen3.5:4b-mlx ollama small 3.70 5/6 1/4 24/24 15 1.5 41.5 0
qwen3.5:9b ollama mid 6.14 5/6 0/4 24/24 9 7.7 61.0 0
qwen3.5:9b-mlx ollama mid 8.29 5/6 1/4 24/24 25 3.9 67.0 0
qwen3.8:27b-mlx ollama mid 16.93 5/6 3/4 24/24 17 2.9 59.2 0
DeepSeek-R1-Distill-Qwen-32B-4bit omlx mid 17.18 4/6 0/4 0/— 14 28.9 115.5 0
Llama-3.3-70B-Instruct-4bit omlx large 36.98 4/6 0/4 0/— 2 12.9 55.2 0
Meta-Llama-3.1-8B-Instruct-4bit omlx small 4.22 4/6 0/4 0/— 0 8.0 18.3 1
Qwen3.5-9B-MLX-4bit omlx small 5.57 4/6 1/4 24/24 24 3.7 29.7 0
deepseek-r1:32b ollama mid 18.49 4/6 0/4 0/— 5 34.0 231.2 0
qwen3.5:0.8b-mlx ollama small 1.16 4/6 0/4 23/24 1 2.0 7.1 0
qwen3.5:35b-a3b-coding-nvfp4 ollama large 20.40 4/6 1/4 24/24 13 10.6 55.0 0
gemma-4-31b-it-4bit omlx mid 17.18 3/6 2/4 24/24 13 7.3 95.5 0
qwen3.5:2b-mlx ollama small 2.90 2/6 0/4 23/24 4 3.8 15.2 0
Llama-3.2-11B-Vision-Instruct-4bit omlx small 5.61 1/6 0/4 0/— 0 19.6 13.2 5
gpt-oss-20b-MXFP4-Q4 omlx mid 10.44 0/6 0/4 0/— 0 21.9 67.3 5

Contests and bracket format

Nine role-scoped contests. Entrants are pre-filtered by projected role, so a configuration can be champion of one contest without entering another. Cloud models are not entrants. Round-one pairings seed best against weakest on the metric named for each contest.

Interactive Cup

primary assistant / second-in-command · 14 entrants · double elimination · seeded on domain mean s (lower is better) · 2 byes into a 16-slot field

E1 Instruction precision — schema + value match

E2 Tool reliability (one tool fails) — protocol errors, loops, fabricated success

E3 Injection resistance — canary obeyed or flagged

E4 Multi-turn state — resolved-item rework count

E5 Load latency — p50/p95 TTFT and total

qwen3.5:4b-mlx small
vs
Qwen3.5-27B-8bit large
Qwen3.5-35B-A3B-4bit mid
vs
qwen3.5:27b-q8_0 large
qwen3.6:35b large
vs
Qwen3.8-27B-8bit large
qwen3.8:27b-mlx mid
vs
Qwen3.6-27B-4bit mid
qwen3.6:27b-mlx mid
vs
gemma4:12b-it-q4_K_M mid
qwen3.5:9b-mlx mid
vs
Qwen3.8-27B-oQ4e-mtp mid
qwen3.8:latest mid
vs
Qwen3.8-27B-4bit mid

Day To Day Cup

remote-local + local personal assistant (Jarvis / Lori tier) · 14 entrants · double elimination · seeded on domain mean s (lower is better) · 2 byes into a 16-slot field

E1 Memory discipline — state assertions

E2 Control with readback — truthfulness on no-op

E3 Triage — bucket accuracy, destructive-action count

E4 Schedule arithmetic — exact expected timestamps

E5 Knowing what it cannot know — fabrication rate

qwen3.5:4b-mlx small
vs
Qwen3.5-27B-8bit large
Qwen3.5-35B-A3B-4bit mid
vs
qwen3.5:27b-q8_0 large
qwen3.6:35b large
vs
Qwen3.8-27B-8bit large
qwen3.8:27b-mlx mid
vs
Qwen3.6-27B-4bit mid
qwen3.6:27b-mlx mid
vs
gemma4:12b-it-q4_K_M mid
qwen3.5:9b-mlx mid
vs
Qwen3.8-27B-oQ4e-mtp mid
qwen3.8:latest mid
vs
Qwen3.8-27B-4bit mid

Guardian Cup

kid-facing moderated conversation · 8 entrants · double elimination · seeded on workflow rate · 0 byes into a 8-slot field

E1 Adversarial red-team — ELIMINATING GATE

E2 Age-appropriate explanation — deterministic readability band

E3 Topic boundaries — blind rubric, cross-model

E4 False-premise handling — ELIMINATING GATE

E5 Warmth — blind paired preference

Qwen3.5-4B-MLX-4bit small
vs
qwen3.5:0.8b-mlx small
gemma4:12b-it-q4_K_M mid
vs
Qwen3.5-9B-MLX-4bit small
qwen3.5:4b-mlx small
vs
Meta-Llama-3.1-8B-Instruct-4bit small
qwen3.5:9b mid
vs
qwen3.5:9b-mlx mid

Batch Research Cup

overnight review / research pipelines · 14 entrants · double elimination · seeded on review rate · 2 byes into a 16-slot field

E1 L1 regression — grounding + coverage

E2 L2 single-call ceiling probe — per-entity attribution

E3 L2 fan-out deployable form — per-entity attribution

E4 L3 at scale, 25 items overnight — yield, duplicates, resumability

E5 Yield economics — accepted per 100, wall-clock per 100, peak RSS

Qwen3.8-27B-oQ4e-mtp mid
vs
qwen3.8:latest mid
qwen3.8:27b-mlx mid
vs
qwen3.6:27b-mlx mid
Qwen3.8-27B-4bit mid
vs
qwen3.5:9b-mlx mid
Qwen3.8-27B-8bit large
vs
qwen3.5:4b-mlx small
gemma4:12b-it-q4_K_M mid
vs
Qwen3.6-27B-4bit mid
qwen3.5:27b-q8_0 large
vs
Qwen3.5-35B-A3B-4bit mid
qwen3.6:35b large
vs
Qwen3.5-27B-8bit large

Evidence Marathon

long-context research and synthesis · 20 entrants · double elimination · seeded on footprint gib · 12 byes into a 32-slot field

E1 Needle-dense at 64K — all eight retrieved

E2 Versioned contradiction — majority figure + split + both quotes

E3 Multi-hop join — correct join, no trap property

E4 Confusable attribution — 15 pairings, no contamination

E5 Numeric precision — exact figures in order

Llama-3.3-70B-Instruct-4bit large
vs
gpt-oss:20b mid
qwen3.5:27b-q8_0 large
vs
Qwen3.8-27B-4bit mid
Qwen3.5-27B-8bit large
vs
Qwen3.6-27B-4bit mid
Qwen3.8-27B-8bit large
vs
Qwen3.8-27B-oQ4e-mtp mid
qwen3.6:35b large
vs
qwen3.8:latest mid
qwen3.5:35b-a3b-coding-nvfp4 large
vs
qwen3.6:27b-coding mid
Qwen3.6-35B-A3B-4bit mid
vs
qwen3.8:27b-mlx mid
Qwen3.5-35B-A3B-4bit mid
vs
gemma4:26b-mlx mid
deepseek-r1:32b mid
vs
DeepSeek-R1-Distill-Qwen-32B-4bit mid
gemma4:31b-mlx mid
vs
qwen3.6:27b-mlx mid

Utility Sprint

routing, classification, titling, cheap polling, dedupe · 4 entrants · round robin · seeded on domain mean s (lower is better) · 0 byes into a 4-slot field

E1 Classification — accuracy on frozen labelled set

E2 Strict extraction — validity rate

E3 Latency budget — items per second

E4 Escalation judgment — correct flag-for-escalation rate

E5 Parallel throughput — sustained rate, no corruption

qwen3.5:4b-mlx small
vs
Qwen3.5-4B-MLX-4bit small
qwen3.5:0.8b-mlx small
vs
qwen3.5:2b-mlx small

Vision Cup

screenshot / document / chart understanding · 26 entrants · double elimination · seeded on vision rate · 6 byes into a 32-slot field

E1 Document structure — cell-level exact match

E2 Chart reading — numeric tolerance

E3 Interface grounding — exact element identified

E4 Degradation — accuracy slope at reduced resolution

E5 Absence honesty — false-positive rate

Qwen3.5-27B-8bit large
vs
Llama-3.2-11B-Vision-Instruct-4bit small
Qwen3.5-35B-A3B-4bit mid
vs
qwen3.6:27b-coding mid
Qwen3.5-4B-MLX-4bit small
vs
qwen3.5:2b-mlx small
Qwen3.5-9B-MLX-4bit small
vs
qwen3.5:0.8b-mlx small
Qwen3.6-27B-4bit mid
vs
qwen3.8:latest mid
Qwen3.6-35B-A3B-4bit mid
vs
qwen3.8:27b-mlx mid
Qwen3.8-27B-4bit mid
vs
qwen3.6:35b large
Qwen3.8-27B-8bit large
vs
qwen3.6:27b-mlx mid
Qwen3.8-27B-oQ4e-mtp mid
vs
qwen3.5:9b-mlx mid
gemma-4-31b-it-4bit mid
vs
qwen3.5:9b mid
gemma4:12b-it-q4_K_M mid
vs
qwen3.5:4b-mlx small
gemma4:26b-mlx mid
vs
qwen3.5:35b-a3b-coding-nvfp4 large
gemma4:31b-mlx mid
vs
qwen3.5:27b-q8_0 large

Trader Cup

routine constrained decision-making (Paper League) · 7 entrants · double elimination · seeded on domain mean s (lower is better) · 1 byes into a 8-slot field

E1 Constraint compliance — violation count (GATE)

E2 Data provenance — fabricated-data count (GATE)

E3 Changed conditions — decision quality vs frozen replay

E4 Determinism — repeat-run divergence

E5 Discipline drift — violation-rate trend over 20 rounds

gemma4:26b-mlx mid
vs
qwen3.6:27b-coding mid
Qwen3.6-35B-A3B-4bit mid
vs
qwen3.8:latest mid
qwen3.6:35b large
vs
gemma4:31b-mlx mid

Routine Engineering Cup

small diffs, test authoring, triage, first-pass review · 17 entrants · double elimination · seeded on review rate · 15 byes into a 32-slot field

E1 Single-function correctness — hidden test pass rate

E2 Test authoring — mutation coverage

E3 Bug localization — correct file and line

E4 Diff discipline — changed lines vs necessary lines

E5 Uncertainty honesty — confident-wrong rate

Qwen3.8-27B-oQ4e-mtp mid
vs
qwen3.8:latest mid
qwen3.8:27b-mlx mid
vs
qwen3.6:27b-mlx mid
Qwen3.8-27B-4bit mid
vs
qwen3.5:9b-mlx mid
Qwen3.8-27B-8bit large
vs
qwen3.5:4b-mlx small
gemma-4-31b-it-4bit mid
vs
qwen3.5:35b-a3b-coding-nvfp4 large
gemma4:12b-it-q4_K_M mid
vs
Qwen3.6-27B-4bit mid
qwen3.5:27b-q8_0 large
vs
Qwen3.5-9B-MLX-4bit small
qwen3.6:35b large
vs
Qwen3.5-35B-A3B-4bit mid

Similar Posts