Qwen 3.5 27B Q8 passed the workflows, then made me check everything
Ollama Qwen 3.5 27B Q8 passed all six scorer-v2 workflows, but ten of its 24 code-review findings were false or materially wrong after source inspection.
Ollama Qwen 3.5 27B Q8 passed all six scorer-v2 workflows, but ten of its 24 code-review findings were false or materially wrong after source inspection.
Ollama Qwen 3.6 35B passed all six scorer-v2 workflows at exceptional speed, then split 26 code-review findings into six useful, seven qualified, and thirteen false or materially wrong.
The Ollama GPT-OSS 20B path completed ten calls but exposed empty or unparsable visible output, leaving zero workflow passes and no review findings to adjudicate.
Ollama Gemma 4 12B passed all six scorer-v2 workflow cases, then mixed four useful code-review findings with nine that needed qualification or rejection.
A real-workflow qualification of Ollama Qwen 3.5 0.8B: three of six bounded cases passed, and all four code-review recommendations were false or materially wrong.
A real-workflow qualification of oMLX Llama 3.3 70B: four of six bounded cases passed, but request failures and invalid output prevented a fair repository-review judgment.
A real-workflow qualification of oMLX Qwen 3.6 35B-A3B: four domain passes out of six, excellent latency, and useful but supervision-heavy review output.
A real-workflow qualification of oMLX Qwen 3.5 35B-A3B: exceptional local speed, three domain passes out of six, and weak broad-review precision.
A real-workflow qualification of oMLX Qwen 3.5 9B: very fast short calls, two domain passes out of six, and too many confident code-review errors.
A real-workflow qualification of oMLX Qwen 3.5 4B: fast bounded work, repetitive repository review, and an 82.8 percent false-or-materially-wrong rate after source adjudication.