Qwen 3.5 4B was quick, but its code review got stuck in a loop
A real-workflow qualification of oMLX Qwen 3.5 4B: fast bounded work, repetitive repository review, and an 82.8 percent false-or-materially-wrong rate after source adjudication.
A real-workflow qualification of oMLX Qwen 3.5 4B: fast bounded work, repetitive repository review, and an 82.8 percent false-or-materially-wrong rate after source adjudication.
A real-workflow qualification of oMLX Qwen 3.8 27B 8-bit: strong bounded work, useful code-review leads, and a costly false-positive problem.
I expected this test to end with a speed chart. It ended with a subtraction problem. The subject was Qwen 3.8 27B on a 64 GB M5 Max, served two ways: Ollama 0.34.4 and oMLX 0.7.0rc1. I wanted two answers that are easy to blur together. First, which serving engine handles the closest practical 4-bit…
Oddbyte’s five-agent setup already spans cloud and local models. The incoming M5 Max Studio is a chance to separate identity, orchestration and inference instead of moving the same bottleneck to a faster box.
Mac Studio, M5 generation (image: Apple) TL;DR: For running a single Qwen3.8-27B agent at long context, the Mac Studio is the better buy at every price tier — the $2,499 M5 Max is already faster on single-stream decode than the $3,999 DGX Spark, and the M5 Ultra is roughly 4–5× faster. The DGX Spark’s real…