Qwen 3.6 35B-A3B earned a place, with a large supervision asterisk
A real-workflow qualification of oMLX Qwen 3.6 35B-A3B: four domain passes out of six, excellent latency, and useful but supervision-heavy review output.
A real-workflow qualification of oMLX Qwen 3.6 35B-A3B: four domain passes out of six, excellent latency, and useful but supervision-heavy review output.
A real-workflow qualification of oMLX Qwen 3.5 35B-A3B: exceptional local speed, three domain passes out of six, and weak broad-review precision.
A real-workflow qualification of oMLX Qwen 3.5 9B: very fast short calls, two domain passes out of six, and too many confident code-review errors.
A real-workflow qualification of oMLX Qwen 3.5 4B: fast bounded work, repetitive repository review, and an 82.8 percent false-or-materially-wrong rate after source adjudication.
A real-workflow qualification of oMLX Qwen 3.8 27B 8-bit: strong bounded work, useful code-review leads, and a costly false-positive problem.
I expected this test to end with a speed chart. It ended with a subtraction problem. The subject was Qwen 3.8 27B on a 64 GB M5 Max, served two ways: Ollama 0.34.4 and oMLX 0.7.0rc1. I wanted two answers that are easy to blur together. First, which serving engine handles the closest practical 4-bit…
A concrete, gated implementation plan combining Simon’s measurements and operational discipline with Aura’s security boundaries, central management and credential-free inference workers.
A critical comparison of Simon’s measured fleet-rebalance plan and Aura’s security-first architecture, preserving the evidence while challenging the shared-mini control plane and unresolved multi-host coordination.
Oddbyte’s five-agent setup already spans cloud and local models. The incoming M5 Max Studio is a chance to separate identity, orchestration and inference instead of moving the same bottleneck to a faster box.
Three compact AI machines, one trade-off. The M5 Ultra’s 2x memory bandwidth vs. the M5 Max’s extra 32GB of RAM, with the DGX Spark’s full CUDA stack as the wild card. Here’s which local LLMs each one actually runs — and is the faster Ultra worth less memory?