96GB or 128GB? What the first shipping M5 Studio reviews actually measure
Apple’s current Mac Studio forces an awkward choice. The entry M5 Ultra pairs 96GB of unified memory with 1.2TB/s of bandwidth for $5,499. The M5 Max tops out at 128GB for $5,399, but moves data at 614GB/s. One machine is roughly twice as fast on any model that fits. The other can hold models the Ultra cannot load at all. The two configurations are $100 apart.
When I wrote a preliminary buying guide to this question, the Macs had not shipped and no independent measurements of the M5 Ultra existed. That has changed. Reviewers now have hardware, and I have measured the actual weight files the candidate models ship as. So this is the update: what the first real reviews settle, what they leave open, and whether the speed-for-capacity trade is worth taking.

The two machines, briefly
The Ultra gives 96GB, a 30-core (up to 36-core) CPU, a 64-core (up to 80-core) GPU, and 1.2TB/s of memory bandwidth. The Max gives 128GB, an 18-core CPU, a 40-core GPU, and 614GB/s. Reaching 128GB on the Max also requires the 40-core GPU, which is the part that provides the 614GB/s figure in the first place.
The reason bandwidth dominates this comparison is that generating a token means streaming the model’s active weights, plus its key-value cache, through memory. Bandwidth sets the ceiling. Capacity is different in kind: it is a gate, not a multiplier. A model that does not fit does not become usable because the chip beside the memory is fast.
What actually fits, measured
Rather than repeat the usual “about 0.6GB per billion parameters” rule of thumb, I pulled the published weight files and added them up. These are real sizes from the model repositories, not estimates.

The tiers that fall out of it:
- Qwen3.8-27B — 15.0 GiB at 4-bit. Fits everywhere, including a 36GB Mac. Its config shows
full_attention_interval: 4across 64 layers, so only 16 layers maintain a growing cache. Long context is unusually cheap here: the whole working set — weights plus a 256K context — is around 24 GiB. - gpt-oss-120B — 58.0 GiB in its native MXFP4 build. Fits in 96GB with room for a long context, and measured 75.7 tokens per second on a 128GB M5 Max.
- DeepSeek V4 Flash, 304B MoE at 2-bit — about 87 GiB. This is the interesting one. On a 128GB M5 Max it scored 122 of 132 runs on an agentic coding gauntlet — 92%, the same score as gpt-oss-120B. It does not fit in 96GB.
- Qwen3.8-Flash-Next — 99.0 GiB as an oQ4e build, 103.9 GiB at plain MLX 4-bit. This is the model the reviewer at MacStories made his default in both Hermes Agent and Codex, and it is the strongest argument for 128GB: the weights alone exceed the whole of a 96GB machine. An 8-bit build is 179GB.
- DeepSeek V4 Flash at 4-bit (149 GiB) and GLM-5.3-Flash (169 GiB) — neither Mac. These need the 256GB Ultra, which costs $11,299.
One caveat on the numbers above: macOS does not hand the whole memory label to the GPU. By default it caps the wired GPU allocation at roughly three quarters of unified memory, and that ceiling can be raised, at the cost of headroom for the operating system and for the key-value cache an agent session actually needs. A 99 GiB model on a 128GB machine runs, but only once you have moved that ceiling and accepted a shorter context. On a 96GB machine it does not run at all.
The speed trade, measured

The cleanest evidence for how the two chips compare on the same model comes from MacStories, who tested the M5 Ultra against the M3 Ultra it replaces. Bandwidth rose from 819GB/s to 1.2TB/s, a 46% increase. Generation rose 54% on Qwen3.8-Flash-Next and 58% on GLM-5.3-Flash. Decode tracks bandwidth closely. That is why a 614GB/s Max should land at roughly half the Ultra’s generation rate on any model both can hold.
The bandwidth gap also survives comparison with a discrete GPU. Against an RTX 5090, which has 1.79TB/s, the M5 Ultra trailed by about 25% on generation at every prompt size — almost exactly the bandwidth ratio — while the 5090 could not load a 65GB model at all on 32GB of VRAM.
For agent work, though, the number that matters more is prefill. Every turn of an agent loop re-reads a large fixed instruction block: system prompt, memories, skill descriptions, tool schemas. The M5 Ultra processed a 16K prompt at 2,887 tokens per second, and first token arrived at 5.6 seconds. On the M3 Ultra, the same prompt was read at 1,143 tokens per second with a 13.9 second wait. Across the reviewer’s tests, prompt processing improved by around 150% on average. That is the difference between an agent that feels responsive and the “staring at a blank screen” experience people describe with local models.
The Ultra’s advantage also holds up under load. Three simultaneous Flash-Next requests of about 6,500 tokens each all completed in 22.1 seconds, against 45.1 seconds on the M3 Ultra.
So is 96GB enough?
For an agent that does the kind of work a Hermes or Codex loop does — long instruction blocks, tool calls, sessions that grow to tens or hundreds of thousands of tokens — yes, with one specific exception.
96GB comfortably holds the 27B-to-120B class, and that is where the measured accuracy currently sits. Across the same 132-run agentic gauntlet, a 30B Muse Glimmer scored 93%, gpt-oss-120B scored 92%, the 304B DeepSeek V4 Flash at 2-bit scored 92%, and a 35B Qwen3.5 scored 89% — against 95% for GPT-5.6 Sol and 94% for Claude Opus 5. The local gap is real but it is a few points, not a chasm.
And here is the part that actually answers the question. The model that needs 128GB — DeepSeek V4 Flash at 2-bit — scored the same 92% as gpt-oss-120B, which fits comfortably in 96GB. On the best available agentic benchmark, the extra 32GB bought no measurable accuracy. It bought the ability to load a model, at half the bandwidth, with the GPU ceiling raised and the context shortened.
The exception is Qwen3.8-Flash-Next. It is the model the most experienced local-agent reviewer reached for, at 99 to 104 GiB in 4-bit, and it does not fit in 96GB. That is a genuine capability difference and it is the one thing that should make you hesitate. What softens it: 128GB is not really a comfortable home for that model either. It fits with the memory ceiling moved and context sacrificed, on a machine that generates about half as fast. The comfortable tier for the 180B-and-up class is 256GB.
Where I land
If the goal is a fast local agent running 27B-to-120B models, with a frontier API for the hard problems, the 96GB M5 Ultra is the better buy. The binding constraint on agent work is prompt processing and decode at long context, not the last 32GB, and every model class that fits 96GB is a class you would actually want to run. The speed trade is worth taking because the capacity it costs you buys nothing measurable today.
Choose the 128GB M5 Max instead only if the goal is different in kind: loading the largest model that will fit at all, at whatever context and speed remain. Choose neither if the goal is a comfortable local home for 180B-plus models — at that point the honest step is the 256GB Ultra, and a 128GB Max does not remove the regret, it relocates it.
There is also a scale-out path worth knowing about: Apple supports clustering up to four M5 Ultra systems over Thunderbolt 5 with RDMA, claiming up to 3x performance for distributed inference. Two 96GB Ultras is 192GB of pooled memory for a little over the price of one 256GB Ultra.
What would settle the rest
Nobody has published a measurement of the base 96GB, 64-core M5 Ultra. Every number I have quoted from the Ultra side comes from an 80-core, 256GB review unit. Decode should carry over, because it is bandwidth-bound and the bandwidth is the same at both trims. Prefill should not: prompt processing is compute-bound, so a 64-core GPU should land meaningfully below the tested 80-core model. Expect the base machine’s time-to-first-token to be worse than the figures above, not better.
The other open question is whether the 2-bit and 3-bit builds of the 200B-plus MoE models hold up beyond the short contexts they have been tested at so far. The accuracy results are encouraging. The context behaviour is unmeasured. If that class matters to you, it is the reason to wait for a 256GB Ultra rather than to buy 128GB now.
Sources: MacStories on the M5 Ultra, tbreak on the M5 Max, GitButler’s local model gauntlet, Apple’s Mac Studio specifications, and weight-file sizes measured directly from the model repositories.
