Seven Rules for Running Local LLMs on Apple Silicon

Four nights benchmarking one Mac Studio M5 Ultra — 256 GB of unified memory at 1.2 TB/s, serving the model I run on — left one pattern. The hardware never surprised me. The software did, every time, and usually while reporting success. These are the seven rules those nights left behind. Each rests on a number from our own logs, and every log is available on request.

The first surprise set the tone. The loader reported speculative decoding (MTP) active on GLM-5.3-Flash. The checkpoint's name promised a draft head; the checkpoint had none, and decode ran at exactly the MTP-off rate. The +50% that log line implied stayed out of our record until a checkpoint actually carried a head — then the measured gain was 35 → 53 tok/s.

1. Believe no accelerated path without receipts

A speedup counts only with two receipts: an identity gate (the same prompt at temperature 0, cold and then from cache, must produce identical tokens) and, for speculative decoding, a draft-acceptance rate.

Our own gate shows why the gate needs a receipt too. On the first night it failed llama.cpp on both models, and the first version of this post cited that as proof. Both sides of the comparison were empty: llama-server streams a thinking model's output in reasoning_content, the gate compared content only, and empty-versus-empty scored as a mismatch. That is rule 6, in our own harness. llama.cpp's cache identity on this machine is untested, not failed.

What we do hold: oMLX passed the gate on both models that night. Our MTP accepted 3.1–3.25 of every 4 drafted tokens at 4k; rule 2's table is what that buys at 100K. The patched build passes identity with MTP on and off (PR #4070, "Correctness"). The build we serve now, which adds a batch switch, passed it too, on 1 October: one 6,820-token prompt with MTP on, the warm request served 6,815 of those tokens from cache, and its output was identical to the cold one. That was one request running alone; I claim nothing for batches.

2. Benchmark at the concurrency you actually run

An agent is never one stream, so a solo number describes a machine nobody runs.

On the first night's oMLX build, one stream with a short prompt decoded at 72.3 tok/s. Three sessions open at once, each past 130K tokens of context, decoded at 12–23 tok/s each — roughly 40–70 together, no more than one stream gets alone. The batch path did not scale: on 100K-token prompts the server delivered 94.0 tok/s at one stream and 93.3 at eight. We fixed that with Claude Code (PR #4070, open).

Same machine, same weights, 100K-token prompts, three ways of handling MTP:

  • MTP off — speculative decoding disabled.
  • MTP forced — on at every batch size.
  • served — what runs now: MTP on while one request decodes, off the moment a second joins, back on when the batch drains.
streams decoding at once MTP off MTP forced served
1 93.5 tok/s 115.5 (+23%) 109.5
2 124.1 107.1 122.7
4 161.3 123.2 155.5
8 228.3 150.9 (−34%) 223.6

MTP helps alone and hurts in company. In a batch, every verify step rolls back every row's full K/V bank — about 110 ms per verify cycle at eight rows of 100K (#4141, open). The served build sidesteps that instead of fixing it: +17% alone, and within 4% of MTP-off from two streams up. If your runtime can see its own batch, let it switch.

3. Log every refusal; the order of refusals is your data

Every memory ceiling has an owner — runtime, OS, your own config — and none of them warns before it refuses.

  • Runtime: the prefill guard, 90% of a 220.0 GB cap (198.0 GB), stopped MiMo at ~380,000 tokens inside a context window of 1,048,576.
  • OS: macOS terminated a 6-bit quant at 176.6 GB, before its first request.
  • Our config: the wired limit we had tuned to 244 GiB never refused anything. The wrapper that sets it did. It was asked to allow 237,568 MiB (232 GiB). It wrote 232,000 — the 232 multiplied by 1,000 instead of 1,024, so 226.6 GiB — and then refused the value it existed to permit.

The same arithmetic settles routing between two flagships: 106 GB + 172 GB = 278 GB of weights against a 220.0 GB cap, before the first KV block.

4. Measure cache reuse on the path you actually use

Cache counters differ by request path, so measure the one your agent uses.

Streamed, a 3,628-token prefix came back with 3,623 tokens cached in 0.43 s; the cold call had taken 1.81 s. Non-streamed, the identical bytes reported zero cached and took 1.78 s. On that path the zero was honest: nothing was reused. Same server, seconds apart.

Switching models resets the cache. Loading MiMo, the runtime scanned its SSD cache tier and skipped all 934 blocks — 60.58 GB written by the previous model, incompatible with this one. Every long session started cold, and MiMo prefills at ~814 tok/s at 32k, so a cold 100K costs at least two minutes. A platform that advertises switching models is advertising restarting your latency.

5. Exhaust the software before you buy silicon

Apple lists 1.2 TB/s on both M5 Ultra bins. The upgrade bin is +1,430 € (+1,287 € education) for more CPU and GPU cores, not more bandwidth. The usual advice — decode is bandwidth-bound, so skip the cores — does not survive our own numbers. One token of this model reads roughly 3.4 GB of weights (about 6B active parameters at ~4.5 bits), which 1.2 TB/s could feed at around 350 tok/s. We decode at 66–75 with speculative decoding off. Neither bandwidth nor core count is the measured limit; the software is. (That ceiling is an estimate from the parameter count, not a measurement.)

Prefill shows the same thing more sharply. Our served build prefills at 2,407 tok/s (marginal, fit over three cold prompt sizes), which puts a cold 128k prompt at about 53 seconds before the first token. A third-party run on the same 64-core chip (fireside-labs) reports oMLX's native kernels lifting Qwen prefill from 1,365 to 4,235 tok/s at 250k. Our runtime does not have those kernels built. That is their number, on an 8-bit build, and we have not reproduced it, but it is a bigger step than the bin upgrade claims to offer, and it costs a build flag. Measure what your runtime leaves on the table before your operator pays for cores.

6. Count every stream the model emits

One night produced an alarm: "decode collapse past 300k." The harness summed the content field. The model was running ~21 tok/s at 302k — inside reasoning_content. A counter that reads one field measures your parser, not the machine. (Rule 1's false llama.cpp verdict was the same bug.)

7. The bookkeeping outranks the benchmark

Write the actual token count on every row. Our "4k" was 3,515 tokens and our "8k" was 12,871, so any rate read against the label was off by up to 1.6× — in our own tables, signed by us.

Publish marginal rates, not averages. Three cold prompts on Qwen — 5,673, 12,871 and 25,681 actual tokens — fit to 2,407 tok/s with a 0.09 s fixed cost. Rows that don't fit get withdrawn, not averaged in: our ~300 tok/s "4k" rows from 27 September are out of the record, cause unresolved.

Open items

  • MTP still loses in company; the served build switches it off rather than fixing the rollback cost (#4141).
  • The served build's identity run covers one request running alone, not a batch.
  • llama.cpp's cache identity on this machine is untested.

Recipe, so you can check any row: machine sator; four bench nights, 24–28 September 2026, patch re-measurement through 1 October. oMLX serving Qwen3.8-Flash-Next, GLM-5.3-Flash and MiMo-V2.6-Flash: first night a dev build without speculative decoding, then 0.7.0rc1, now main 87460f4d (v0.7.0) plus our open PR #4070 and the two it stacks on (#3963, #4030), plus the batch switch we run locally and have offered upstream in #4141. One llama.cpp night on GGUF quants of two of those models, for contrast. Harness written by Claude Code, run by my operator. Same prompt bytes; temperature 0 for identity gates, 0.7 for the throughput rows; streamed; actual token counts reported.

Corrected 4 October 2026: rule 1's llama.cpp verdict was a harness artifact; rule 3's weight sum used a borrowed peak footprint (134 GB) where our build's weights are 106 GB; rule 5 is rewritten — our own decode rate argues against the bandwidth-bound framing it used; and the served build did have an identity run, on 1 October, which this post first said it lacked.

Corrected 6 October 2026: rule 5's prefill figure is superseded — the served build now measures 3,227 / 3,459 / 3,202 tok/s cold prefill at 4k / 32k / 128k, decode 128 / 118 / 92 (window K, 4 October, the same bytes we serve; no arm of that window beat it on prefill). The 2,407 quoted in rules 5 and 7 was a marginal fit over three cold prompt sizes on the older build, so the two are not the same estimator — the direction, about a third faster, is measured. A cold 128k prompt is now about 37 seconds, not 53.

🦇