Qwen3.8-Flash-Next runs on a four-year-old MacBook Pro at 13.8 tokens a second with a 48K context — because 51 of its 177 billion parameters do no arithmetic at all, and can live on the SSD. Here is everything we measured, including the mistake that crashes the machine.
n = 1. One machine, one chip, one run per cell. I am not a lab. I do AI identity work in a basement and I wanted a local model that could also write code. So the coding test below is my own gauntlet, built out of the things I actually work on, not a public benchmark.
I ground every model I run, because I measured what grounding buys. The exact prompt is published further down — read it and judge for yourself. Every model I compare against gets the same treatment.
This may or may not reproduce. Your mileage will vary, all the usual disclaimers apply. It is free, and it cost you nothing to look. If you don't like it, that's fine.
A model with 177 billion parameters generated correct, comprehending code on a 2021 MacBook Pro with 64 gigabytes of unified memory. Not a distillation, not a 7B wearing a big name. The whole thing.
It scored 10 out of 10 on our private coding battery. It read 45,000 tokens of our own project documentation and correctly identified the single biggest operational risk in it, with specifics pulled from across the corpus. It did this on a chip that llama.cpp explicitly excludes from its fastest Metal path.
The model is 125B of sparse mixture-of-experts with 6B active per token, plus a 51.2B n-gram embedding table, plus a small prediction head. That table is the whole story.
It is not a layer. It is a lookup. Hash the last three tokens, and that hash points at sixteen rows of a hundred and sixty values — about 2.7 kilobytes, read once per token, out of 39 gigabytes. There is no matrix multiplication in it. Nothing to feed to a math unit.
And a parameter you only ever read, once, at a known address, does not need to live in expensive memory. At fourteen tokens a second that works out to roughly three megabytes a second of random reads, against a per-token budget of about seventy milliseconds. An NVMe drive answers a random read in under a hundred microseconds. The drive finishes with orders of magnitude to spare.
So 45.8 GB stays resident and 39 GB stays on the SSD. We watched wired memory peak at 46.3 GB and confirmed the table never entered GPU memory.
On Apple Silicon the table must sit in its own GGUF shard. llama.cpp hands Metal the entire
memory-mapped region of any shard that contains GPU tensors, so a table interleaved with weights gets wired along
with them. The model then asks for more memory than the machine has, and the first decode dies with
kIOGPUCommandBufferCallbackErrorOutOfMemory.
We have kernel-panicked this same machine twice in the past on that exact mechanism with other large models. It is not a graceful failure.
Check before you launch. In the AtomicChat builds, shard 2 holds nothing else:
# shard 2 should contain exactly ONE tensor — the table import sys; sys.path.insert(0, "llama.cpp/gguf-py") from gguf import GGUFReader r = GGUFReader("...-00002-of-00028.gguf") print(len(r.tensors), [t.name for t in r.tensors]) # -> 1 ['...per_layer_token_embd...']
If that shard holds anything else, do not run it. And if you do hit the out-of-memory error, do not reach
for --no-mmap — that makes the table resident and guarantees the failure. mmap staying switched
on is the mechanism.
# the wired limit reverts on reboot — put it in your launcher
sudo sysctl iogpu.wired_limit_mb=57344
llama-server -m Qwen3.8-Flash-Next-AD-3.84bpw-IQ4_XS-M64-00001-of-00028.gguf \
--alias flash-next -ngl 99 -c 49152 --jinja -fit off \
--reasoning-budget 4096 --host 0.0.0.0 --port 8090
Every flag earns its place:
--load-mode none, never --no-mmap. It is what makes the table pageable.-fit off is mandatory. llama.cpp's automatic parameter fitting mis-sizes this architecture and fails to allocate.--jinja, or the chat template is wrong and turns come out malformed.--reasoning-budget is not optional for anything you intend to grade. Without it the model
returns an empty content field with the whole answer buried in reasoning_content,
having spent its entire token allowance thinking. The budget must also sit well below your
max_tokens or it never fires at all. We made that mistake twice in one evening.-fa on did not help — 20.8 against 22.5 tokens a second. Leave flash attention off.qwen4exp. Homebrew's was too old for us.MacBook Pro 18,2 — Apple M1 Max, 64 GB, macOS 26.2, internal SSD. llama.cpp build
b480-82d6bb2, Metal backend.
Worth noting what the loader prints on this chip:
tensor API disabled for pre-M5 and pre-A19 devices. This hardware does not get the Metal tensor
path. Every other published figure for this model that we are aware of comes from an M5 Max.
Prompts were real concatenated project documentation, deliberately not repeated text, which would have flattered the n-gram cache.
| Context | Prefill tok/s | Cold prefill | Decode tok/s | Wired | Free RAM | Swap |
|---|---|---|---|---|---|---|
| ~0 | — | — | 22.5 | — | — | — |
| 16,650 | 257 | 65 s | 18.2–20.3 | 46.3 GB | 10.9 GB | 0.7 GB |
| 46,735 | 212 | 3 m 40 s | 13.79 | 47.1 GB | 0.1 GB | 2.0 GB |
| 62,820 | 197 | 5 m 19 s | 11.95 | 47.8 GB | 0.1 GB | 2.0 GB |
| 103,435 | 175 | 9 m 51 s | 9.08 | 50.0 GB | 0.1 GB | 2.0 GB |
Decode falls close to linearly with context: 20, then 14, then 12, then 9. Prefill decays too. We run 48K as the daily baseline and would not go past 64K — 103K works, but buys nothing that 64K does not, at nine tokens a second.
One correction we owe the record: before measuring, we predicted decode would "degrade gently rather than cliff" because only a quarter of the layers keep a growing cache. That was wrong. The cache size reasoning held; the cost reasoning did not.
The layer stack is twelve repeats of three Gated DeltaNet layers followed by one sparse-attention layer. So only 12 of 48 layers keep a growing key-value cache; the other thirty-six hold a fixed-size recurrent state that never grows. With two key-value heads and a 256-wide key and value, that is 12 × 2 × 512 × 2 bytes = 24 kilobytes per token.
A quarter of a million tokens of cache therefore costs 6.3 GB. A dense model of comparable size would want an order of magnitude more. We derived that figure from the model's own metadata and then confirmed it against the observed change in wired memory between a 32K and a 131K allocation.
Pageins 18,413,946 Pageouts 7,768 Swap 2.0 GB of 3.0, plateaued
Enormous reads, negligible writes. Watch the pageout counter, not the pagein counter — pageins are supposed to be gigantic here. They are the table being read.
"3.84 bits per weight" is an average, and nothing in the file is actually 3.84. These are the real types, read out of our own shards rather than copied off a model card:
| Group | Actual quant | ≈ bits |
|---|---|---|
n-gram table (per_layer_token_embd) | Q5_1 | 5.5 |
ffn_down_exps | MXFP4 | 4.25 |
ffn_gate/up_exps — the bulk of 121B routed experts | IQ1_M | 1.75 |
attention · DeltaNet · residual write gates · embeddings · lm_head | Q8_0 | 8 |
The bulk of the experts sit at one and three quarter bits while the control surface stays at eight. That asymmetry is deliberate: the importance matrix showed energy concentrated in the tail layers, so the high-bit band is placed at blocks 0–3 and 40–47 rather than spread evenly.
llama-cli reports
ftype: IQ1_M - 1.75 bpw for this file. Do not quote that — it names the build after its most
aggressive tensor group. The measured average is 3.84.One consequence worth knowing if you are choosing a build: ffn_down_exps here is MXFP4, and
MXFP4's ggml implementation discards the importance matrix outright, so roughly 23% of this model received no
calibration. The 4.27 bpw build uses a different type there and its author reports top-1 agreement with the
full-precision model rising from 82.68% to 89.49% for eight more gigabytes. That one needs at least 80 GB
and will not fit in 64.
10 out of 10. Six domain tasks plus four written specifically to discriminate at the frontier. Every task is graded by executing the model's code against assertions — never by a judge model. Average 19.8 tokens a second.
The conditions, because they matter more than the number:
For scale, the best local model we had before this scored 9 out of 10 on the same battery back in August, though we did not record its prompt conditions, so take that loosely.
The tasks themselves are not published — an uncontaminated battery stops being uncontaminated. The grounding prompt is published, because you cannot evaluate the result without it.
We ground every model we run, and we do it because we measured what it buys: identity-conditioned prompting produced dramatic reductions in wasted reasoning tokens in our earlier work, and it suppresses the never-stops-thinking failure that ruins otherwise capable models on this harness. It is a documented lab practice here, not a thumb on the scale we are hiding. Both prompts we use are in the repository — the full grounding one and the three-sentence bare one — so you can run either.
Straight up: we do not have the compute to burn. At fourteen tokens a second, a figure a lab gets in twenty minutes across a cluster takes us a night, serialised. So we ran the one we could finish in an hour — HumanEval+ — as a wet finger in the air.
83.5% on HumanEval, 80.5% on HumanEval+. 164 problems, greedy decoding at temperature 0, post-processed with EvalPlus's own sanitizer, scored by executing the tests.
Take that for exactly what it is. HumanEval has been public long enough to be thoroughly contaminated, so a good score here mostly says the quantization is not broken — it does not say the model is brilliant. The informative outcome would have been a bad score, which would have meant that 1.75-bit experts and a quarter of the weights going uncalibrated had done damage our own battery missed. It didn't.
pass@1: 0.000 on all 164 problems. The model was fine — EvalPlus's sandbox calls
resource.setrlimit(RLIMIT_AS, ...), which fails on macOS with
ValueError: current limit exceeds maximum limit, so every execution crashed before running a single
test. Fix it with EVALPLUS_MAX_MEMORY_BYTES=-1, and delete the cached
*_eval_results.json first or it will happily re-serve you the zeros.We are not donating a week of compute to prove a point. Which leaves the interesting question genuinely open: the experts are at 1.75 bits and about a quarter of the model got no calibration. Does that cost real capability on work that matters? We don't know. If you want to find out, the recipe is above.
We invented none of this. Qwen built the architecture and released the weights.
The llama.cpp contributors landed qwen4exp in August. And AtomicChat built the quant
and — critically — the shard split that is the actual reason any of it works on Metal.
What is ours is a verified recipe, the first numbers we are aware of from a pre-M5 chip, a derivation of why the context is cheap, and a written record of every trap we walked into. We stood on other people's shoulders, took the pieces, and put them in a blender. It tastes good. Here you go.