Honest field notes from the workshop — what we built, what the numbers actually said, and where we hit the wall. Wins and dead ends both; the dead ends are usually the more useful read.
Qwen3.8-Flash-Next running on a 2021 M1 Max at 13.8 tok/s with a 48K context — because 51 of its 177 billion parameters do no arithmetic and can live on the SSD. The full prefill and decode curve at five context sizes, the adaptive-quant breakdown read out of the files, 10/10 on our own gauntlet, and the mistake that kernel-panics the machine.
Read → Field Report · InferenceGetting the FreeToken MoE runtime up on a Turing RTX 5000 — the two fixes that mattered, and 36–40 tok/s on a 35B-A3B NVFP4 model from a six-year-old card. With the numbers and the patches.
Read → Writeup · Mixture-of-ExpertsA prototype that predicts which experts fire next and prefetches ahead of the router across a three-tier cache — byte-identical output at half the VRAM, coherent at a third. The design, the results, the wall we hit, and why we gave the whole thing away.
Read →