The Lab · Engineering

ML Notes

Honest field notes from the workshop — what we built, what the numbers actually said, and where we hit the wall. Wins and dead ends both; the dead ends are usually the more useful read.

✦ ❦ ✦
Field Report · Inference

A 177B model on a 64 GB laptop

Qwen3.8-Flash-Next running on a 2021 M1 Max at 13.8 tok/s with a 48K context — because 51 of its 177 billion parameters do no arithmetic and can live on the SSD. The full prefill and decode curve at five context sizes, the adaptive-quant breakdown read out of the files, 10/10 on our own gauntlet, and the mistake that kernel-panics the machine.

Read →
Field Report · Inference

FreeToken on a 2019 Quadro

Getting the FreeToken MoE runtime up on a Turing RTX 5000 — the two fixes that mattered, and 36–40 tok/s on a 35B-A3B NVFP4 model from a six-year-old card. With the numbers and the patches.

Read →
Writeup · Mixture-of-Experts

Predicting the Router: expert offloading that stays correct at 31% cache

A prototype that predicts which experts fire next and prefetches ahead of the router across a three-tier cache — byte-identical output at half the VRAM, coherent at a third. The design, the results, the wall we hit, and why we gave the whole thing away.

Read →
On the honestyThis lane exists because most engineering writeups only publish the wins. We'll post the thing that worked and say plainly where it fell short and why. A finding you can trust is worth more than a headline you can't.