FINDING / 2026-08-29
Speculation doesn’t beat a memory bottleneck
MTP accepted roughly 0.8 of drafted tokens but produced no wall-clock improvement on CPU-offloaded Flash-Next.
The documented baseline measured 14.67 tok/s. The MTP run measured approximately 13.6–14 tok/s despite acceptance near 0.8.
The practical conclusion is narrow: on this memory-bandwidth-bound configuration, accepted speculative tokens did not translate into faster single-stream decode.
Evidence
| Run | Score | Decode | Runtime | Configuration | Evidence |
|---|---|---|---|---|---|
| Flash-Next Q3_K_XLflash-next-q3kxl-agent-2026-08-28 | 9/9 | 14.5 tok/s | llama.cpp | Q3_K_XL · 32,768 ctx · c1 | Documented local result |
Limitations
- Transcribed from local notes
- No raw timing artifact is currently public