METALBENCH

FINDING / 2026-08-29

Speculation doesn’t beat a memory bottleneck

MTP accepted roughly 0.8 of drafted tokens but produced no wall-clock improvement on CPU-offloaded Flash-Next.

The documented baseline measured 14.67 tok/s. The MTP run measured approximately 13.6–14 tok/s despite acceptance near 0.8.

The practical conclusion is narrow: on this memory-bandwidth-bound configuration, accepted speculative tokens did not translate into faster single-stream decode.

Evidence

RunScoreDecodeRuntimeConfigurationEvidence
Flash-Next Q3_K_XLflash-next-q3kxl-agent-2026-08-289/914.5 tok/sllama.cppQ3_K_XL · 32,768 ctx · c1Documented local result

Limitations

  • Transcribed from local notes
  • No raw timing artifact is currently public