Documented local result
Flash-Next IQ3_XXS
llama.cpp · IQ3_XXS
41/42quality
16.92tok/s decode
UNRAID LAB / 128 GB DDR4 / RTX 3060
Measured local inference, deterministic grading, and the failures that matter. No vendor numbers. No LLM judges.
RUN INDEX / 003
Initial results are note-derived and clearly marked. Missing raw artifacts limit reproduction, not disclosure.
llama.cpp · IQ3_XXS
FreeToken · NVFP4
llama.cpp · Q3_K_XL
FINDING 01
MTP accepted roughly 0.8 of drafted tokens but produced no wall-clock improvement on CPU-offloaded Flash-Next.
Read the finding →