METALBENCH

UNRAID LAB / 128 GB DDR4 / RTX 3060

The fast model isn’t the reliable one.

Measured local inference, deterministic grading, and the failures that matter. No vendor numbers. No LLM judges.

Current verdictQwen for speed. Flash-Next for strict work.

RUN INDEX / 003

Latest evidence

Initial results are note-derived and clearly marked. Missing raw artifacts limit reproduction, not disclosure.

Documented local result

Flash-Next IQ3_XXS

llama.cpp · IQ3_XXS

41/42quality
16.92tok/s decode
Documented local result

Qwen3.6 NVFP4

FreeToken · NVFP4

7/9quality
61.5tok/s decode
Documented local result

Flash-Next Q3_K_XL

llama.cpp · Q3_K_XL

9/9quality
14.5tok/s decode

FINDING 01

Speculation doesn’t beat a memory bottleneck.

MTP accepted roughly 0.8 of drafted tokens but produced no wall-clock improvement on CPU-offloaded Flash-Next.

Read the finding →
0%measured speedup