Qwen3.8-27b outperforms larger models on benchmarks, but results are based on bf16 weights; practical runs use 4-bit quantization due to hardware limits, meaning the downloaded artifact differs from the measured one.
Read the original at old.reddit.com→qwen3.8 27B has seriously impressive benchmarks on its model card, but that's for the unquantised version. Almost everyone here will run one quant or another. Are there good benchmarks for how those quants perform? ...
Original headline: "we benchmark models nobody actually runs"
Coverage timeline
- Aug 17, 21:53 UTC r/LocalLLaMA lead source we benchmark models nobody actually runs