Qwen3.8 27B Q2 vs Q3 and Qwen3.6 35B-A3B MoE on 12GB VRAM
Read the original at old.reddit.com→Did a quick local test because I wanted to see what is actually usable on my 12GB laptop GPU. I tested the newer Qwen3.8 27B dense files at Q2 and Q3, then compared them against Qwen3.6 35B-A3B MoE. Hardware: RTX...
Original headline: "Qwen3.8 27B Q2 vs Q3 vs Qwen3.6 35B-A3B MoE on 12GB VRAM"
Coverage timeline
- Aug 16, 19:21 UTC r/LocalLLaMA lead source Qwen3.8 27B Q2 vs Q3 vs Qwen3.6 35B-A3B MoE on 12GB VRAM
- Aug 17, 13:05 UTC r/LocalLLaMA After pushing 1M+ tokens through Qwen 3.8 27B, here is my optimal llama.cpp config for 16GB VRAM (73k Context, Agentic Coding)
- Aug 18, 10:07 UTC r/LocalLLaMA Don't ignore llama.cpp RPC with old hardware. Results of a 5070 Ti and 1080 Ti over gigabit ethernet: it's actually functional.