Qwen3.8-2.4T-A95B runs locally on RTX 5090 + RTX 5060 Ti at about 0.80 tokens per second using llama.cpp with Unsloth GGUF quantization and 512 routed experts, 10 active per token.
Read the original at old.reddit.com→I managed to get Qwen3.8-2.4T-A95B running locally with llama.cpp on mu PC just for fun, cause why not. I was using the Unsloth Qwen3.8-2.4T-A95B-UD-Q1_0 GGUF quantization. The full GGUF is about 397 GiB. The model...
Original headline: "EXPERIMENT: Qwen3.8-2.4T-A95B running locally on an RTX 5090 + RTX 5060 Ti at ~0.80 tok/s"
Coverage timeline
- Aug 13, 18:52 UTC r/LocalLLaMA lead source EXPERIMENT: Qwen3.8-2.4T-A95B running locally on an RTX 5090 + RTX 5060 Ti at ~0.80 tok/s
- Aug 15, 14:19 UTC r/LocalLLaMA If you would have told me half a year ago that a local model running in my office would be able to one-shot a Super Mario clone, I would have called you nuts. Qwen3.8-27B is a different beast.
- Aug 16, 04:28 UTC r/LocalLLaMA 5090: Windows or Linux for Qwen3.8.27b
- Aug 16, 17:57 UTC r/LocalLLaMA Qwen 3.8 2.4T at 288k tokens/s on Nvidia GB300 NVL72