Qwen 3.8 27B runs slowly on vLLM with RTX 6000 Pro, regardless of thinking effort settings.
Read the original at old.reddit.com→Title pretty much says it all. I’ve deployed Qwen 3.8 27B using vLLM on an RTX 6000 Pro (tried multiple vLLM releases and launch recipes), but I can't get it into a usable state because of crazy long reasoning...
Original headline: "Anyone managed to get Qwen 3.8 27B running smoothly on vLLM? Can't get rid of endless thinking"
Coverage timeline
- Aug 16, 05:47 UTC r/LocalLLaMA lead source Anyone managed to get Qwen 3.8 27B running smoothly on vLLM? Can't get rid of endless thinking
- Aug 16, 19:14 UTC r/LocalLLaMA Qwen 3.8-27b unusable long thinking?
- Aug 17, 08:32 UTC r/LocalLLaMA How to stop Qwen3.8-27b from overthinking
- Aug 17, 15:36 UTC r/LocalLLaMA Qwen 3.8 27B Overthinking, It has to be done, it has to be overthinking to punch Opus 4.6