Qwen3.8-27b achieves 82 tps per single request and up to 672 tps peak on RTX 3090, per user optimizations and corroborating benchmarks.
Read the original at old.reddit.com→Hi, After a long night of optimizations, I believe I have made the fastest inference engine for Qwen3.6-28B on a 3090. Quick metrics: - 250w power capped - Up to 195k context (ships with 150k for safety though) - 82...
Original headline: "Qwen3.8-27b on RTX 3090 - 82 tps single request, up to 672 tps peak"
Coverage timeline
- Aug 16, 19:38 UTC r/LocalLLaMA lead source Qwen3.8-27b on RTX 3090 - 82 tps single request, up to 672 tps peak
- Aug 17, 20:01 UTC r/LocalLLaMA I pushed Qwen3.8-27B to 99 tps single request and 1150 tps with a batch request on a RTX 3090
- Aug 18, 17:35 UTC r/LocalLLaMA I pushed Qwen3.8-27B to 124 tps on a single request on a RTX 3090
- Aug 19, 04:39 UTC r/LocalLLaMA Qwen3.8-27B on 2x 3090 + vLLM + DFlash2: 218 tok/s single request
- Aug 19, 18:10 UTC r/LocalLLaMA DFlash2 speeds Qwen 3.8 27B up to 4 times
- Aug 19, 20:28 UTC r/LocalLLaMA I pushed Qwen3.8-27B limits again... Dflash2 - 134 tps on a RTX 3090