Full 1M context on a single RTX5090 with DDR5 desktop setup using vLLM CPU/RAM offloading; ~800 tps per prompt and 15+ tps decode
Read the original at old.reddit.com→First of all, obviously I took some help from AI to type this post and this is the topic that enabled me to accomplish all that:...
Original headline: "[Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding]"
Coverage timeline
- Aug 4, 14:06 UTC r/LocalLLaMA lead source [Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding]
- Aug 4, 15:00 UTC r/LocalLLaMA Deepseek V4 Flash 2-bit quant is the first model I can run locally that achieves 100% in this SQL benchmark
- Aug 5, 00:10 UTC r/LocalLLaMA DeepSeek-V4-Flash on SM89 4x48gb 4090s with DSpark
- Aug 5, 10:55 UTC r/LocalLLaMA I updated my localy run benchmark with DeepSeek V4 Flash 0731
- Aug 5, 17:34 UTC r/LocalLLaMA DeepSeek V4 Flash 0731 at 10–17 t/s (nothink) on MacBook M5 Pro **64GB***, partly via SSD streaming
- Aug 6, 09:48 UTC r/LocalLLaMA Best llama cpp flags to run Deepseek-flash 0731
- Aug 6, 11:36 UTC r/LocalLLaMA Final optimization: from ~10 tok/s to ~15 tok/s on DeepSeek-V4-Flash-0731 at 128K ctx - 1 RTX 3090
- Aug 7, 11:19 UTC r/LocalLLaMA Anyone running DeepSeek-V4-Flash-0731 on MI325X with vLLM? Mine is behaving completely broken