Benched a 124B on one DGX Spark for a week and published results; 38.7 tok/s on the fastest path found, 2.4x DeepSeek V4 Flash on the same box
Read the original at old.reddit.com→sudoingX on X spent a week on this and posted his wrap-up. The hardware is his. He isn't affiliated with us and we didn't see any of it before he put it out — I work on Ling at inclusionAI. Where he landed on one...
Original headline: "Benched a 124B on one DGX Spark for a week and published all of it — 38.7 tok/s on the fastest path he found, 2.4x DeepSeek V4 Flash on the same box"
Coverage timeline
- Aug 12, 16:34 UTC r/LocalLLaMA lead source Benched a 124B on one DGX Spark for a week and published all of it — 38.7 tok/s on the fastest path he found, 2.4x DeepSeek V4 Flash on the same box
- Aug 12, 22:41 UTC r/LocalLLaMA I ran DeepSeek V4 Flash 284B + DSpark on one RTX PRO 6000. The drafter was faster in RAM than VRAM.
- Aug 13, 16:12 UTC r/LocalLLaMA A 124B emitted 15,128 tokens in a single response on one DGX Spark, decode went 35.62 → 35.68 tok/s across the whole thing
- Aug 15, 10:34 UTC r/LocalLLaMA Deepseek v4 flash Q2 on a single 4090 😅