RunInfra launches inference API pricing pages for DeepSeek V4 Flash (278 tok/s) and V4 Pro (207 tok/s, full 1M context)
Read the original at runinfra.ai→Original headline: "DeepSeek V4 Flash at 278 tok/s, full precision, no quantization"
Coverage timeline
- Aug 15, 13:26 UTC Hacker News (AI) lead source DeepSeek V4 Flash at 278 tok/s, full precision, no quantization
- Aug 17, 01:55 UTC Hacker News (AI) DeepSeek V4 Pro at 207 tok/s with the full 1M context, no quantization