Humaneval benchmark for Deepseek V4 Flash 0731 vs GLM5.3 Flash on 2x DGX Spark setup
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
I have a 2 DGX Spark setup recently and I have been happily running Deepseek V4 Flash 0731. Since the release of GLM5.3 Flash and Qwen 3.8 Flash Next this week, a lot of folks are still waiting to see what model to run given their own hardware situations. I am very interested in running a GLM model locally and looks like nvfp4 would be a good option for my setup, but I have been hearing a lot of conflicting opinions (mostly negative) about GLM5.3 Flash on nvfp4 quant which gave me pause. So I decided to do a simple benchmark myself and hopefully this is useful for folks with the same setup:
Deepseek V4 Flash 0731 recipe : official checkpoint, fp8 kv, 1M context, 4 concurrent streams
GLM 5.3 Flash recipe : NVFP4 quant & Dflash2 drafter, fp8_e4m3 kv, 256k context, 6C
| Model | Thinking Mode | HumanEval Pass@1 (Base) | HumanEval+ (Adversarial Edge Cases) | Total Benchmark Run Time | Local Stream Speed |
|---|---|---|---|---|---|
| GLM-5.3-Flash NVFP4 | Thinking Enabled (high) | 97.0% (159 / 164) | 92.1% (151 / 164) | 20m 52s | ~50 tok/s (DFlash2) |
| DeepSeek-V4-Flash-0731 | Thinking Enabled (high) | 94.5% (155 / 164) | 88.4% (145 / 164) | 38m 16s | ~70 tok/s (MTP-5) |
| GLM-5.3-Flash NVFP4 | Direct Zero-Shot (off) | 93.3% (153 / 164) | 89.6% (147 / 164) | 23m 32s | ~50 tok/s (DFlash2) |
| DeepSeek-V4-Flash-0731 | Direct Zero-Shot (off) | 92.7% (152 / 164) | 87.8% (144 / 164) | 14m 52s | ~70 tok/s (MTP-5) |
So raw numbers tell you GLM 5.3f is a decent upgrade over DSv4f 0731 especially with thinking enabled. Unsloth saying their Q4 quant has around 92% accuracy but looks like nvfp4 still holds up pretty well (97% would have been a SOTA score not that long ago and this is not even max thinking). The major trade off is the 256 context. I am pretty sure 512GB+ VRAM (or 4x sparks) people will be able to run the fp8 model + 1M context without issues and I am jealous 🥹 Regardless, your own experience matters more than any benchmark out there.
[link] [comments]
More from r/LocalLLaMA
-
an unscientific qwen 3.8 flash next and glm 5.3 flash comparison
Aug 30
-
Ran Qwen3.8-Flash-Next (79 GB, 2-bit) at 350K ctx for 3.5 hours on a 128 GB M5 Max — speed vs context depth, 100 turns, one graph
Aug 30
-
Nemotron-3.5-Lightning at 11.77 GiB, a 16 GB option for a model that didn't have one
Aug 29
-
Any current Voice2Voice AI model that runs locally that’s good?
Aug 29
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.