Updated benchmark: Deepseek V4 Flash on SlopCodeBench (local)
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
Howdy - I posted a benchmark here - https://www.reddit.com/r/LocalLLaMA/comments/1vbtiy7/deepseek_v4_flash_on_slopcodebench/
This was using the hosted API - since then I've been playing around with quants
Here is the lastest benchmark - https://github.com/michaelasper/benchmarks/blob/main/deepseek-v4-flash-0731-pi-on-slop-code-bench.md
This uses antirez q2-q4 imatrix quant - i switched from opencode to pi
Very interesting results! Much slower on a macbook m5 max than the hosted API, but switching the harness made up for some of the intelligence lost
Compared with the other reported runs
| Reported run | Serving | Harness | Strict | Isolated | Core |
|---|---|---|---|---|---|
| DeepSeek V4 Flash 0731 (run B) | local quant (antirez, higher cap) | pi 0.84.0 | 5/17 (29.4%) | 6/17 | 10/17 |
| Opus 5 | hosted API | Claude Code | 4/17 (23.5%) | — | — |
| DeepSeek V4 Flash | hosted API | OpenCode 1.18.10 | 3/17 (17.6%) | 6/17 | 11/17 |
| Opus 4.8 | hosted API | Claude Code | 1/17 (5.9%) | — | — |
| Sonnet 5 | hosted API | Claude Code | 1/17 (5.9%) | — | — |
| DeepSeek V4 Flash 0731 (run A) | local quant (unsloth, misconfigured cap) | pi 0.84.0 | 1/17 (5.9%) | 1/17 | 2/17 |
[link] [comments]
More from r/LocalLLaMA
-
Unpopular opinion Qwen 3.8 is hard to understand
Aug 30
-
Qwen3.8-Flash-Next turns 4xR9700 into a local AI powerhouse! 120 t/s TG and 12k t/s PP single request with optimized vLLM
Aug 30
-
Oh so that's where my PCIe lanes went...
Aug 30
-
Uncensored Multi-Model Releases, LongCat-Flash-Lite-Sparse with MTPs and LSAs, Qwen3.8-27B with MTPs, Qwen3.5-122B-A10B with MTPs, Qwen3-Coder-Next and Laguna-S2.1 with Vision, All Available in GGUF Format! Bonus: Links to my llama.cpp Fork for LongCat-Flash-Lite Support and…
Aug 30
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.