r/LocalLLaMA
500 articles archived · Visit source ↗ · RSS
-
-
-
-
-
-
-
-
-
-
-
r/LocalLLaMA community 18h ago
Qwen3.8 Flash Quants
~20–30GB smaller than Unsloth/AesSedai Q4 at similar PPL. After several days of testing I released a set of mainline-compatible imatrix quants for Qwen3.8-Flash-Next. Goal: same quality band as the popular Unsloth / AesSedai Q4 builds, less disk and RAM. Savings are roughly…
32 -
-
r/LocalLLaMA community 19h ago
Someone tested various Models on the Political Compass test...
  submitted by   /u/Thrumpwart [link]   [comments]
4 -
-
-
-
-
r/LocalLLaMA community 1d ago
Different Qwen thinking levels
  submitted by   /u/Tall_Abrocoma_3533 [link]   [comments]
27 -
r/LocalLLaMA community 1d ago
I always wonder how much more speed and/or context they'd be getting..
Nothing personal. I just have too much time on my hands. Probably because I spend none of it inspecting the code my agent writes, just the finished product.   submitted by   /u/_-_David [link]   [comments]
18 -
-
-
-
-
-
-
-
-
r/LocalLLaMA community 1d ago
Hot or not?
Does anyone else add active cooling to their DGX stack? Found mine was getting quite hot under extended load. This helps immensely with that so far. I will be adding some stats as they relate to comphy and DeepSeek flash this weekend. I did not create the original designs but…
27 -
-
-
-
-
-
-
-
r/LocalLLaMA community 1d ago
Local agentic coding Benchmark : Qwen3.8-Flash-Next NVFP4 vs 27B (and the others...)
Using https://huggingface.co/RadixArk/Qwen3.8-Flash-Next-NVFP4 and https://old.reddit.com/r/BlackwellPerformance/comments/1w04xb7/qwen38_flashnext_on_1x_rtx_pro_6000_171_ts_c1_428/ As usual, all the details in…
34 -
r/LocalLLaMA community 1d ago
ds4 branch with GLM 5.3 Flash support
As a happy user of ds4, I'm very excited about this branch. Ran some prompts and it seems to be working well on my M4 Max 128gb! https://x.com/antirez/status/2093349448445243873   submitted by   /u/lakySK [link]   [comments]
32