News / #model-release Tag Model releases 500 articles archived under #model-release · RSS Sign in to follow r/LocalLLaMA community 3h ago an unscientific qwen 3.8 flash next and glm 5.3 flash comparison I stole the reference image from a recent post on r/stablediffusion , and then asked both qwen 3.8 flash next (q4 K XL) and GLM flash (oQ4e MLX) to choose try to reproduce it into a "video game or tech demo" as closely as possible, iterating over a period of (up to) about an… 9 r/LocalLLaMA community 4h ago Ran Qwen3.8-Flash-Next (79 GB, 2-bit) at 350K ctx for 3.5 hours on a 128 GB M5 Max — speed vs context depth, 100 turns, one graph Setup: MacBook Pro M5 Max, 128 GB unified, macOS 26.5.2 · llama.cpp b10686 (Metal, 12 threads, batch 2048, flash-attn, kv-unified, ngram-mod spec decode) · Qwen3.8-Flash-Next UD-Q2_K_XL (Unsloth), 78.9 GB · 358,400-token context slot via YaRN from the native 262,144, fp16 KV.… 13 Simon Willison community 8h ago Introducing Hy4 Preview Introducing Hy4 Preview New open weight text input (no vision) LLM from Chinese company Tencent today: 770B total parameters, 49B active parameters, 1M token context window, 1.56TB on Hugging Face . This is a big size increase from their previous Hy3 in July, which was 295B, 21B… 9 r/LocalLLaMA community 9h ago Humaneval benchmark for Deepseek V4 Flash 0731 vs GLM5.3 Flash on 2x DGX Spark setup I have a 2 DGX Spark setup recently and I have been happily running Deepseek V4 Flash 0731. Since the release of GLM5.3 Flash and Qwen 3.8 Flash Next this week, a lot of folks are still waiting to see what model to run given their own hardware situations. I am very interested in… 33 r/LocalLLaMA community 13h ago Any current Voice2Voice AI model that runs locally that’s good? You guys remember sesame AI? With their really good AI voice model? Obviously ChatGPT has their voice model that’s also really good. Is there any smaller local variant that runs on like consumer grade gpu‘s (12,16 24gb?) I think NVidia released something but I didn’t really… 25 r/LocalLLaMA community 13h ago Tenstorrent Qwen3.7-27b Benchmarks I saw someone here posted about getting a Tenstorrent QuietBox 2, and I wanted to look into it the hardware. It's very difficult to find any benchmarks, but I managed to find some from an employee. The machine it was benchmarked on has 2 p300c's, their top of the line card,… 28 Hacker News — AI on Front Page community 13h ago Hy4 preview Article URL: https://www.tencent.com/tencent-releases-and-open-sources-tencent-hy4-preview/ Comments URL: https://news.ycombinator.com/item?id=49492632 Points: 206 # Comments: 130 13 r/LocalLLaMA community 13h ago 67-84 t/s DeepSeek flash v4 off 2x GX10s Finally achieved usable results with 2 gx10 at over 65 tokens a second sustained. The 2570 prompt eval is really crucial for me as well. Overall stoked 10/10 edit: I followed this setup with 2 ASUS GX10 DGX computers :)… 29 llama.cpp releases dev-tools 14h ago b10687 opencl: use a better matmul path on two Adreno GPU generations ( #27640 ) opencl: default the Adreno xmem F16xF32 GEMM on for X2E kernel_mul_mm_f16_f32_l4_lm is the slowest matmul this backend has on Adreno: on the X2-90 it runs the gpt-oss-20b attention projections at roughly a… 27 r/LocalLLaMA community 14h ago (NInfer Fork) I wanted to have a 1M context Qwen-3.8 27B, tp2, dual 5090s Hey! I forked NInfer (a from-scratch C++20/CUDA inference engine for Qwen models) and added two things: tensor-parallel across two GPUs, and YaRN ×4 rope scaling. Together they let Qwen3.8-27B NVFP4 run a 1,048,576-token context on two consumer 5090s — 27.4 GB per card, no… 21 r/LocalLLaMA community 14h ago This finance-model benchmark card is more useful for what it discloses than for who "wins" The official benchmark card for Ling-3.0-flash-Fin is a useful reminder that the unit being tested is rarely just “the model.” The release says most runs used temperature 1, top_p 0.95 and the highest available reasoning effort. FinFIRST and FinSearchComp Verified used a common… 22 r/LocalLLaMA community 17h ago Qwen3.8 Flash Quants ~20–30GB smaller than Unsloth/AesSedai Q4 at similar PPL. After several days of testing I released a set of mainline-compatible imatrix quants for Qwen3.8-Flash-Next. Goal: same quality band as the popular Unsloth / AesSedai Q4 builds, less disk and RAM. Savings are roughly… 32 r/LocalLLaMA community 19h ago Qwen 3.8 27B at 50 tok/s with 100k Context on a 16GB GPU! (beellama.cpp) I wanted to share my successful setup for running a Qwen 3.8 27B model with a massive context window on a consumer 16GB GPU (RTX 4070 Ti SUPER). The goal was to fit everything into VRAM without sacrificing quality or speed. 🧠 Key Components Model:… 38 r/LocalLLaMA community 20h ago How important is it for Chinese LLMs to reach the Opus 4.8 level? In mid-August, Ramp published spending data collected from 70,000 U.S. companies: Fable 5 ,the most powerful and expensive model in Anthropic’s lineup, accounts for just 11% of what those businesses spend on the company’s tools. The remaining 79% is worth its weight in gold.… 15 r/LocalLLaMA community 22h ago Qwen3.8-Next streaming - 150tps prefill, 3.6 tps decode on M5 Air Out of curiosity, I thought I'd see if I could adapt my DSv4 streaming stack from a few weeks ago to take Qwen3.8-Next. It worked, better than I thought - it actually runs faster on my 32GB M5 than the dense 27b does (admittedly not apples to apples as I decided to use a 3bit of… 22 r/LocalLLaMA community 1d ago Different Qwen thinking levels   submitted by   /u/Tall_Abrocoma_3533 [link]   [comments] 27 r/LocalLLaMA community 1d ago Terminal Bench 4.0 just dropped, GLM-5.3 is at the same level as Fable 5, accounting for margin of error Announcement: https://www.tbench.ai/news/terminal-bench-4-0 Leaderboard: https://www.tbench.ai/ Imo the best aspect in their announcement is their focus on rapidly iterating on TerminalBench to keep the pace up with new model releases to fight benchmark saturation. On a similar… 20 r/LocalLLaMA community 1d ago ~ 2x Speed Boost for Qwen3.8 27B on Apple Silicon https://x.com/koc_z3/status/2093581036756025744?s=46 ~ 2x speed boost for Qwen3.8 27B on Apple Silicon ~ 1.5x speed boost for Qwen3.6 35B AЗB Tested on an M1 Max 64GB Mac using MTPLX with 262K (MAX) Context length. Qwen3.8-27B (Q4): - Decode ~ 21 TPS - Prefill ~ 83 TPS (Peak 111… 6 r/LocalLLaMA community 1d ago Qwen3.8-27B vs Qwen3.8-Flash-Next smaller quant? If you only had 128gb ram which one would be more "intelligent", Qwen3.8-27B (or even 3.6) or a smaller quant of Qwen3.8-Flash-Next (Q4/Q5) ? Mostly for discussions, but also interested in coding. Thanks edit: I have a 128gb Halo. By "intelligent" I mean more intelligent answers… 35 r/LocalLLaMA community 1d ago Qwen3.8-Flash-Next + MTP on Strix Halo: Vulkan Runtime Notes Below are the benchmark results for running Qwen3.8-Flash-Next on Strix Halo using the Vulkan backend of llama.cpp, combined with MTP model. Hardware Item Details CPU AMD Ryzen AI MAX+ 395 (16C/32T) GPU Radeon 8060S (integrated, RADV STRIX_HALO) RAM 128GB unified memory Software… 27 r/LocalLLaMA community 1d ago Linux Ubuntu 26.04 LTS & 26.04.1 upgrade - How much Improvements? People who moved from older version to this new one, how much improvements do you see on inference? Ex: llama.cpp performance? This new version comes with Linux Kernel 7.0. Hope AMD cards gonna enjoy additional improvements with their recent ROCm 10.0 version release.  … 14 r/LocalLLaMA community 1d ago Saved my fiances phone with qwen 3.8 27b this model is really something incredible, saved us like 600 dollars. my fiance is always breaking her electronics and then getting me to fix them. The other day she brought me her phone and it was stuck in a boot loop that wouldn't post. tried normal stuff, managed to get it to… 26 r/LocalLLaMA community 1d ago 50% tg increase with offloading "hot" experts to VRAM I got a 50% performance boost (20 t/s -> 30 t/s) in llama.cpp for MoE models that don’t fit entirely in VRAM—in my case, Qwen 3.8 Flash Next. The idea is simple: instead of offloading entire layers to the GPU, I offload only the “hot” experts. I found that certain groups of… 34 r/LocalLLaMA community 1d ago use llms to auto annotation your dataset locally hi i make tool for this called llmog it's purpose to make llms free to - auto annotation datasets - reclassification existing yolo datasets running totally local using llama cpp or vllm or use external api you'd rather click than code. 🔗 GitHub: mohamed-em2m/llmog: framework… 36 r/LocalLLaMA community 1d ago Qwen 3.8 Flash Next ngram look up table offloaded to SSD and streamed in SGLang Have anyone tried this yet? Looks promising, seems too good to be true with no performance loss.   submitted by   /u/Easy_Werewolf7903 [link]   [comments] 5 r/LocalLLaMA community 1d ago AtomicChat/Qwen3.8-Flash-Next-GGUF is Really Good specs hardware: M4 Max 128GB Studio inference engine: llama.cpp (qwen4exp branch) judge: claude-opus-4-6 AtomicChat/Qwen3.8-Flash-Next-GGUF Qwen3.8-Flash-Next is a great model I benched in my previous post , but it is very tight, since all n-grams / PLE are loaded along with the… 17 r/LocalLLaMA community 1d ago Hot or not? Does anyone else add active cooling to their DGX stack? Found mine was getting quite hot under extended load. This helps immensely with that so far. I will be adding some stats as they relate to comphy and DeepSeek flash this weekend. I did not create the original designs but… 27 r/LocalLLaMA community 1d ago Is it worth running Qwen 3.8 Flash Next on 4x3090 vs 27B? Can someone please tell me if it's worth running Qwen 3.8 Flash Next on 4x3090 yet over 27B? 27B is good but damn it is indecisive. I am getting frustrated watching it get "so close" to solving a problem, only to do another 2 hours of "let me just check/prove/etc" It looks like… 30 r/LocalLLaMA community 1d ago Today I hit 181 toks/s (aggregate) on Qwen3.8-Flash-Next on 2x DGX Sparks Hey all, and hello fellow DGX Spark-ers! Today I managed some pretty crazy numbers: 181 tok/s aggregate on 2× DGX Spark on Qwen3.8-Flash-Next at 512Kcontext (2.8M kvc) I hit 181 tok/s aggregate today across a multi-agent fleet on a 2-node DGX Spark cluster. Single-stream decode… 22 r/LocalLLaMA community 1d ago [Release] SOTA GGUFs for Qwen3.8-27B: GSQ-RCO at 2.5 to 3.0 bpw We're releasing Qwen3.8-27B quantized with our newest methods, GSQ + RCO. Higher-quality models, same file size, now with the search and the quantizer both learned. What's inside: GSQ (Gumbel-Softmax Quantization): post-training scalar quantization that jointly learns grid… 32 r/LocalLLaMA community 1d ago I benchmarked 9 open models on spotting fake sources during agentic search (DeepSeek V4, Qwen 3.8, Nemotron 3 Ultra) I built a benchmark called EchoNet. An agent gets a factual question, then searches a syntheic web I made before answering. Some of that web is seeded with misinformation: one fake page, a fake page ranked first in search results, the same fake claim copied across many pages, a… 27 llama.cpp releases dev-tools 1d ago b10678 model: qwen4exp: reduce number of graph splits ( #27880 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/43734155 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS… 23 r/LocalLLaMA community 1d ago ROCm 10.0: A Decade of Open Compute, Built for the Age of Agentic AI Their last version 7.14 was released just a month ago. llama.cpp PR(waiting for approval) for Version 10.0 https://github.com/ggml-org/llama.cpp/pull/27803 Hope this version comes with more boost & improvements.   submitted by   /u/pmttyji [link]   [comments] 35 r/LocalLLaMA community 1d ago Local agentic coding Benchmark : Qwen3.8-Flash-Next NVFP4 vs 27B (and the others...) Using https://huggingface.co/RadixArk/Qwen3.8-Flash-Next-NVFP4 and https://old.reddit.com/r/BlackwellPerformance/comments/1w04xb7/qwen38_flashnext_on_1x_rtx_pro_6000_171_ts_c1_428/ As usual, all the details in… 34 r/LocalLLaMA community 1d ago Qwen3.8-Flash on RTX3090 + 64GB RAM (but you only need 12GB VRAM) I've got Qwen3.8-Flash-next running on RTX 3090, Ryzen 9 3950X, a PCIe 3.0 motherboard, and 64GB DDR RAM from 2020. IQ4_XS weights, full kvarn5 context, vision on GPU, experts in host RAM, n-grams on disk. MTP works but actually slows decode down even with 80% draft acceptance,… 18 llama.cpp releases dev-tools 1d ago b10672 OpenVINO: Update OV to 2026.3.1, whisper.cpp support, Qwen3.5 on NPU, and new ops ( #27843 ) OpenVINO Backend: Fuse IM2COL + MatMul convolution into OpenVINO convolution ci:ggml-ov: Skip recurrent state rollback tests ci:ggml-ov: Skip recurrent state rollback tests Update… 33 Hacker News — AI on Front Page community 1d ago Htmx 4.0.0 Article URL: https://four.htmx.org/announcements/2026-08-28-htmx-4.0.0-is-released Comments URL: https://news.ycombinator.com/item?id=49478178 Points: 255 # Comments: 56 37 r/LocalLLaMA community 1d ago Qwen3.8-27b q8 KV cache does seem to actually hurt model performance EDIT: Though the issue with q8 kv cache seems to arise from when and how often we run the quantize step, not that kv quantizing can't ever work - see comments --- One of the things I see debated a lot is whether to use kv cache quantization. The idea I see a lot is that q8… 28 llama.cpp releases dev-tools 1d ago b10669 sycl: bind the f16 KV cache in place for the oneDNN SDPA path ( #27468 ) Measured at a live KV length of 34816 (32768 depth plus one 2048 ubatch), on Qwen3.8 27B Q4_K_S: per tensor 4 * 34816 * 256 * 2 B = 71.3 MB staged per call K and V, so 2x = 142.6 MB traffic per call read… 9 r/LocalLLaMA community 1d ago After Meta avocado we get watermelon, due in November Quote: ... developing a new A.I. model intended to be as powerful as Anthropic’s cutting-edge models. ... In July, while developing Watermelon, Meta paused and later resumed a stage of A.I. development called “pretraining,” which delayed its release until at least October, four… 38 r/LocalLLaMA community 1d ago Qwen3.8-Flash-Next (UD-IQ4_XS) on 2x RTX 3060 + 7800X3D, from initial 36 tps prefill to 400 tps and other benchmarks (-sm tensor trap) + VRAM/RAM usage TL;DR: llama.cpp with --load-mode mmap used 21-32 GB of RAM, with ik_llama.cpp using 106-108 GB. -sm tensor killed my prefill, changing to -sm layer went from 36 tps to 135 tps. After that, -ubatch 2048 pushed it up to 400 tps at -c 131072 . I can't fit -ubatch 2048 at -c 262144… 24 The Information — AI news-outlet 1d ago Tencent’s New Flagship AI Model Shows Major Progress Chinese tech giant Tencent Holdings on Friday launched a preview version of its new flagship open-source model, Hy4, which demonstrates a significant improvement in performance from its predecessor. Benchmarks and early feedback on Hy4 suggest that Tencent is emerging as a more… 27 r/LocalLLaMA community 1d ago TontaubeV1 - Open TTS model release for local long-form generation Hey everyone, I am the co-founder of Tontaube. My brother and I just released TontaubeV1, a 2.9B-parameter open-weight TTS model focused on expressive speech, long-form generation, and low-latency local inference. The model is primarily aimed at English and German and supports… 36 The Information — AI news-outlet 1d ago Z.ai’s Latest Model Intensifies Competition for Low-Cost Offerings Chinese AI firm Z.ai ’s new open-source model is generating a lot of buzz because of its near-frontier capabilities at super low costs. This further intensifies the price competition in the global AI model market. The new GLM-5.3 Flash model, officially released earlier this… 29 r/LocalLLaMA community 1d ago claude mods didn't like that, somehow 🤷♀️   submitted by   /u/peculiar-ragdoll [link]   [comments] 37 llama.cpp releases dev-tools 2d ago b10666: tests : run test-save-load-state across all architectures (#27755) tests : run test-save-load-state across all architectures test-save-load-state previously only ran in ctest against a single downloaded model (tinyllamas/stories15M), i.e. only the llama arch. Add a --models DIR mode to test-save-load-state that runs the full save/load suite… 20 Hacker News — AI on Front Page community 2d ago Sovereign Tech Agency invests €500k in Flatpak Article URL: https://modal.cx/blog/announcing-flatpak-sta/ Comments URL: https://news.ycombinator.com/item?id=49474786 Points: 206 # Comments: 120 33 r/LocalLLaMA community 2d ago I reverse-engineered an NPU vendor's engine format (int8 weights stored as two nibble planes) to run GGUFs with no model conversion — now 1.5× faster than the vendor's own runtime I've been running Qwen3-0.6B on the M5Stack LLM-8850 card (Axera AX8850 NPU, 24 TOPS, 8GB LPDDR4x) hosted by a Raspberry Pi 5 — as a llama.cpp backend. The problem: the vendor stack requires converting every model through their compiler, and their closed runtime gets 13.5–14.5… 28 r/LocalLLaMA community 2d ago Ornith 1.5 is actually pretty good hey guys i recently started using ornith 1.5 to rapidly test some tools im working on since qwen 3.8 27b was too slow for my testing loop. This model is actually really good. im getting around 130 tokens / second with mtp and its very good at tool calling. i feel like this is… 32 arXiv — Machine Learning research 2d ago When Privacy Hurts Mergeability: Geometry-Aware Model Merging under Differential Privacy arXiv:2608.26655v1 Announce Type: new Abstract: Model merging promises to construct a single multi-task model from independently fine-tuned task models without accessing the original task data. This makes it attractive when task data cannot be centralized, but released task… 18 Page 1 of 10 · 500 articles Older →