News / #model-release Tag Model releases 500 articles archived under #model-release · RSS Sign in to follow r/LocalLLaMA community 13d ago "Opus 4.8 thinks too much", "Muse Glimmer sits between Gemma and Qwen, that's boring", "Gemma 4 is too lazy" I'm starting to think there's no way to make a reasoning model that won't draw persistent vocal complaints on here. EDIT: Qwen 3.8 not Opus 4.8*, freudian slip lol   submitted by   /u/MerePotato [link]   [comments] 29 r/LocalLLaMA community 13d ago Qwen 3.8 35bA3b wen? Artificial analysis index scores Qwen 3.5 27b: 35 Qwen 3.6 27b: 38 Qwen 3.8 27b: 52 What the hell kind of a jump was that? Even if it is benchmaxxed, the jump is insane. Qwen3.6 35b A3b: 32 That's ~6 points behind its dense 27b model, but is ~5x faster for inference given only… 29 r/LocalLLaMA community 13d ago Qwen3.8 27B > Opus 5 Medium on Artificial Analysis Agentic Index https://preview.redd.it/xh1rloaf4zjh1.png?width=1628&format=png&auto=webp&s=a536fae1b50b327f2bc55d4f4f874f94ae66e867 Thanks Qwen team!   submitted by   /u/secopsml [link]   [comments] 14 r/LocalLLaMA community 13d ago Qwen3.8 27B = GPT-5.6 Luna compressed into 27B How crazy is that?   submitted by   /u/kevinlch [link]   [comments] 37 r/LocalLLaMA community 13d ago Artificial Analysis' Qwen3.8-27B benchmarks put it neck and neck with DeepSeek V4 and GPT-5.6 Luna Max   submitted by   /u/anderspitman [link]   [comments] 30 Hacker News — AI on Front Page community 13d ago Qwen3.8 27B scores 52 on Artificial Analysis Article URL: https://artificialanalysis.ai/models/qwen3-8-27b Comments URL: https://news.ycombinator.com/item?id=49334544 Points: 245 # Comments: 112 26 Hacker News — AI on Front Page community 13d ago Cursor launches Origin, GitHub alternative Article URL: https://cursor.com/changelog/origin-code-hosting Comments URL: https://news.ycombinator.com/item?id=49334209 Points: 268 # Comments: 215 25 r/LocalLLaMA community 13d ago Why do people like coding harnesses like opencode etc instead of an IDE? Just curious - I like to be able to see and manage the scripts my agent is working on. I find stuff like Claude Code and Open Code useful for doing stuff on my linux box but I don't understand why people would use that instead of a dedicated IDE where you can actually see the… 33 r/LocalLLaMA community 13d ago Ling 3.0 Tiny is the strongest, fastest and greatest model on my low end PC! This Ling 3.0 Tiny 8b param with 1.3b active is the fastest, smartest model I can run on my poor old pc, with 4gb vram. It actually runs lightning fast, like 36 token / sec, as smart as Qwen 3.5 9b / Gemma 12, (Very close), and even faster because of 1.3b active parameters. The… 17 r/LocalLLaMA community 13d ago Qwen 3.8 27B Overthinking, It has to be done, it has to be overthinking to punch Opus 4.6 Yes, it sucks to waste time waiting on 16K+ reasoning tokens alone. But here's the thing, this is only a 27B model trying to perform on par with 1T+ parameter models. Something has to be sacrificed, and that sacrifice is the amount of reasoning or trajectory tokens. This isn't… 4 r/LocalLLaMA community 13d ago Deepseek Harnness - why is feels better Guys, could someone smarter than me explain what makes Deepseek Harness so efficient? I run it with local Qwen 3.8 (Q6). I tried Opencode/Openchamber (my favourite so far), Pi agent and Hermes. New Qwen seems to overthing by default but this could be minimized with some effort.… 16 llama.cpp releases dev-tools 13d ago v0.1.1 Release v0.1.1 19 llama.cpp releases dev-tools 13d ago b10470 ci : push release tag explicitly in release.yml ( #27261 ) Add a "Create and push git tag" step to the release job, right before the "Create release" step. The tag is created with git tag and pushed with the deploy key already configured by the Clone step, instead of relying on… 21 r/LocalLLaMA community 13d ago llama.cpp version v0.1.0 has been released llama.cpp is apparently moving to semantic versioning instead of just sequential build numbers (like b10456). The first semantic version tag was created today: https://github.com/ggml-org/llama.cpp/releases/tag/v0.1.0 Congrats to llama.cpp on version v0.1.0!   submitted by… 29 r/LocalLLaMA community 13d ago EXL3 seems to be fading from the r/LocalLLaMa consciousness, and while I suspected it, I'm surprised at this point in time. EXL3 is an alternative to llama.cpp. And while there is extensive tooling for llama.cpp, EXL3's primary deployment ( TabbyAPI ), has a OpenAI compatible API so it shouldn't matter. Why won't this tool matter to you? If you have a GPU with under 24 GB of VRAM, the value kind of… 27 r/LocalLLaMA community 13d ago After pushing 1M+ tokens through Qwen 3.8 27B, here is my optimal llama.cpp config for 16GB VRAM (73k Context, Agentic Coding) Following up on my previous post about my budget server setup (Intel N100 + RTX 5060 Ti 16GB), a few of you asked for a deeper dive into my actual inference config and real-world agentic performance. Like many of you, I was refreshing the page waiting to download Qwen 3.8 27B… 31 r/LocalLLaMA community 13d ago tencent/EVIE-Preview-4.5B · Hugging Face Overview EVIE-Preview-4.5B is a state-of-the-art multilingual Visual Document Retrieval (VDR) model built upon Qwen3.5-4B . It employs ColBERT-style late interaction with native 128-dimensional multi-vector token embeddings (4.54B parameters, BF16). By combining native… 37 Hacker News — AI on Front Page community 13d ago GPT 5.6 Sol is the best "vision" model OpenAI ever released Article URL: https://blog.roboflow.com/openai-gpt-5-6/ Comments URL: https://news.ycombinator.com/item?id=49329575 Points: 216 # Comments: 108 16 r/LocalLLaMA community 13d ago 100$ worth of gpu runs qwen 3.8 27b at 7.39 t/s Qwen 27b Q3_K_M 2x rx 580 8gb (~50$ each in my country, edge cases 60$ per gpu) gives us 16gb vram We used it on an old already existing ddr3 motherboard with 2 gpu slots(you can buy it ror around 200$ with 32 gb of ddr3 ram, a workstation xeon cpu and a workstation motherboard,… 20 r/LocalLLaMA community 13d ago Unpopular opinion : Qwen 3.8 27b is not an overthinker Yes it uses a ton more reasoning tokens than 3.6 did But test in on the same tasks with the other chinese models, glm 5.3, deepseek v4 flash and pro, etc it's really similar, and they are needed The reality is, we're just frustrated because our hardware do not allow most of us… 19 r/LocalLLaMA community 13d ago Petition to add a rule for people to add their DAMN quant levels to their posts Every time I see a post about a newly released model, whether it be a comparison or shitting on it, I have to dig through the endless comments to see what quants they used and what their specs were. Its quite a common occurrence here in this sub to ask someone that's saying a… 25 llama.cpp releases dev-tools 13d ago v0.1.0 Release v0.1.0 17 r/LocalLLaMA community 13d ago Qwen3.8-27B Q8_0 on Strix Halo is seriously impressive Sorry for the slop, but I was impressed by this model as I have been testing Qwen3.8-27B Q8_0 locally on my ROG Flow Z13 (Ryzen AI Max+ 395, 128 GB unified memory) and this model was the only one who could made this short simulator (and I have tested a lot of models). Prompt:… 16 r/LocalLLaMA community 13d ago Long Review: Qwen 3.8 27B is VERY good at tapping into it's real-world knowledge. It's "overthinking" brings it to Sonnet level performance with the potential for Opus level results. Hi all! I finally just got around to testing out Qwen 3.8 27b. I'm using Unsloth's UD-Q8_K_XL quant as a sit-in replacement to Qwen 3.6 27b, same quant size. Wow -- this thing isn't messing around. I have many baseline test prompts to gauge the 'intelligence' and usability of… 28 llama.cpp releases dev-tools 13d ago tmp-testing-0 ci : make release workflows use a deply key 19 llama.cpp releases dev-tools 13d ago b10456 sycl: fix thread/block count in quantized cpy kernel launches ( #27160 ) Adjusts the thread/block count to be proportional to the size of the quant, reducing under/over subscription. Largest perf improvement is the q4_0 -> f32 path, with, on a Arc 70, throughput goes from 20.21… 10 r/LocalLLaMA community 13d ago How many tokens/second output are you getting with Qwen3.8-27B? Trying to get a feel for where I stand. If you can list your relevant hardware and model used, that would be awesome. Here's mine: Model: Qwen3.8-27B-heretic-ara, Q5_K_M GGUF T/s : ~30-32 t/sec (I think, I'll verify in a bit) Hardware: 3090 GPU | 64 GBs DDR4 RAM | AMD 7950x CPU… 17 Vercel — AI dev-tools 13d ago GPT-5.6 Sol is 50% off on AI Gateway for the next month GPT-5.6 Sol , the flagship of OpenAI's GPT-5.6 series, is 50% off on AI Gateway through September 18. The discount applies on the OpenAI provider to all token types, tiers, regions, and modes, and it is available only on requests running directly through AI Gateway (not BYOK).… 6 r/LocalLLaMA community 13d ago [audio.cpp] Release 0.6: dots.tts, MiniMax-H3 text2audio (up to 3x realtime), MiniMax-Music3 (preview), and more new audio models. 5+ demos included. Hi all :) audio.cpp release 0.6 has been out for a little while, so this is more of an update on what landed and what has been improving around it. 0.6 added 5 new model families: dots.tts, NeuTTS-2e, MuScriptor (Music to MIDI), MiniMax-H3, and SenseVoice-Small, bringing… 25 r/LocalLLaMA community 13d ago Simon Willison: Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things I look to Simon for a broad survey of current LLM tech. Here's his review of playing with Qwen 3.8 27B . His comment on Mastodon was "I can't remember the last time I've had this much fun playing with a local model that runs on my own computers". BTW, the "wildly overthinking"… 30 r/MachineLearning community 13d ago It only took 200 update steps to flip Qwen2.5-7B-Instruct from denying sentience to developing a robust identity of being a "sentient machine" [P] First, I want to clarify that I am not claiming that LLMs are sentient. Basically all of my behavioral descriptions are anthropomorphizations to make communicating my results easier. For fun, I decided to post-train Qwen2.5-7B-Instruct to develop a generalizing self-belief of… 4 r/LocalLLaMA community 13d ago Qwen 3.8 27b vs 3.6 27b - how good is with a Turtle library. Prompt: Provide complete working code for a realistic looking tree in Python using the Turtle graphics library and a recursive algorithm. Difference between 3.6 and 3.8 is huge!   submitted by   /u/Healthy-Nebula-3603 [link]   [comments] 6 Simon Willison community 13d ago Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things Friday's big release was Qwen 3.8 27B , an Apache 2 licensed 27B parameter vision-capable LLM from Alibaba's Qwen research lab. I've been looking forward to this one: 27B is an excellent size for running a model on a reasonably specced laptop, and its predecessor Qwen 3.6 27B… 17 Hacker News — AI on Front Page community 13d ago Anthropic's 'watermark' text adulteration in Claude is a perversion of writing Article URL: https://daringfireball.net/2026/08/anthropics_watermark_text_adulteration_in_claude_is_a_perversion_of_writing Comments URL: https://news.ycombinator.com/item?id=49324087 Points: 359 # Comments: 344 4 r/LocalLLaMA community 13d ago Dario Amodei defends his policy proposals, warns open weights won't decentralize power, endorses pre-launch vetting, says real accomplishments will earn trust   submitted by   /u/f0urxio [link]   [comments] 17 r/LocalLLaMA community 13d ago Anyone else get a kick out of Qwen 3.8 27B Reasoning Dialogue? I've been paying attention to the reasoning because I'm still evaluating the model, and I've just noticed that sometimes I get a kick out of the way this model's internal monologue seems to play out sometimes. Like I've seen it get genuinely frustrated with itself and express… 27 r/LocalLLaMA community 13d ago Qwen 3.8 9b?   submitted by   /u/Thatisverytrue54321 [link]   [comments] 25 r/LocalLLaMA community 14d ago Qwen3.8-27b on RTX 3090 - 82 tps single request, up to 672 tps peak Hi, After a long night of optimizations, I believe I have made the fastest inference engine for Qwen3.6-28B on a 3090. Quick metrics: - 250w power capped - Up to 195k context (ships with 150k for safety though) - 82 tps single request, 417 tps sustained with 64 concurrent -… 23 r/LocalLLaMA community 14d ago Qwen3.8 27B Q2 vs Q3 vs Qwen3.6 35B-A3B MoE on 12GB VRAM Did a quick local test because I wanted to see what is actually usable on my 12GB laptop GPU. I tested the newer Qwen3.8 27B dense files at Q2 and Q3, then compared them against Qwen3.6 35B-A3B MoE. Hardware: RTX 5070 Ti Laptop, 12GB VRAM Backend: llama.cpp CUDA Settings: 4k… 12 r/LocalLLaMA community 14d ago Qwen 3.8 2.4T at 288k tokens/s on Nvidia GB300 NVL72 https://developer.nvidia.com/blog/serve-qwen3-8-2-4t-a95b-a-2-4t-parameter-model-with-configurable-reasoning-on-nvidia-gb300-nvl72/ 4k tokens per second per GPU of which there are 72. 350 tokens per second per user "Without additional model tuning, the model achieves a… 13 r/LocalLLaMA community 14d ago Qwen 3.8 distillations https://x.com/i/status/2088993948983246906 Not tested by me in any way :)   submitted by   /u/jacek2023 [link]   [comments] 13 r/LocalLLaMA community 14d ago Koboldcpp v1.119 released   submitted by   /u/Fcking_Chuck [link]   [comments] 7 r/LocalLLaMA community 14d ago Qwen3.8-27B Hybrid IQ4_XS quantization for 16GB gang   submitted by   /u/Johnny_Rell [link]   [comments] 11 r/LocalLLaMA community 14d ago Newer commits removed the Qwen 35B In this commits, the 35B model was removed. Looks like it's confirming the 35B model won't get released. I think they need to be made aware how big the 35 moe is widely used. Think need to make noise on theyre X, huggingface and online places. If they dont know there's no need… 24 r/LocalLLaMA community 14d ago Quick PSA: Qwen3.8-27B reasoning effort vs reasoning budget in llama.cpp If you are using llama-server with their web-ui for testing, keep in mind, that the reasoning selector is just a reasoning budget aka a hard cap and has, at least to my knowledge, nothing at all to do with Qwen3.8-27B's native reasoning effort capability! Selecting any value for… 34 llama.cpp releases dev-tools 14d ago b10454: ci : fix dry-run reporting in make-release job [no ci] (#27167) This commit fixes the reporting in the make-release CI job when --dry-run is used. It will currently incorrectly report that all checks pass even if there are steps that fail. Refs: #26839 (comment) 34 Hacker News — AI on Front Page community 14d ago Claude: System Prompts Article URL: https://platform.claude.com/docs/en/release-notes/system-prompts Comments URL: https://news.ycombinator.com/item?id=49319556 Points: 259 # Comments: 129 20 r/LocalLLaMA community 14d ago Huihui-ai Qwen 3.8 Ablit Available At hugging face, this is the model I use for most of my analysis that most of my services have to not be refused. previous versions are quite good. It seems this might have dropped today and pulling right now!   submitted by   /u/Frizzy-MacDrizzle [link]   [comments] 38 r/LocalLLaMA community 14d ago Qwen3.8-27B-int4-AutoRound (18GB) - with working MTP spec decode   submitted by   /u/BusinessMud9586 [link]   [comments] 25 r/LocalLLaMA community 14d ago Qwen 3.8 27b with DSH(DeepSeek Harness) is Amazing!! Experiences so far and perfomance. https://preview.redd.it/wkg27e152qjh1.png?width=853&format=png&auto=webp&s=2e3f8b11ea6393041f501e95c5835f9bea0245dd So ive been trying different harnesses and coding agents with the new qwen 3.8 , and after trying out many ive been mostly impressed by deekseek harness , paired… 33 Page 10 of 10 · 500 articles ← Newer