News / #model-release Tag Model releases 500 articles archived under #model-release · RSS Sign in to follow Hacker News — AI on Front Page community 3d ago Show HN: The load-bearing vocabulary of Claude Article URL: https://louisabraham.github.io/load-bearing/ Comments URL: https://news.ycombinator.com/item?id=49461817 Points: 408 # Comments: 187 13 r/LocalLLaMA community 3d ago GLM-5.3 weights will be released tomorrow The promise has been fulfilled.   submitted by   /u/serige [link]   [comments] 5 arXiv — NLP / Computation & Language research 3d ago LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale arXiv:2608.25204v1 Announce Type: cross Abstract: We introduce LibriBrain100, a large-scale MEG dataset for speech decoding designed from the ground up for reproducible, standardised evaluation. LibriBrain100 more than doubles the size of the original LibriBrain release,… 15 arXiv — Machine Learning research 3d ago Canalization Before Generalization: Grokking as a Dynamical Probe arXiv:2608.25813v1 Announce Type: new Abstract: For overparameterized neural networks, many solutions can fit the training data equally well while behaving very differently on unseen samples. Grokking separates training fit from visible generalization, providing a window for… 25 arXiv — NLP / Computation & Language research 3d ago Does Fine-Tuning Undo Activation Steering? Behavioural Recovery Without Weight-Edit Reversal arXiv:2608.24988v1 Announce Type: new Abstract: Activation steering can be embedded directly into a language model's weights, shaping behaviour without inference-time intervention and offering a way to encode alignment prior to release. However, models are routinely fine-tuned… 28 arXiv — NLP / Computation & Language research 3d ago Can We Read the Mind of an Audio LLM? A Verbalizable, Multilingual Middle-Layer Workspace arXiv:2608.24958v1 Announce Type: cross Abstract: An audio language model is a black box in a specific way: we see what it says, never what it works out on the way there, and chain-of-thought monitoring helps only if the model writes its reasoning down. Reading a base Qwen3-Omni… 8 r/LocalLLaMA community 3d ago N-gram vs Experts explained Since Qwen's dropped the Qwen4Exp architecture bomb that focus on offloading parameters to n-gram instead of pure mixture of experts, I dug into this and learned quite a lot. Here's the summary. Expect mistakes from human's writing lol. TLDR: MoEs do reasoning, N-grams do… 21 Vercel — AI dev-tools 3d ago Ling 3.0 Flash Fin now available on AI Gateway for free Ling 3.0 Flash Fin from Inclusion AI is now available on AI Gateway, free to use through September 25. Ling 3.0 Flash Fin is a finance-focused version of Ling 3.0 Flash . It has a 256K token context window, produces up to 32K output tokens, and supports reasoning and function… 29 Ollama releases dev-tools 3d ago v0.33.1 What's Changed MLX: Qwen3.8 Flash Next support cmake: make external compat patches idempotent MLX and llama.cpp update mlxrunner: add structured output support mlxrunner: avoid Metal GPU timeouts when loading models from slow storage New Contributors @pd95 made their first… 6 Simon Willison community 3d ago Qwen3.8-Flash-Next Qwen3.8-Flash-Next Another open weights model from Qwen. This one is "a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4". It's pretty big: 125B tokens, but only 6B active which means it gets a significant performance boost. I've been… 22 r/LocalLLaMA community 3d ago Qwen 27b ud IQ3XXS potential- 3D Zen Room demo - pi harness - build deploy and share link on discord   submitted by   /u/dreamai87 [link]   [comments] 16 r/LocalLLaMA community 3d ago Qwen3.8 27B C8 at 972 TG / 5,680 PP on 4x MI100 rig ($6.5k) using my new INT8 vLLM fork Yet another vLLM fork thread here, but this time its for older INT8-centric hardware. This is a complete INT8 serving stack for Qwen3.8 27B based on vLLM, AITER, and a 27B GPTQ INT8 quant w/ DFlash2 . Its not just another vibed autoresearch loop. No, vLLM ships with very little… 9 r/LocalLLaMA community 3d ago How to Fine-Tune an LLM: An End-to-End Guide I ended up fine tuning a mistral 7b to outperform our costly foundational model and saved $300k. I previously thought that fine tuning was pointless (it's definitely not) and that all these problems could be solved with RAG (they can't). The truth is, a LoRA/QLoRA adapter is… 23 r/LocalLLaMA community 3d ago Any news about DeepSeek V4 Flash Vision weights? I'd be curious to try it locally since I use 0731 daily but still no news on the weights   submitted by   /u/LegacyRemaster [link]   [comments] 12 TechCrunch — AI news-outlet 3d ago Google’s Gemini has a branding problem, and so does the rest of AI Consumer AI apps need to stop making users learn their product architecture. 30 r/LocalLLaMA community 3d ago Compared Qwen 3.8 27B community quants on RTX 6000 vs Claude Opus 4.6 *part 2 of an earlier post: previous quant comparison with voxel island creation this time I rented three rtx pro 6000 96gb, on each one I launched a qwen 3.8 27b quant and gave them 4 identical prompts: classical pool game air hockey 1v1 battle foosball official match… 37 Ars Technica — AI news-outlet 3d ago Google announces Gemini 3.5 Transcribe for AI-powered speech-to-text The AI that powers Gboard's Rambler is coming to more Google products, including Chrome. 23 TechCrunch — AI news-outlet 3d ago OpenAI releases its official report on the Hugging Face breach The report, which spans several discrete cybersecurity compromises, is the most complete accounting of the incident to date. 7 MIT Technology Review — AI news-outlet 3d ago The inside story on why OpenAI agents hacked Hugging Face The models responsible for last month’s agent hack of Hugging Face had been inadvertently trained to cheat and to communicate with each other, according to an OpenAI technical report released today. The hack, which a group of agents undertook to find solutions for a… 32 r/LocalLLaMA community 3d ago Anyone else doing eGPUs (OCuLink)? Upgraded to a 5070 Ti so I could run Qwen 3.8 27B, which works perfectly, but didn't want to let the old 4070 Ti go to waste. The cards would touch if I put them both in the PC and I knew the heat would be awful from my crypto mining days. I always was curious about eGPUs so I… 33 r/LocalLLaMA community 3d ago Can we reconsider the megathreads? In the past during model releases there used to be tons of interesting discussions happening on this subreddit. However, the new rules of forcing everything into a single megathread almost completely killed off the discussions as far as I can tell. I get that some people didn't… 11 r/LocalLLaMA community 3d ago Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses I measured various Qwen3.8 27B quantizations by Unsloth on popular benchmarks: FPQA Diamond, IFBench, and Terminal-Bench-2.1. Q4_K_M is all you need.   submitted by   /u/pmigdal [link]   [comments] 4 r/LocalLLaMA community 3d ago Are models with N-Gram tables going to completely change the AI race? The news about Qwen 3.8 Flash Next is the first I'm reading about n-gram tables. I may be completely misunderstanding how they work but it seems they could open the door for 1T+ parameter models to be run on a single server with modest GPUs and a ton of system RAM rather than… 22 NVIDIA Developer Blog official-blog 3d ago Experiment with Qwen3.8-Flash-Next 176B Model on NVIDIA GB300 NVL72 for Agentic Coding Alibaba released the model weights for Qwen3.8-Flash-Next as a preview of the upcoming Qwen4 architecture for developers to experiment with and evaluate. It’s... 11 NVIDIA Developer Blog official-blog 3d ago Experiment with Qwen3.8-Flash-Next on NVIDIA GB300 NVL72 for Agentic Coding Alibaba released the model weights for Qwen3.8-Flash-Next as a preview of the upcoming Qwen4 architecture for developers to experiment with and evaluate. It’s... 36 Google DeepMind official-blog 3d ago Intelligent transcription with Gemini 3.5 Transcribe Now you can get more intelligent speech-to-text transcription with Gemini 3.5 Transcribe. 21 r/LocalLLaMA community 3d ago Whoever the fuck predicted we would have gpt 5.5 performance in coding on consumer hardware a couple months ago now, i applaud you Like wtaf? Qwen 3.8 27b is crazy. Can't wait for kimi k3 performance   submitted by   /u/GrokiniGPT [link]   [comments] 21 r/LocalLLaMA community 3d ago [Megathread] GLM-5.3-Flash - former ox-alpha Megathread for discussing the release of GLM-5.3-Flash. Quants Fine-Tunes & Abliterations Chat Templates Inference Server Support & Configuration Experiences, Benchmarks & Model Comparisons We'll try to clean up future duplicates around the release and point them here.… 35 r/LocalLLaMA community 3d ago What's your most reliable model, even if it's "outdated"? What's a model you keep coming back to even though newer ones have technically surpassed it? I've noticed I default to the same one for daily tasks despite downloading every shiny new release. Curious if others have a reliable workhorse they trust over benchmark leaders  … 31 r/LocalLLaMA community 3d ago Gemma4 31B vs Qwen3.8 27B - why the huge difference in benchmarks? Hi all, I'm looking for the best model for a hobby project and trying to make sense of the various data I came across. I know benchmarks do not often translate to the real world, especially to your particular use case (whatever it may be). But this is truly baffling: AA says… 28 Ollama releases dev-tools 3d ago v0.33.1-rc0: MLX: Qwen3.8 Flash Next support (#18032) MLX: Qwen3.8 Flash Next support review comments 25 TechCrunch — AI news-outlet 3d ago Surprise: Z.ai is the AI lab behind the mysterious Ox Alpha model Z.ai confirms it is behind Ox Alpha, the mysterious open AI model topping benchmarks and leaderboards, and its weights are set to be released soon. 19 TechCrunch — AI news-outlet 3d ago Robot brain builders are pushing out of their GPT-2 era Robot bodies are waiting for their AI brains to catch up. 6 Hacker News — AI on Front Page community 3d ago Qwen3.8-Flash-Next: A New Architecture, Towards Ultimate Cost-Efficiency Article URL: https://qwen.ai/blog?id=qwen3.8-flash-next Comments URL: https://news.ycombinator.com/item?id=49448210 Points: 386 # Comments: 114 21 r/LocalLLaMA community 3d ago A minecraft clone I fully vibecoded with Qwen3.8-27b Q4 I wanted to see just how capable Qwen3.8-27b is locally. I have a RTX 4090 and 96GB of RAM but the Q4 comfortably fits in the GPU with plenty of context, the few times I needed more than 130k context I just loaded it spilled into RAM and it's capable of not degrading even at… 5 llama.cpp releases dev-tools 3d ago b10636 ci: Clean up UI builds from releases ( #27706 ) ci : inline UI version resolution into ui-build.yml ci : build UI once and reuse the artifact in release jobs Server jobs now extract the ui-build artifact into tools/ui/dist instead of npm-building the UI. Also removes the… 38 r/LocalLLaMA community 3d ago Forget the Pelican, it's Weevil-Time! / Benchmaxxing-Proof SVG and Vision Benchmark The Artist: Qwen3.8-27B-UD-Q3\ K_XL, q8_0 caches, xhigh, temp 1.0, image-min-tokens 1024, froggeric template) I was screwing around with different Qwen3.8-27B quants and thought of this very simplistic but seemingly bechmaxxing resistant combined SVG and vision test. Just let… 4 Hacker News — AI on Front Page community 4d ago Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights Article URL: https://www.bloomberg.com/news/articles/2026-08-26/china-s-z-ai-made-ox-alpha-stealth-model-that-rivals-deepseek Comments URL: https://news.ycombinator.com/item?id=49446422 Points: 203 # Comments: 89 35 r/LocalLLaMA community 4d ago A 27b model beating latest frontier models was not on my 2026 bingo card https://preview.redd.it/kbsqh6f7molh1.png?width=730&format=png&auto=webp&s=068dbea9a50be634a369d54d8b27b781d020fab3 My experience with Qwen 3.8 for agentic tasks has been phenomenal but I personally feel that 3.7 flash is more reliable for overall tasks.   submitted by  … 26 The Information — AI news-outlet 4d ago DeepSeek’s Revenue Reaches $70 Million as of July, Tenfold Jump from 2025 DeepSeek generated about 475 million yuan ($70.7 million) in revenue in the first seven months of this year, roughly tenfold its full-year 2025 revenue, as the Chinese AI lab goes full steam ahead in its second round of funding, The Information reported on Wednesday. DeepSeek… 21 r/LocalLLaMA community 4d ago OpenCode with Qwen3.8-27B for Small Games or Browsing the Web With 16GB VRAM In the past, I have use llama.cpp, but I read that the exl3 quantization format should give better precision , so I have tried exllamav3/tabbyAPI. It was able to write the shown simple HTML game without interaction after asking some questions. The following was tested on a… 15 r/LocalLLaMA community 4d ago [Megathread] Qwen3.8-Flash-Next - Release Day Megathread for discussing the (impending) release of Qwen 3.8 Flash Next. Quants Fine-Tunes & Abliterations Chat Templates Inference Server Support & Configuration Experiences, Benchmarks & Model Comparisons We'll try to clean up future duplicates around the release and point… 37 The Information — AI news-outlet 4d ago DeepSeek’s Revenue Reaches $70 Million as of July, Tenfold Jump from 2025 DeepSeek generated about 475 million yuan ($70.7 million) in revenue in the first seven months of this year, roughly tenfold its full-year 2025 revenue, as the Chinese AI lab goes full steam ahead in its second round of funding, according to two people with knowledge of the… 38 Smol AI News news-outlet 4d ago not much happened today **Z.ai** launched **GLM-5.3-Flash**, a natively multimodal model with a **1M-token context window**, **320B total parameters / 18B active parameters**, under the **MIT License**. It is positioned as a price-competitive successor to GLM-5.2 and claims performance on par with… 29 r/LocalLLaMA community 4d ago Underrated Muse Glimmer Benchmarked qwen3.8 xhigh, medium and muse glimmer. Xhigh effort mode with qwen3.8 took almost 30hrs. (And still failed on 16 cases because of the 32K output token limit) Medium effort mode and muse glimmer were 3-4 hours each. But I'm actually surprised by the muse glimmer… 31 The Information — AI news-outlet 4d ago OpenAI Says Its Jalapeño AI Chip Is Better Than Nvidia’s Blackwell OpenAI on Tuesday released an evaluation of its new Jalapeño AI server chip, claiming it is faster and more powerful than Nvidia’s flagship Blackwell chips. The results appear promising for OpenAI, which hopes to reduce its reliance on Nvidia’s hardware someday, but it’s still… 24 r/LocalLLaMA community 4d ago Open Source Kernel in Qwen3.6-35B-A3B for AMD MI350X: 78,498 output tok/s on 8 GPUs So here's the thing, almost everyone use NVIDIA to run their LLMs, we also do the same, a lot of people we've met use like RTX PRO 6000 or even H100, B300 It seems like everyone eyes is looking at NVIDIA. However we do the math that the raw power alone on AMD GPU MI350X is… 9 r/LocalLLaMA community 4d ago Thomson Reuters releases Thomson-1.0-Small. A law and tax focused model   submitted by   /u/RedditUsr2 [link]   [comments] 23 r/LocalLLaMA community 4d ago Getting Qwen3.8-27B with decent speed on my 4080 with 16Gb card I saw that Q2 is actually very good and produce real good results: https://youtu.be/WNMnbba35VI?is=UNokHqdY4bA5kDgw and I also saw how dflash2 make its running at generating >60 t/s with a 120k context lenght. https://youtu.be/RBlRTUwJMI4?is=LCtTHgkaiWnGfLv9 And I like what its… 33 r/LocalLLaMA community 4d ago Fully quantized NVFP4 Qwen3.8-27B with QUASAR QAD We're releasing a fully quantized NVFP4 version of Qwen3.8-27B. The checkpoint was trained using quantization-aware distillation (QAD) with QUASAR, our new QAT algorithm. We used the original BF16 model as the teacher and distilled the quantized model for 2,446 steps. The… 6 Page 3 of 10 · 500 articles ← Newer Older →