News / #model-release Tag Model releases 500 articles archived under #model-release · RSS Sign in to follow r/LocalLLaMA community 2d ago Ornith 1.5 is actually pretty good hey guys i recently started using ornith 1.5 to rapidly test some tools im working on since qwen 3.8 27b was too slow for my testing loop. This model is actually really good. im getting around 130 tokens / second with mtp and its very good at tool calling. i feel like this is… 32 arXiv — Machine Learning research 2d ago When Privacy Hurts Mergeability: Geometry-Aware Model Merging under Differential Privacy arXiv:2608.26655v1 Announce Type: new Abstract: Model merging promises to construct a single multi-task model from independently fine-tuned task models without accessing the original task data. This makes it attractive when task data cannot be centralized, but released task… 18 arXiv — NLP / Computation & Language research 2d ago TelecomGPT-R1: A Unified Open-Source Reasoner for the Telecom Stack arXiv:2608.26126v1 Announce Type: new Abstract: Telecommunications is a high-leverage domain for large language model (LLM)-based reasoning because routine engineering workflows require joint grounding in normative specifications, operational telemetry, vendor-specific fault… 5 arXiv — NLP / Computation & Language research 2d ago Agents Don't Paginate: First-Chunk Selection for LLM Tool Responses arXiv:2608.26130v1 Announce Type: new Abstract: Coding agents built on large language models (LLMs), such as Claude Code, Cursor, OpenAI Codex, GitHub Copilot, and Aider, receive tool responses that routinely exceed the agent's per-turn token budget. The standard remedy,… 32 arXiv — NLP / Computation & Language research 2d ago Vowel Signs Are Not Letters: A Pre-tokenization Ceiling on Multilingual Tokenizer Fertility arXiv:2608.26449v1 Announce Type: new Abstract: Byte-level BPE tokenizers that use the HuggingFace ByteLevel pre-tokenizer inherit GPT-2's word regex, where a word is defined as \p{L}+, one or more Unicode letters. In abugida scripts, vowels are written as combining marks; this… 35 r/LocalLLaMA community 2d ago Ornith-1.5-35B-A3B on 8 GB VRAM: I think I've found my sweet spot A few days ago I posted asking what people considered the best local model for an 8 GB VRAM GPU . At the time, my personal sweet spot was Qwen3.6-35B-A3B , for agentic coding with Pi.dev. Well… Thanks to the suggestions in that thread, I think I've found something even better.… 25 r/LocalLLaMA community 2d ago I am Concerned if Nvidia Acquires Llama.CPP, Dev Team and HF, Anybody else? I dont know about others, but Nvidia is aiming (potentially) to close the lid on older GPUs since they want to push their new technology. Llama and team has been the to go places for older GPUs like V100s. Knowing how Nvidia have tried killing these GPUs of relevancy concerns me… 25 OpenAI official-blog 2d ago Supporting Thailand’s next generation of AI startups OpenAI and Thailand’s MHESI launch an eight-week accelerator helping 10 health, wellness, and education startups turn AI prototypes into trusted products. 26 r/LocalLLaMA community 2d ago Do you think a few Qwen3.8-27B models working together could score as well as Fable-5 on LiveCodeBench Hard? Has anyone tested this? Ensemble of small Qwen models claiming Fable 5-level coding performance. A new paper claims that running several Qwen3.8-27B models together matches Fable 5’s accuracy on LiveCodeBench. The authors also say their setup paired with GPT Terra reaches Fable… 11 r/LocalLLaMA community 2d ago yall are sleeping on qwen 3.8 27b q2 + q2 dflash + q5 kv ok bit more context: it's actually a QAT Q2 for Qwen 3.8 27 B: https://huggingface.co/sdkyuan/qwen3.8-27B-qat-q2_0-gguf QAT Q2 for DFlash model: https://huggingface.co/HermiHg/Qwen3.8-27B-DFlash2-Q2_K_S-MIX-GGUF Q5 KV seems to cause 0 problems for me; I've used it up to 200K… 36 ThursdAI news-outlet 2d ago NVIDIA Buys Hugging Face! GLM-5.3-Flash, Qwen4 Preview, Gemini Omni 1.1, and the Datacenter Debate w/ Andy Masley From CoreWeave - join Alex and ThursdAI co-host, covering the last week of the summer in AI, with 4 Flash models, Datacenter debate & more AI news 11 Vercel — AI dev-tools 2d ago Hy4 Preview now available on AI Gateway Hy4 Preview from Tencent is now available on AI Gateway. Hy4 Preview is an open-source Mixture-of-Experts model with 770B total parameters aimed at long-horizon coding, document analysis, game development, and scientific reasoning. It serves a context window of 1M tokens. To use… 32 Simon Willison community 2d ago Breaking Claude Code Opus 5 Auto Mode Breaking Claude Code Opus 5 Auto Mode Anthropic are putting a great deal of faith in Claude Code's auto mode for protecting their coding agent users against prompt injection attacks. They recently made that the default and have made bold claims about its effectiveness. Johann… 10 r/LocalLLaMA community 2d ago Appreciation Post - thomsonreuters/Thomson-1.0-Small With the lack of support from Qwen regarding the smaller 9B and 35B MOE models. Like myself, not everyone is looking for an agentic coding model, I particularly use it for RAG and reviewing and require high reasoning across different documents & came across this Finetune:… 22 LangChain releases dev-tools 2d ago langchain==1.4.0a1 Initial release fix(langchain): name the content type MCP conversion could not handle release(langchain): 1.4.0a1 test(langchain): skip MCP tests on a pydantic older than mcp supports test(langchain): drive MCP tests through FastMCP's own utilities fix(langchain/mcp): review… 25 LangChain releases dev-tools 2d ago langchain-fireworks==1.6.1 Changes since langchain-fireworks==1.6.0 release(fireworks): 1.6.1 ( #39975 ) fix(fireworks): drop reasoning history blocks ( #39973 ) chore(model-profiles): refresh model profile data ( #39844 ) 4 r/LocalLLaMA community 2d ago Agent Arena Code - Very good result (preliminary) for GLM and Qwen! Everything has changed in two months: DS4 0731 flash was the start of a wave that is taking open weights to paradise. It is easy to think that Qwen 4 and GLM 6 will be on par with Mythos.   submitted by   /u/LegacyRemaster [link]   [comments] 22 r/MachineLearning community 2d ago py-evoFE: Automated Evolutionary Feature Engineering for Tabular ML in Python (Genetic Algorithms + Scikit-Learn + Polars) [P] Hey everyone! I’m excited to announce the release of py-evoFE (v0.3.0) — an open-source Python library that uses genetic algorithms to automatically discover, combine, and optimize feature transformations for tabular datasets. GitHub: https://github.com/tanopereira/py-evoFE… 34 Ars Technica — AI news-outlet 2d ago Elon Musk’s xAI used child porn to train Grok models, lawsuit says xAI accused of training Grok on real and AI-generated child pornography. 30 llama.cpp releases dev-tools 2d ago b10661 ci : build only the ggml-hip backend for windows-rocm release ( #27753 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/43506547 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel… 15 Ollama releases dev-tools 2d ago v0.33.2-rc1 app: list account cloud models for Claude ( #18077 ) 16 llama.cpp releases dev-tools 2d ago b10660 model: add Qwen3.8-Flash-Next (qwen4exp) ( #27742 ) gguf: add qwen4exp (Qwen3.8-Flash-Next) arch and converter Adds the GGUF-side plumbing for HF model_type qwen4_exp: MODEL_ARCH.QWEN4EXP plus tensors for the low-rank hyper-connection variant (hc_ norm/down/up/inject) and the… 29 r/LocalLLaMA community 2d ago Over 200k context on 16GB VRAM with Qwen 3.8 27B UD-IQ3_XXS I was using UD-Q3_K_XL until now with more than 140000 context. Quality wise it's very good, very few erroneous tool calls. Then I saw many others here reporting good results with IQ3_XXS, so I gave it a try. The downside is prompt processing speed went down from 700-800 tk/s to… 7 r/LocalLLaMA community 2d ago llama.cpp support for Qwen3.8-Flash-Next has been merged finally I can download the GGUF   submitted by   /u/jacek2023 [link]   [comments] 23 LangChain releases dev-tools 2d ago langchain-core==1.6.1 Changes since langchain-core==1.6.0 revert: release(core): 1.6.2 ( #39971 ) release(core): 1.6.2 ( #39967 ) fix(core): shore up indexing in genai v1 streaming content ( #39964 ) fix(core): make StructuredTool JSON-serializable ( #39631 ) chore(deps): bump minor and patch… 21 r/LocalLLaMA community 2d ago Bill Gates is looking to meet with Chinese President Xi Jinping later this year, eager to propose global efforts to mitigate the growing risks posed by artificial intelligence. He believes China might agree to restricting potentially dangerous AI model releases if US took…   submitted by   /u/f0urxio [link]   [comments] 38 llama.cpp releases dev-tools 2d ago b10659 ci : bundle HIP runtime DLLs with Windows ROCm release ( #26973 ) Copy amdhip64_7, amd_comgr and rocm_kpack next to the binaries so the correct HIP runtime loads over the driver's copy in System32. Fixes #26929 . Website: https://llama.app Attestations:… 34 llama.cpp releases dev-tools 2d ago b10658 spec : add DFlash2 support (local convolution + candidate selector) ( #27342 ) ( #27816 ) spec : add DFlash2 support (local convolution + candidate selector) ( #27342 ) support DFlash2 Add p_min in DFlash2 Assisted-by: Claude Opus 5 Revert unnecessary changes Assisted-by: Claude… 35 r/LocalLLaMA community 2d ago No, Engrams won't let you run 1T models locally. It does something even better. Ever since Qwen 3.8 Flash Next dropped, there's a misconception going around that N-gram tables will let people run 1T+ parameter models on a single server with 980B parameters offloaded to SSD. I'm here to disappoint you: it won't. But what it will actually do for local models… 31 r/LocalLLaMA community 2d ago [audio.cpp] Release 0.7: 62 audio model families (85+ variants), Arena UI for model comparison, MiniMax Music 3, FireRed TTS3/Audio, ControlFoley, Personaplex, and more audio.cpp 0.7 is out :) This release adds a lot of new audio models and a new way to compare them locally. Audio.cpp is now at 62 model families and 85+ model variants. And it keeps growing! The biggest user-facing change is the new Arena UI . Instead of testing one model at a… 22 LangChain releases dev-tools 2d ago langchain==1.3.18 Changes since langchain==1.3.17 release(langchain): 1.3.18 ( #39966 ) fix(langchain): preserve content-block shape in PIIMiddleware redaction ( #39894 ) fix(core): shore up indexing in genai v1 streaming content ( #39964 ) 9 Hacker News — AI on Front Page community 2d ago Gemini Omni 1.1 Flash Article URL: https://blog.google/innovation-and-ai/technology/developers-tools/build-with-gemini-omni-1-1-flash/ Comments URL: https://news.ycombinator.com/item?id=49467922 Points: 213 # Comments: 150 31 Google DeepMind official-blog 2d ago Gemini Omni 1.1 Flash lets you build with more control Gemini Omni 1.1 Flash lets you build with more control Aug 27, 2026 | x.com Facebook LinkedIn Mail Omni now delivers studio-quality video production, including the ability to extend a scene, first and last frame interpolation, crisp 4K upscaling, faster prototyping, and more.… 26 r/LocalLLaMA community 2d ago I used local Qwen 27b to build a harness and replace OpenCode Sharing my harness for running local LLMs that I built using Qwen 3.x 27B (> 90% locally built) under my supervision - not vibe-coded. Its free, no telemetry, and open-source (AGPL). Works on Windows, Linux (sorry, no Mac yet). I use it for my own coding + mixed workflows. How… 13 r/LocalLLaMA community 2d ago We’re the Team Behind Apodex 1.1 — Ask Us Anything! Hi r/LocalLLaMA ! We’re Apodex , the team behind Apodex 1.1 , our new model family built to scale agentic intelligence for complex work. We’re excited to be here and answer your questions directly. Apodex 1.1 is designed around sustained, verifiable progress toward real-world… 18 LangChain releases dev-tools 2d ago langchain-anthropic==1.7.0 Changes since langchain-anthropic==1.6.1 release(anthropic): 1.7.0 ( #39963 ) feat(anthropic): support top-level param for skills via container ; updates thinking display mode ( #39962 ) feat(anthropic): support 1.0 sdk ( #39938 ) fix(anthropic): auto-append… 28 r/LocalLLaMA community 2d ago Request: unsloth Please re-quantize Qwen3.6 35 A3B and 27B using UD 3.0 UD 3.0 seems to be a massive improvement over UD 2.0 Some of us still want to run the older Qwen models but would benefit from UD 3.0 UD 2.0 vs 3.0 is like the difference between a full quant. So Q3 UD 3.0 is similar to Q4 UD 2.0.   submitted by   /u/Fancy-Snow7 [link]… 14 r/LocalLLaMA community 2d ago GLM-5.3 Flash Unsloth GGUF now available   submitted by   /u/ElementNumber6 [link]   [comments] 30 r/LocalLLaMA community 2d ago llama: model_loader: add TENSOR_READ_LAZY by ngxson · Pull Request #27794 · ggml-org/llama.cpp Qwen 3.8 Next Flash (Qwen 4) engrams don't need to be in VRAM/RAM   submitted by   /u/jacek2023 [link]   [comments] 8 Vercel — AI dev-tools 2d ago Cursor is now available in the AI SDK harness layer The AI SDK harness layer now supports Cursor through the official @ai-sdk/harness-cursor adapter. The harness layer lets your application run different coding agents through the same HarnessAgent interface, so you can switch agents without changing your application code. Pass… 18 r/LocalLLaMA community 2d ago GLM-5.3-Flash @ DGX Station GB300: ~206 tok/s (single stream), 1M context Hey all! I'm finally doing some cool stuff with my "thinking heater" (h/t u/-TV-Stand- ). I'm still experimenting with GLM-5.2 (in anticipation of 5.3 coming tomorrow, I hope!) and things are very cool so far. With the release of GLM-5.3-flash, I decided to play with it on the… 25 llama.cpp releases dev-tools 2d ago b10655 Feature: Added LIGHTNING_INDEXER support for Deepseek V4 ops on Vulkan Backend ( #27453 ) vulkan: add LIGHTNING_INDEXER op vulkan: updated lightning_indexer.comp and ggml-vulkan.cpp with 128-lane dot-product reduction moved from a shared-memory tree to subgroupAdd. vulkan:… 14 r/LocalLLaMA community 2d ago Qwen3.8-Flash-Next: Time to Update Those Benchmarks specs hardware: M4 Max 128GB Studio inference engine: oMLX & lllama.cpp insights it still very early, so had to disable oMLX K/V caching, qwen4_exp architectureis not yet supported + the obvious n-grams with which the whole 4 bit quant takes ~100G, so pretty tight nevertheless,… 35 r/LocalLLaMA community 2d ago Moderation !== Censorship All this mega threads and censoring posts killing LocalLLaMA vibe. And yeah, I liked more when we had 20+ posts about new model.   submitted by   /u/inkberk [link]   [comments] 20 llama.cpp releases dev-tools 2d ago b10646 metal : fix memory leaks due to missing autoreleasepools ( #27758 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/43372015 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel… 20 r/LocalLLaMA community 2d ago What are the minimum specs required to run Qwen3.8-Flash-Next? How much system RAM? How much VRAM? How much SSD space? Ideally list for q3/4 but q2 might also work since I have seen 3.8 27B perform well even on q2. Currently I have 5070 Ti with 16GB VRAM and 48GB system RAM. I can upgrade system RAM to 96GB is that will allow it to run.… 19 llama.cpp releases dev-tools 2d ago b10645 llama : add --n-cpu-ffn option ( #26622 ) common : dedupe --n-cpu-moe / --spec-draft-n-cpu-moe override loops common : add --n-cpu-ffn to CPU-offload dense FFN weights of first N layers common : generalize llm_ffn_block_regex over the FFN regex, drop TODO Website:… 34 The Information — AI news-outlet 3d ago Exclusive: Gibson Dunn Hires Paul Weiss Tech Lawyer Ashtor Gibson Dunn, a Los Angeles-headquartered law firm well-known for its litigation practice, has hired tech lawyer Jonathan Ashtor from New York law firm Paul Weiss, according to a release reviewed by The Information. The hire is the latest in a series of senior moves among the… 36 r/LocalLLaMA community 3d ago llama : add --n-cpu-ffn option by John-194 · Pull Request #26622 · ggml-org/llama.cpp tl;dr faster dense models for low VRAM people option similar to the existing --n-cpu-moe It puts user specified amount of FFN sublayers for dense models. PR by u/Stainless-Bacon   submitted by   /u/jacek2023 [link]   [comments] 7 r/LocalLLaMA community 3d ago Qwen3.8-Flash-Next better then DeepSeek V4 Pro   submitted by   /u/Normal-Phone7762 [link]   [comments] 30 Page 2 of 10 · 500 articles ← Newer Older →