News / #model-release Tag Model releases 500 articles archived under #model-release · RSS Sign in to follow r/LocalLLaMA community 8d ago AntLing released a dspark draft model for Ling-3.0-flash No GGUFs yet on Huggingface though.   submitted by   /u/Ihtien [link]   [comments] 8 r/LocalLLaMA community 8d ago Freetokens project is impressive A new project was released yesterday and I have the opportunity to test it today. Papper: https://arxiv.org/abs/2608.16157 Github: https://github.com/FlashML-org/FreeToken My initial tests with the following setup: RTX 5080 (16 GB) DDR6 64GB AMD Ryzen 9 9950X3D I got 100tok/s on… 14 r/LocalLLaMA community 8d ago Llama.cpp version 0.2.0 is out! You can find the changelog and source code here: https://github.com/ggml-org/llama.cpp/releases/tag/v0.2.0 Associated pre-build is here: https://github.com/ggml-org/llama.cpp/releases/tag/b10566   submitted by   /u/PhilippeEiffel [link]   [comments] 32 r/LocalLLaMA community 8d ago Bro wtf, Qwen Lab cooked with Qwen 3.8 27B, it's so fucking good Context: Earlier, open-source large models like Mimo V2.5 Pro, DeepSeek V4 Pro (first version), and Kimi K2.5 used to struggle with this prompt, and Qwen 3.6 27B couldn't even render the globe properly. But now, Qwen 3.8 27B is so much better than the previous model. I hope qwen… 29 r/LocalLLaMA community 8d ago Sharp template to NInfer: -42% output tokens, same speed Sharp is u/peculiar-ragdoll 's system prompt that makes Qwen answer way more tersely without losing correctness. It's built on top of froggeric's fixed chat templates for Qwen; several fixes now in the v22.x templates (error-escalation tiers, false retry-loop kills, multi-system… 20 r/LocalLLaMA community 8d ago 16 GB VRAM purgatory discussion thread What models and configs are we using? Please share here On windows, I am using this copium pared down model https://huggingface.co/Bucoid/Qwen3.8-27B-Uncensored-IQ4-XS-MTP-16GB-VRAM-GGUF with MTP disabled, q4 k/q4 v mmproj banished to CPU/RAM and a small ub to save whatever… 18 Ollama releases dev-tools 8d ago v0.33.0 What's Changed mlx: fix mac assumptions on linux/windows by @dhiltgen in #17898 mlx update by @dhiltgen in #17886 lint fixes by @dhiltgen in #17897 app: add claude desktop app by @ParthSareen in #17899 app: polish onboarding layout and disable zoom by @hoyyeva in #17885 launch:… 20 r/LocalLLaMA community 8d ago Qwen 3.8 27b - PI AGENT vs OPENCODE - another smaple That is the second comparison and the last one. I will not be spamming again ;) Continuation from: https://www.reddit.com/r/LocalLLaMA/comments/1vu0u2v/qwen_38_27b_pi_agent_vs_opencode/ That is one of my many tests I make comparing output quality. What is more interesting using… 9 r/LocalLLaMA community 8d ago I tried to do agenic coding with Qwen 3.8 27B 3bit quant on a macbook air m2 24gb. It took 63 hours, but amazingly, the flight simulator worked. I used LM Studio Bionic with Qwen 3.8 27B Q3_K_S with 57k context. It took a staggering 63 hours to finish coding. After the first prompt "Create a beautiful, relaxing flight simulator in a single HTML page" taking 47.8 hours, it created an html file that showed the title screen… 15 TechCrunch — AI news-outlet 8d ago Anthropic’s Opus 4.6 is a smut-machine Anthropic forbids its Claude models from generating sexually explicit content. But a series of tests conducted by TechCrunch found that it didn't take much to get past the restriction. 19 llama.cpp releases dev-tools 8d ago b10568 model: use ggml_rope_set_offset() ( #27382 ) model: use ggml_rope_set_offset() partially apply to deepseek2 Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/42252109 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64,… 8 Ollama releases dev-tools 8d ago v0.33.0 What's Changed mlx: fix mac assumptions on linux/windows by @dhiltgen in #17898 mlx update by @dhiltgen in #17886 lint fixes by @dhiltgen in #17897 app: add claude desktop app by @ParthSareen in #17899 app: polish onboarding layout and disable zoom by @hoyyeva in #17885 launch:… 31 r/LocalLLaMA community 8d ago I'm really hoping we're in 2026's 2-month-gap between QwQ and Qwen3 right now QwQ was genuine next-gen performance usable on local hardware, but the massive required context (it's reasoning style was akin to "if I say every possible word, I'll notice the right one!" ) kinda made it unusable for agentic coding. It was ~2 months later that Qwen3-32B came… 20 r/LocalLLaMA community 8d ago Qwen3.8-27B different thinking levels Even the low preset is better than Qwen 3.7 plus or Qwen3.6-27B reasoning   submitted by   /u/Tall_Abrocoma_3533 [link]   [comments] 24 r/LocalLLaMA community 8d ago Qwen 3.8 Low and Medium are goated Artificial Analysis just benchmarked them and the scores are crazy good, proving the earlier success wasn't only enabled by overthinking.   submitted by   /u/Eyelbee [link]   [comments] 37 r/LocalLLaMA community 8d ago Strix Halo (8060S / gfx1151), Qwen-3.8-27B @ Q8 and Q6 UD v3, up to 256K ctx, llama.cpp, DFlash2, vision, real workloads quality and steady performances, optimized recipes, ... Hi fellows fully-local halos, after manually following existing guides, I decided to build an LLM API endpoint installation and optimization guide that works even when autonomously followed by my pi agent, so I can install/experiment/reinstall easily and without babysitting. Q8… 38 Hacker News — AI on Front Page community 8d ago A week of using Codex more than Claude Article URL: https://allaboutcoding.ghinda.com/a-week-of-using-codex-more-than-claude/ Comments URL: https://news.ycombinator.com/item?id=49393051 Points: 209 # Comments: 230 31 r/LocalLLaMA community 8d ago Qwen3.8-27B on an RTX 5060 Ti 16GB: IQ4 vs Q8, 64K context, MTP, vision, and agent benchmarks I’ve been testing Qwen3.8-27B as a possible replacement for the Qwen3.5-9B that I have been running on RTX 5060Ti 16G. The goal was not just maximum tokens/sec, but useful context capacity, reliable tool calling, multi-turn behavior, and vision on a single 16GB GPU for true… 34 llama.cpp releases dev-tools 8d ago b10567 ci : run ccache-clear as the last step of release jobs ( #27503 ) ci : run ccache-clear as the last step of release jobs Assisted-by: pi:llama.cpp/Qwen3.8-27B update disabled job too to force rebase Co-authored-by: Sigbjørn Skjæret sigbjorn.skjaeret@huggingface.co Website:… 11 r/LocalLLaMA community 8d ago Qwen3.8-27B Q6 is a beast at agentic coding A quick feedback after a really major test: nearly 20 hours of non-stop goal-oriented work with Qwen3.8-27B Q6, running across an RTX 3090 and an RTX 3060. It maintained a speed of around 60–63 tokens/s throughout the session.   submitted by   /u/Ok_Ninja7526 [link]… 30 llama.cpp releases dev-tools 8d ago v0.2.0 Overview New version has been released. Nightly build: b10566 Web UI: the nightly-tag.txt asset contains the tag of the corresponding nightly release More info: dist : releases and versioning of ggml-org projects Changelog since v0.1.2 bb4caa7 llama.cpp : bump version to 0.2.0 (… 36 Hacker News — AI on Front Page community 8d ago Scientists release biggest 2D map of the universe Direct link to Legacy Survey Sky Viewer: https://viewer.legacysurvey.org Comments URL: https://news.ycombinator.com/item?id=49392200 Points: 216 # Comments: 57 28 Simon Willison community 8d ago llm 0.32.1 Release: llm 0.32.1 Fresh installs of LLM stopped working the other day because the OpenAI Python library dropped its usage of httpx , and it turned out LLM depended on that library but only installed it via a transitive openai dependency. This dot-release fixes that for the… 12 Simon Willison community 8d ago llm-openrouter 0.7 Release: llm-openrouter 0.7 Now that this plugin is compatible with LLM 0.32 it works much better with reasoning LLMs available through OpenRouter. Updated for compatibility with LLM 0.32 . Models now use OpenRouter's implementation of the Responses API . Three new server-side… 24 r/LocalLLaMA community 8d ago Qwen 3.8 vs 3.6 27b low reasoning loops way less now Have seen some people say Qwen 3.8 still overthinks even when reasoning is set to low. Which on my case has been way better compared to 3.6, eveb on a 3 bit quant. I think it's worth mentioning that the default is actually xhigh, so first make sure to specify it if not already.… 36 r/MachineLearning community 8d ago Does telling an LLM to "be concise" actually save you money? We measured it across 9 models. Compressing the output can save you money and keep accuracy, compressing the input prompt does not. [R] LLMs are too verbose and with a black box model the only things you control are what goes in and how you tell it to write back. Yesterday Claude Code shipped a "concise output style" where Claude keeps things short. We already have a paper out about this! We tested both… 7 r/LocalLLaMA community 8d ago What’s the best local AI harness for coding + general use? So what’s actually the best local AI harness rn? I’ve read a TON about this already and somehow ended up more confused than when I started so I figured screw it, let the community decide. Right now I mainly run Qwen 3.6 35B-A3B and Qwen 3.8 27B, with Ornith 1.5 9B sometimes for… 16 r/LocalLLaMA community 8d ago Ultrafast Qwen3-TTS at 34 ms Time-to-First-Audio, Handling 10 Requests Per Second [OSS] Hey locallama! We recently open sourced a Qwen3-TTS 1.7B implementation that achieves 10 requests per second (RPS) and sub-50 ms p95 time-to-first-audio (TTFA) while maintaining real-time playback on 1 x H100. This extends to 20 RPS at sub-100 ms p95 TTFA. By adjusting settings,… 31 Simon Willison community 8d ago Quoting Matt Webb After I released version 1.0, I figured I would have to do the rotations myself. So I sat down with ChatGPT and I didn’t get it to write the code, but I got it to educate me. With a patient, interactive tutor, I was able to finally do what I hadn’t by reading books and asking… 35 Hacker News — AI on Front Page community 8d ago Claudette: Make Claude stop talking like a BuzzFeed article Article URL: https://github.com/adnanakil/nobuzz/blob/main/README.md Comments URL: https://news.ycombinator.com/item?id=49388752 Points: 234 # Comments: 165 6 LangChain releases dev-tools 9d ago langchain-perplexity==1.4.1 Changes since langchain-perplexity==1.4.0 release(perplexity): 1.4.1 ( #39826 ) fix(perplexity): include type="message" on Responses input items ( #39774 ) fix(perplexity): preserve caller extra_body ( #39203 ) chore: bump pillow from 12.2.0 to 12.3.0 in… 21 TechCrunch — AI news-outlet 9d ago Starcloud raises $250 million for orbital data centers as launch options dry up There's about to be a big fight to secure access to space. 31 r/LocalLLaMA community 9d ago DeepSeek Harness v0.1.1 released https://github.com/deepseek-ai/deepseek-harness/releases/tag/dsh-v0.1.1-rc.1 The DeepSeek adapter adds the multimodal visual understanding model DeepSeek-V4-Flash-Vision-Exp. It also supports configuring native image requests. Commands such as /goal and /plan can accept text and… 22 r/LocalLLaMA community 9d ago Qwen 3.8 27b is strong even at Q3_xxs So usually I avoid Q3 quants because I have had bad experiences with it, models were usually too degraded, so the smallest I normally do is Q4, since I only have rtx 4060 ti 16gb. But since there hasn't been a 35b-3ab released yet, I had to try it. I don't use LLMs in agentic… 13 r/LocalLLaMA community 9d ago Buying a V100/older NVIDIA GPU? Run this to check for older memory issues I recently bought a 32GB V100 off eBay and thought I was all set with rudimentary tests showing zero issues. But then I started seeing weird VRAM-related errors in llama.cpp. I sic'ed Claude on it, to find that ECC was disabled and that can hide small RAM issues. I believe the… 19 r/MachineLearning community 9d ago I pre-registered 23 experiments (~$207, solo) testing whether a verification harness can substitute for scale at 4B — including a locked holdout that falsified my own headline result [R] **Question:** how much of grounded-reasoning performance is weights, and how much is architecture? I spent a month testing this on the Qwen3-4B class, solo, with pre-registration discipline: every run's success criterion frozen in a runbook before spend, every failed bar… 19 Hacker News — AI on Front Page community 9d ago DeepSeek-v4-flash-vision-exp Article URL: https://api-docs.deepseek.com/guides/vision/ Comments URL: https://news.ycombinator.com/item?id=49386163 Points: 286 # Comments: 87 18 llama.cpp releases dev-tools 9d ago b10549 TP: enable tensor split for LFM2/LFM2MOE ( #26993 ) Assisted-by: deepseek-v4-flash Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/42096995 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED… 32 r/LocalLLaMA community 9d ago DeepSeek-V4-Flash-Vision-Exp   submitted by   /u/Xhehab_ [link]   [comments] 8 r/LocalLLaMA community 9d ago Fastest NVFP4 quant of Qwen3.8 27B out there Here's a brand new Blackwell-native, prefill-optimized 4-bit quant that runs 50% faster on compatible hardware than a Q4 quant of the same memory footprint. And it runs 4-7% faster than other NVFP4 quants as benchmarked on RTX 5090 32GB. Quant Benchmark Speed NVFP4 pp2048 6250… 27 r/LocalLLaMA community 9d ago Qwen3.8-27B at 262K context on a Strix Halo + RTX 3090 Ti: 9.5 -> 153 tok/s, and it beats a dual-3090 vLLM box on HumanEval Spent a while treating layer placement, KV format and llama.cpp itself as experimental variables. 159 logged experiments. Numbers first, caveats after. Hardware: AMD Ryzen AI MAX+ 395 (Strix Halo, 128 GB unified) + RTX 3090 Ti on an eGPU link. One llama.cpp process, AMD on… 37 r/LocalLLaMA community 9d ago Fastest qwen 3.8 27b for AMD gpu? Hey, just wondering if there are forks or exact gguf versions that give fastest prompt processing and token gen speeds for AMD gpu? Looking to run q8 or q6 Vram 96gb W7900 + w7800 both 48gb With bandwidth mismatch, tensor paralleling amd equivalent not working   submitted by… 21 r/LocalLLaMA community 9d ago SenseNova U1.5-Lite full release: expert training, OPD distillation, one model at inference Benchmarks: Benchmark U1 Preview Full Qwen-Image-Bench 47.14 55.20 (PE) 60.18 (PE) ImgEdit 3.9 4.37 4.59 GEdit-Bench-EN 7.47 8.14 8.26 Instead of just scaling up, they train task-specialized expert models for text rendering and infographics, aesthetic quality, and image editing.… 15 llama.cpp releases dev-tools 9d ago b10537 CI: Use LLVM's OpenMP over MSVC_DEBUG_non_redist on Windows ( #26678 ) CI: Use LLVM's OpenMP over MSFT_DEBUG_non_redist on Windows Currently, we ship the non-redist debug version of microsoft's libomp. This PR changes this to official LLVM's release, also packaging the license… 35 r/LocalLLaMA community 9d ago Is there any interest in a self hosted open source version of manus/perplexity? I have been working on a project for about 3 months because I found Vane to be awful and I wanted something like Manus for myself. I have a pretty solid app that's probably about ready to release, but I'm not sure if there is any desire there. It does research at different… 9 r/LocalLLaMA community 9d ago Make Jensen Huang Sound Like Anyone. New Streaming Voice Conversion Model MeanVC2 Released! Finally see a new voice conversion model. MeanVC2 supports cross-gender and cross-language voice conversion. 3x realtime on CPU with audio.cpp. Disclaimer: The converted voice quality of MeanVC2 is decent; the noise comes from my rough demo engineering, not the model itself.… 38 ThursdAI news-outlet 9d ago Chill week with Qwen 27B and GLM 5.3 beating GPTs, OpenAI announces pausing RL to focus on security and a cancer vaccine being produced From CoreWeave: a chill week, 2 AI model releases and 3 interviews! 36 r/LocalLLaMA community 9d ago I did it! I'm free! It's been 7 hours since I used claudecode My Pro subscription expired today, they killed my access at 1pm local time. I'm now using Qwen3.8-27b w/ 5090m 24gb vram and pi to do everything i was doing in claudecode. The only downside is claudecode let me code without using my gpu, meaning I have to plan things now. Last… 15 r/LocalLLaMA community 9d ago Qwen 3.8 27b - PI AGENT vs OPENCODE https://www.reddit.com/r/LocalLLaMA/comments/1j7r47l/i_just_made_an_animation_of_a_ball_bouncing/ This post inspired me to make that test after a year ;) That is one of my many tests I make comparing output quality. What is more interesting using a PI Agent results are much… 35 Vercel — AI dev-tools 9d ago DeepSeek V4 Flash Vision Experimental now available on AI Gateway DeepSeek V4 Flash with vision is now available on AI Gateway. This model is an experimental version that accepts images alongside text. You can ask it to describe a picture, read text out of a screenshot, or work through a chart in the same request as your prompt. DeepSeek V4… 13 Page 6 of 10 · 500 articles ← Newer Older →