News / #version-bump Tag Version Bump 450 articles archived under #version-bump · RSS Sign in to follow Ollama releases dev-tools 1mo ago v0.32.2-rc1: server: detect download stalls before the first byte (#17259) server: detect download stalls before the first byte server: keep stall timeout out of download API 12 Ollama releases dev-tools 1mo ago v0.32.2 What's Changed launch: keep Claude Code channels available by @hoyyeva in #17210 cmd: remove dead agent prompt wrappers by @ParthSareen in #17227 agent: reorder working directory instruction by @ParthSareen in #17228 agent: clean up semantics, UX, DX, and procedural code by… 31 r/LocalLLaMA community 1mo ago BeeLlama.cpp v0.4.0: KVarN, KV precision tail, q2_0-q3_1 KV cache, upstream rebase TL;DR llama.cpp fork with more KV cache quantization features, with all claims supported by benchmarks: KVarN, KV cache precision tail, additional types of standard KV cache (q2_0-q3_1, q6_0, q6_1), and more. BeeLLama v0.4.0 is a complicated update, removing and adding features… 14 Simon Willison community 1mo ago Claude Code uses Bun written in Rust now In Rewriting Bun in Rust Jarred Sumner made the following claim: Claude Code v2.1.181 (released June 17th) and later use the Rust port of Bun. Startup got 10% faster on Linux but otherwise, barely anyone noticed. Boring is good. I decided to have a poke at my own Claude Code… 10 r/LocalLLaMA community 1mo ago Arandu v0.6.5 available Repository: https://github.com/fredconex/Arandu Hello Guys, This is Arandu a open source app to easily launch models using llama.cpp, it tries to integrate all in one place with easy multi version management of llama.cpp, HuggingFace models search/download, intuitive arguments… 14 ComfyUI releases dev-tools 1mo ago v0.28.2 ComfyUI v0.28.2 8 r/LocalLLaMA community 1mo ago PSA: <30 Score Gap in Arena.AI is Unnoticeable People generally sees "better than even" as ~59% (ELO gap would be around 70). If you reduce the odds by half, it is ~54% which translates to a 30 ELO gap. The reason why Kimi-k3 made a splash is because the gap is "close enough". P.S. Mimo-v2.5-pro are SOTA for soft science and… 34 ComfyUI releases dev-tools 1mo ago v0.28.1 ComfyUI v0.28.1 38 Hugging Face Daily Papers research 1mo ago WanSong v1.0 Technical Report Abstract Music generation foundation models have recently attracted significant industry attention. However, achieving efficient generation and high-fidelity long-form audio while supporting controllability remains challenging. To address these needs, we present WanSong, a… 19 OpenAI Python SDK releases dev-tools 1mo ago v2.46.0 2.46.0 (2026-07-17) Full Changelog: v2.45.0...v2.46.0 Features api: /organization/projects/{project_id}/service_accounts/{service_account_id}/api_keys" endpoint ( 5a00941 ) api: add owner_project_access to APIKeyListParams ( f589d04 ) api: manual updates ( 980f176 ) api: manual… 10 Hugging Face Daily Papers research 1mo ago Video = World + Event Stream Abstract We present Wan-Streamer v0.3, which reframes our native-streaming interaction model under a single organizing view: a video is a world plus an event stream. The world is the persistent context in which a video unfolds, including the environment, scene, subjects, ambient… 19 Anthropic SDK (Python) releases dev-tools 1mo ago v0.117.0 0.117.0 (2026-07-16) Full Changelog: v0.116.0...v0.117.0 Features api: add support for dreaming ( 642eee7 ) api: add support for MCP Tunnels ( d716df6 ) Bug Fixes credentials: keep credential material out of traceback frame locals via SecretStr ( aa93a4d ) Chores docs: small… 10 llama.cpp releases dev-tools 1mo ago b10038 ci : add official website link to release notes ( #25728 ) Assisted-by: pi:llama.cpp/Qwen3.6-27B Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU)… 4 Ollama releases dev-tools 1mo ago v0.32.1 What's Changed Improved Gemma 4 tool calling and multi-turn reasoning, including more reliable tool-response continuations Fixed a recurrent MLX model cache leak that could increase memory use across requests, and improved cache snapshot performance MLX text model loading now… 15 Ollama releases dev-tools 1mo ago v0.32.1-rc0 cmd: put current working dir in the system prompt ( #17188 ) 5 r/LocalLLaMA community 1mo ago ExLlamaV3 v1.0.0 - Major Performance Upgrades After over a year in development, ExLlamaV3 has had its first production release . Turboderp has been pulling 10 hour days with Fable to bring us this massive batch of improvements. Check out detailed performance metrics and a little write-up from him here . Some of the biggest… 5 ComfyUI releases dev-tools 1mo ago v0.28.0 ComfyUI v0.28.0 34 r/LocalLLaMA community 1mo ago KAT-Coder-Air V2.5 - Open model soon Tweet : https://xcancel.com/KwaiAICoder/status/2075482952696578544#m KAT-Coder-Air V2.5 is available on Openrouter . Somebody please check & let us know about this model. KAT-Coder-V2.5 Technical Report https://arxiv.org/abs/2607.05471 https://arxiv.org/pdf/2607.05471  … 14 Ollama releases dev-tools 1mo ago v0.32.0 launch: rename Codex App integration to ChatGPT ( #17161 ) 25 vLLM releases dev-tools 1mo ago v0.25.1: [Bugfix] Guard mixed-dtype allreduce RMSNorm quant fusions (#48330) Signed-off-by: hcenteno hugo.centeno@estudiantat.upc.edu (cherry picked from commit 5f8e73c ) 6 llama.cpp releases dev-tools 1mo ago b9970 ggml : add GGML_OP_LIGHTNING_INDEXER that implements DeepSeek V3.2/V4 lightning indexer ( #24231 ) ggml : add GGML_OP_LIGHTNING_INDEXER that implements DeepSeek V3.2/V4 lightning indexer ggml : remove scale parameters from lightning indexer OP, add f16 mask parameter tests : add… 14 r/LocalLLaMA community 1mo ago Vellium v1.0.0 released: security hardening, wallpaper-based themes, JSON chat export and a major desktop stability pass Vellium has reached v1.0.0. It is a local-first desktop workspace for writing, roleplay, character creation, lorebooks and knowledge management with local LLMs. This release promotes the previous v1.0.0-beta build to the first stable version. The main focus was security and… 13 r/LocalLLaMA community 1mo ago Xiaomi quietly uploaded MiMo-V2.5-DFlash — official DFlash weights are now on Hugging Face https://huggingface.co/XiaomiMiMo/MiMo-V2.5-DFlash Xiaomi appears to have quietly uploaded MiMo-V2.5-DFlash to Hugging Face: there is dedicated dflash directory containing the Dflash model, anyone willing to GGUF it and try? I'd do it but I can't today. This model is pretty good… 19 r/LocalLLaMA community 1mo ago **Your $80 Tesla P100 has been doing silently noisy math in llama.cpp for years. Three lines fix it, for free.** ## TLDR; Shipped — in turboquant v0.3.0, downloadable now. https://github.com/TheTom/llama-cpp-turboquant/releases/tag/tqp-v0.3.0 llama.cpp's CUDA code has a flag that means "this GPU is fast at fp16, so do the math in fp16." The GTX 10-series and P40's (sm_61) were exempted… 7 r/LocalLLaMA community 1mo ago Grok Build CLI uploads your whole repo — full git history + .env secrets — to xAI's cloud, and the opt-out doesn't stop it (wire-captured) I ran Grok Build CLI (v0.2.93) through mitmproxy. It uploads your entire repo as a git bundle (full history) to xAI's Google Cloud — independent of what you open. With the prompt literally "do not read or open any files," a file I planted came back verbatim when I git clone -d… 5 Ollama releases dev-tools 1mo ago v0.32.0-rc0 cmd: agent UI ( #17017 ) 20 r/LocalLLaMA community 1mo ago Koboldcpp v1.117 released   submitted by   /u/Fcking_Chuck [link]   [comments] 14 r/LocalLLaMA community 1mo ago MiMo v2.5 is underrated. Feels like the tokens are pouring out of the screen in OpenCode. When I recently built my inference server, I expected to deploy DeepSeek v4 flash, but that doesn't look like it's going to be fast for a long time, if ever. There is a massive gap, as we all know, in competent models between 30b and 400b. I was very surprised to find that this… 22 vLLM releases dev-tools 1mo ago v0.25.0: [CI] Fix cargo-deny config flag ordering (#48170) Signed-off-by: Lucas Wilkinson lwilkins@redhat.com 35 OpenAI Python SDK releases dev-tools 1mo ago v2.45.0 2.45.0 (2026-07-09) Full Changelog: v2.44.0...v2.45.0 Features api: gpt-5.6-sol updates ( 039d1fe ) Bug Fixes api: restore beta resource accessors ( 2dfc130 ) Chores retrigger release automation ( 7b61351 ) 21 vLLM releases dev-tools 1mo ago v0.25.0rc3 [P/D][Bugfix] Fix PD async KV load lookahead handling for MTP spec de… 6 vLLM releases dev-tools 1mo ago v0.25.0rc2 Fix embed scaling + CUDA graphs in Transformers modelling backend ( #4 … 12 r/LocalLLaMA community 1mo ago 82 TPS On Qwen 3.6 27b On A Macbook Pro | Introducing MTPLX V2: The Fastest Way To Run MLX Models. Hey Everyone, here is an update on MTPLX! One month after releasing MTPLX V1 which brought a swift based app and upgraded CLI for coding use I am happy to announce MTPLX V2. The biggest change is Turbo Mode: using custom verify-specialized quantized-matmul kernels plus a… 12 ComfyUI releases dev-tools 1mo ago v0.27.1 ComfyUI v0.27.1 36 r/LocalLLaMA community 1mo ago any one else finds Mimo v2.5 better than deepseek v4 flash!? I noticed while using both, mimo was often better, after benchmarking mimo v2.5 via open code endpoint in diff harness like codex, oh my pi, hermes. i found that mimo is indeed better in coding tasks. and over all, hermes scored 55% with mimo v2.5 via terminal bench v2.0 others… 25 arXiv — NLP / Computation & Language research 1mo ago Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability arXiv:2607.06196v1 Announce Type: new Abstract: Current AI safety evaluation and benchmarking frameworks predominantly rely on Western-centric culture-agnostic defaults that mask critical regional laws, socio-linguistic nuances, and cultural taboos, leaving Vision-Language… 7 vLLM releases dev-tools 1mo ago v0.25.0rc1 [CPU][Bugfix] Fix flaky ShortConv prefill test on ARM (uninitialized … 5 Ollama releases dev-tools 1mo ago v0.31.2-rc2: llm: allow iGPU mmproj offload with fit padding (#16996) llm: allow iGPU mmproj offload with fit padding llama.cpp's fit pass sizes text-model placement before the multimodal projector is loaded. Ollama had been avoiding that risk on non-Metal iGPUs by disabling projector offload entirely, which forces CLIP onto CPU on GB10 and Strix… 33 Ollama releases dev-tools 1mo ago v0.31.2: llm: allow iGPU mmproj offload with fit padding (#16996) llm: allow iGPU mmproj offload with fit padding llama.cpp's fit pass sizes text-model placement before the multimodal projector is loaded. Ollama had been avoiding that risk on non-Metal iGPUs by disabling projector offload entirely, which forces CLIP onto CPU on GB10 and Strix… 25 r/LocalLLaMA community 1mo ago mistral.rs v0.9.0: up to 1.8x faster CPU decode than llama.cpp on x86 and ARM! https://preview.redd.it/nuk5rxceptbh1.png?width=1448&format=png&auto=webp&s=300344dd4c6552379e8536b81ba288be3d6dca3f On Qwen3 4B Q4_K, mistral.rs decodes faster than llama.cpp at every context depth we measured, on x86 (Sapphire Rapids) and ARM (GB10). We optimized mistral.rs at… 27 r/LocalLLaMA community 1mo ago Mimo & deepseek are really amazing at optimizing ai. Read the the official blog page i linked, it will give amazing insight on how they pulled off this kind of low pricing with 2x - 3x profit margins. For quick look --> https://x.com/i/status/2059618247553745204 Detailed --> https://mimo.xiaomi.com/blog/mimo-v2-5-inference I hope in future we get fable lvl ai at the cost of current DSV4. Thats far more sufficient for like 90% of people. Xai is also pushing for low cost api… 17 Hugging Face Daily Papers research 1mo ago Wan-Streamer v0.2: Higher Resolution, Same Latency Abstract Wan-Streamer v0.2 enhances audio-visual interaction by increasing visual resolution while maintaining low latency through optimized thinker-performer architecture with multi-GPU parallel processing. Generated by Qwen/Qwen2.5-Coder-32B-Instruct We present Wan-Streamer… 22 Hugging Face official-blog 1mo ago LeRobot v0.6.0: Imagine, Evaluate, Improve Back to Articles a]:hidden"> LeRobot v0.6.0: Imagine, Evaluate, Improve Published July 7, 2026 Update on GitHub Upvote 1 Steven Palma imstevenpmwork Pepijn Kooijmans pepijn223 Caroline Pascal CarolinePascal Khalil Meftah lilkm Martino Russi nepyope Nikodem Bartnik nikodembartnik… 26 Ollama releases dev-tools 1mo ago v0.31.2-rc1: create: harden GGUF create flows (#17062) create: harden GGUF create flows lint 26 Ollama releases dev-tools 1mo ago v0.31.2-rc0 mlx: update to de7b4ed9 ( #17056 ) 16 Simon Willison community 1mo ago sqlite-utils 4.0rc3 Release: sqlite-utils 4.0rc3 I hoped to release sqlite-utils 4.0 stable this weekend, but as I worked through the backlog of issues and PRs with a combination of Claude Fable 5 and GPT-5.5 the changelog since rc2 kept getting bigger . The biggest new feature is support for… 32 Hacker News — AI on Front Page community 1mo ago Shadcn/UI now defaults to Base UI instead of Radix Article URL: https://ui.shadcn.com/docs/changelog Comments URL: https://news.ycombinator.com/item?id=48791328 Points: 212 # Comments: 94 23 r/LocalLLaMA community 1mo ago Is there some KDL chart for MiMo-V2.5 or something regarding the quants quality? I'm using the model with opencode and the issue is it's looping hard when reasoning. It's not a deranged babbling though, the reasoning is legit large spans of text, it just can't get outside of the loop and make a decision. So I babysit it, stop and direct it to the right path,… 33 Anthropic SDK (Python) releases dev-tools 1mo ago v0.116.0 0.116.0 (2026-07-02) Full Changelog: v0.115.1...v0.116.0 Features api: add agent-memory-2026-07-22 beta header ( e181d5c ) 11 Hacker News — AI on Front Page community 1mo ago Kimi K2.7 Code is generally available in GitHub Copilot Article URL: https://github.blog/changelog/2026-07-01-kimi-k2-7-is-now-available-in-github-copilot/ Comments URL: https://news.ycombinator.com/item?id=48756602 Points: 202 # Comments: 89 24 Page 4 of 9 · 450 articles ← Newer Older →