News / #model-release Tag Model releases 500 articles archived under #model-release · RSS Sign in to follow r/LocalLLaMA community 4d ago Qwen3.8-27B IQ3_XXS wrote a correct multilayer TMM on a 16 GB Quadro — after 100 minutes, 3 compactions, and 108k output tokens https://preview.redd.it/i8rjx0ar5mlh1.png?width=2160&format=png&auto=webp&s=c2588bc7b2519ea71b176ca73faf566dfc585496 I wanted to see whether a heavily quantized 27B model running entirely on an older 16 GB workstation GPU could do more than the usual coding demos. FFT felt too… 8 Vercel — AI dev-tools 4d ago GLM 5.3 Flash now available on AI Gateway GLM 5.3 Flash from Z.ai is now available on AI Gateway. The model is a faster, cheaper sibling of GLM 5.3 built for coding and agent tasks that run across many steps. GLM-5.3 Flash is a multimodal model that supports text and vision input, with a 1M token context window and a… 4 Vercel — AI dev-tools 4d ago Qwen 3.8 Flash now available on AI Gateway Qwen3.8-Flash from Alibaba is now available on AI Gateway. It takes text and images as input, serves a context window of 1 million tokens, and can return up to 65k tokens in a response. Alibaba points it at coding, tool use, and multi-step agent work. To use Qwen3.8-Flash, set… 24 Vercel — AI dev-tools 4d ago Muse Image now available on AI Gateway Muse Image from Meta Superintelligence Labs is now available on AI Gateway. It is their first image model and a separate family from Muse Spark, returning images rather than text. Send a prompt and get an image back, or send an image with an instruction and get it changed. One… 30 Vercel — AI dev-tools 4d ago Gemini 3.5 Transcribe now available on AI Gateway Gemini 3.5 Transcribe from Google is now available on AI Gateway. It takes audio and returns text, in two variants: google/gemini-3.5-transcribe transcribes a complete recording in a single request. google/gemini-3.5-transcribe-live transcribes audio over a WebSocket, returning… 18 r/LocalLLaMA community 4d ago M5 Ultra 96GB vs M5 Max 128GB — is 2x bandwidth worth losing 32GB of RAM, with Qwen3.8-Flash-Next dropping tomorrow? I’ve been going back and forth on this for a week and I can’t settle it, so I’m hoping someone here has hands-on numbers. The two configs (German prices, dealer quote, incl. VAT): Config Price Mac Studio M5 Max, 128GB / 512GB SSD €5,859 Mac Studio M5 Max, 128GB / 1TB SSD €6,189… 32 Simon Willison community 4d ago EVE Online: The Move to Python 3 Begins! EVE Online: The Move to Python 3 Begins! EVE Online has been one of the most interesting case studies in Python at scale for over twenty years now. They've been running on Stackless Python since their launch in 2003, and their last major upgrade was 16 years ago, to Stackless… 22 Ollama releases dev-tools 4d ago v0.33.0 What's Changed Claude Desktop Turn individual Ollama models on or off for use in Claude, directly from the menu bar Choose from your available Ollama models from within Claude; cloud models appear only when you're signed in A new Apps view manages app integrations with copyable… 6 r/LocalLLaMA community 4d ago Qwen 3.8 27b has ThreeJs locked down. Generated on a 3090 Qwen 3.8 27b Q4 Thinking high.   submitted by   /u/Both_Opportunity5327 [link]   [comments] 19 r/LocalLLaMA community 4d ago Peak Portable Personal Datacenter Portable rig for Qwen3.8-27B-BF16 200K+ token prompts. My work Panasonic Toughbook + the T1 + power brick + headphones all fit in my lunchbox. Need the BF16 for huge context highly sensitive document OCR, image analysis, aggregation and summarization. I've done a ton of testing… 35 r/LocalLLaMA community 4d ago 35B-A3B tool calling benchmark: Original Qwen vs. KAT Coder, Ornith and Tiel-Coder With hopes of a Qwen3.8-35B-A3B release now mostly dashed, many people including myself are looking at fine-tunes and other variants of Qwen3.6-35B-A3B to run on VRAM-limited hardware. I decided to try to benchmark some of the top contenders: KAT-Coder, Ornith 1.5 and the very… 35 TechCrunch — AI news-outlet 4d ago Claude Cowork finally remembers what you told the app in chat Anthropic is giving Claude a shared memory across chat and Cowork, so users no longer have to repeatedly brief the AI on projects, preferences, and other context. 17 r/LocalLLaMA community 4d ago Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. 👀 Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate: Ideal 4-bit quant ≈ 82 GB (58 GB main weights + 24 GB n-gram tables) Real-world quants likely land in the 80–90 GB range. The big n-gram table is sparsely accessed → excellent candidate for system RAM offload. This… 36 llama.cpp releases dev-tools 4d ago b10625 chat : scope qwen3-coder workarounds ( #27679 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/42919681 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework… 21 The Information — AI news-outlet 4d ago SpaceX Announces $100 Billion Starship Launch Site in Louisiana SpaceX announced on Tuesday plans to build a $100 billion launch site for its Starship rocket on the coast of Louisiana. SpaceX plans to begin construction of the site, which the company has named “Starbase, Louisiana,” in 2027 and is targeting 2029 for the first Starship launch… 22 Vercel — AI dev-tools 4d ago Vercel applications are protected from Next.js August 2026 security vulnerabilities Summary Two vulnerabilities affecting Next.js were disclosed in the August 2026 Security Release. Next.js applications hosted on Vercel are protected and require no customer action. Next.js August 2026 vulnerabilities Next.js disclosed the following critical vulnerabilities:… 23 r/LocalLLaMA community 4d ago Mac Studio M5 Max Cost Analysis At $10k, you could get - 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan) - 5.7B tokens with DeepSeek V4 Pro OpenRouter - 100B tokens with DeepSeek V4 Flash OpenRouter As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait… 26 r/LocalLLaMA community 4d ago Qwen 4 architecture: What do we know? My best bet is the embedding-offloaded linear where the 51b n-grams track semantics and context, like 3.8 with its loss of real world knowledge bolted back on. Otherwise? Sparse full-attention with dense routing? 6b active route to 'heavy' layers when the n-gram gets stuck.… 5 r/LocalLLaMA community 4d ago Apple releases M5 ultra at 1.2TB/s bandwith lpddr5x probably, the m7 ultra if is using ddr6 should be at 1.8 Tb/s   submitted by   /u/Last-Owl-8342 [link]   [comments] 4 The Information — AI news-outlet 4d ago Citadel Securities, DTCC Partner LayerZero Launches New Exchange LayerZero, a blockchain infrastructure startup that counts Citadel Securities, DTCC, and Intercontinental Exchange among its partners, said it’s launching a new blockchain-based exchange targeting financial institutions this fall. Unlike major crypto exchanges, LayerZero’s new… 15 r/MachineLearning community 4d ago How we built a SOTA search engine using PostgreSQL, pgvector, and Qwen3 embeddings [P] I wrote a technical breakdown of how search works on Papers with Code. The system combines keyword and semantic search, which produced better results than either approach alone. The stack includes: PostgreSQL with pgvector Qwen3-Embedding-0.6B for text embeddings Hugging Face… 31 r/LocalLLaMA community 4d ago Qwen 3.8 Flash Next day 0 support from unsloth Prepare your disk space guys   submitted by   /u/jacek2023 [link]   [comments] 18 Hacker News — AI on Front Page community 5d ago Qwen 3.8-Flash-Next releasing tomorrow (125B a6B) Article URL: https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next Comments URL: https://news.ycombinator.com/item?id=49432317 Points: 201 # Comments: 86 5 r/LocalLLaMA community 5d ago New: Llama.cpp adaptive speculation for faster inference We have been working on some performance optimisations for Qwen3.8 and other models. The main new feature that we introduced is adaptive speculation for Llama.cpp What is it? MTP and DFlash work well to speed up inference work, especially for dense models. However, different… 36 r/LocalLLaMA community 5d ago Qwen3.8 flash next   submitted by   /u/RuthlessCriticismAll [link]   [comments] 25 r/LocalLLaMA community 5d ago Qwen3.8-Flash-Next tomorrow   submitted by   /u/rerri [link]   [comments] 15 llama.cpp releases dev-tools 5d ago v0.3.0 Overview llama.cpp 0.3.0 introduces the dots3-note multimodal model (with a new DSA-ISWA KV cache), MTP support for GLM-4.5-Air, and tensor-split ( -sm tensor ) plus multi-sequence rollback fixes for DeepSeek 4. ggml is bumped to v0.22.0 (meta-backend tensor split, per-op Metal… 23 r/LocalLLaMA community 5d ago Glm 5.3 flash? glm 5.3 flash While awaiting the release of the version 5.3 weights, this theory is gaining ground. OxAlpha is new GLM.   submitted by   /u/LegacyRemaster [link]   [comments] 13 llama.cpp releases dev-tools 5d ago b10621 llama.cpp : bump version to 0.3.0 ( #27696 ) llama.cpp : bump version to 0.3.0 ci : update release default desc scripts : add prompt for generating release summary Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/42818481 macOS/iOS:… 32 r/LocalLLaMA community 5d ago tencent/WeMM-Embedding 9B/4B/2B WeMM-Embedding-9B is a universal multimodal embedding model built on Qwen3.5. It accepts text, images, videos, visual documents, and interleaved multimodal inputs, and returns a 4,096-dimensional L2-normalized embedding. Audio input is not supported.… 10 llama.cpp releases dev-tools 5d ago b10618 grammar : parse - in char classes as literal hyphen ( #27591 ) grammar : accept "-" escape in character classes gbnf_escape_char_class() escapes '-' as "-" but parse_char() rejected that escape, so generated tool-call grammars failed to parse. Assisted-by: Claude Code… 21 r/MachineLearning community 5d ago Reviewing 4 papers for AAAI 2027 and none have code, Reject? [D] I got my batch of four papers for AAAI 2027. All four papers make empirical claims, none include code, data, or anything I can actually check. Just the PDF and the checklist. AAAI-27's own rules say code/data should be provided at submission, and "we'll release it after… 8 Ollama releases dev-tools 5d ago v0.33.0 What's Changed Claude Desktop Turn individual Ollama models on or off for use in Claude, directly from the menu bar Choose from your available Ollama models from within Claude; cloud models appear only when you're signed in A new Apps view manages app integrations with copyable… 4 Vercel — AI dev-tools 5d ago Introducing Run SDK: secure eval for your agents Agents increasingly write TypeScript programs to coordinate tools and process their results. Once those programs touch real applications, some steps require authentication, while others need human approval. Executing that code with eval gives it the same access as the… 34 LangChain releases dev-tools 5d ago langchain==1.3.17 Changes since langchain==1.3.16 release(langchain): 1.3.17 ( #39893 ) fix(langchain): frame custom HITL rejection reasons ( #39773 ) chore(deps): bump minor and patch dependencies ( #39869 ) 33 OpenAI official-blog 5d ago Introducing the Admin plugin for ChatGPT Work and Codex Use the Admin plugin for ChatGPT Work and Codex to analyze workspace usage, manage members and permissions, adjust limits, and act on admin requests. 34 Vercel — AI dev-tools 5d ago Wan 3.0 now available on AI Gateway Wan 3.0 from Alibaba is now available on AI Gateway as alibaba/wan-v3.0-video . One model covers text to video, image to video, first and last frame conditioning, and reference-based generation, and it takes image, video, and audio as references. Clips run up to 30 seconds at… 29 r/LocalLLaMA community 5d ago [2608.16157] FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution Source of Claims: https://x.com/Andy_ShuoYang/status/2090856976880472439 Your gaming PC can now serve frontier models at interactive speed using official checkpoints without extreme quantization! Qwen3.6 35B → 8GB RTX 4060 laptop @ 39 tok/s DeepSeek-V4-Flash 284B → RTX 5090… 22 The Information — AI news-outlet 5d ago Alabama Starts Probe Into OpenAI Over Hugging Face Hack Alabama Attorney General Steve Marshall launched an investigation into OpenAI’s practices related to its AI models’ recent cyberattack against Hugging Face. The investigation is examining whether OpenAI’s “inability or unwillingness to ensure the safety of its products” violates… 18 r/LocalLLaMA community 5d ago I just tried DeepSeek Harness and it escaped from its workspace folder It worked pretty well, digging through and analyzing some local files. Claude code regularly stops at some point and fails to continue while DSH worked for 2 h, recognized that it could benefit from reading more context and ... bummer: It left the project folder (although DSH… 34 The Information — AI news-outlet 5d ago Meta Plans to Launch ‘Hatch’ AI Agent Platform in Coming Weeks Meta Platforms plans to launch its consumer version of the OpenClaw AI agent, dubbed Hatch internally, as soon as the next several weeks and is targeting October for its latest AI model, called Watermelon, according to internal documents reviewed by The Information. Hatch is… 22 r/LocalLLaMA community 5d ago The journey of letting Qwen 3.6/3.8 autonomously coding a c compiler. Hi, Back in late march I begun playing around with Qwen 3.6 27b and found like everyone else that it's notoriously good at tool calls, where every model I tried before just derailed after a few turns it kept going and felt quite reliable outside of typical behaviors of smaller… 22 r/LocalLLaMA community 5d ago Qwen-3.8-27B, Nemotron-3.5-Lightning-30B-A3B, Ornith-1.5-35B-A3B, Muse-Glimmer-30B oQ8e comparison Ornith does really well. TielCoder ( https://llm-bench.io/benchmarks/cmt7kp2zj002r01lcmpchvlko ) might be even a bit better in coding. Will give it a try soon. Details of the comparison see here:… 35 r/LocalLLaMA community 5d ago JetBrains local AI (using Qwen3.6 27B) Sounds quite interesting, a big IDE provider optimizing for local AI with their coding harness. Especially that they picked Qwen3.6 over Qwen3.8 because of the thinking needs. Haven't read the full article yet, but sounds really cool.   submitted by   /u/Danmoreng [link]… 36 The Information — AI news-outlet 5d ago Inside Musk’s First Address to Cursor: Grok Is Falling Behind Earlier this month, on the day SpaceX announced it had completed its $60 billion acquisition of Cursor, Elon Musk called into an all-hands video meeting with the coding startup’s staff. Musk told them that SpaceX’s AI unit had fallen behind competitors and that he wasn’t used to… 36 r/LocalLLaMA community 5d ago Planning to spend ~$100 benchmarking differnet Qwen3.8-27B quants and kv cache and looking for input before I start TL;DR: I'm planning to spend around $100 on cloud GPUs to benchmark Qwen3.8-27B with a focus on questions that actually matter when running it locally: different quant levels/providers, 8-bit vs 16-bit KV cache, GGUF vs EXL3, context length tradeoffs, and token efficiency on… 29 r/LocalLLaMA community 5d ago Do not blindly delete your older models, some are still precious I have deleted tons and tons of older models to make space since I can't afford storage anymore. Easily 10TB... Anyways, I have been considering deleting DeepSeekV3.2 but decide to run it one more time. I have a problem I have been brainstorming about and have chatted locally… 27 r/LocalLLaMA community 5d ago This is what Qwen 3.8 27b is capable of Try it here: https://ocean.blackbeardlabs.dev/ Model: Qwen 3.8 27b Q8_X_KL Unsloth Hardware: 3 x RTX3090 Harness: DeepSeek Harness Prompt: /goal I want you to create a **JavaScript + Node.js WebGL project** that renders a highly realistic real-time ocean in the browser. Use… 14 r/LocalLLaMA community 5d ago Qwen 3.8 27B in 9th position on code arena. Gemma 4 31B is 80th.   submitted by   /u/tarruda [link]   [comments] 16 Simon Willison community 5d ago llm-anthropic 0.27 Release: llm-anthropic 0.27 This release of the Anthropic plugin for LLM mainly provides compatibility with the recently released anthropic v1.0.0 Python library, which switches from httpx to httpx2 . OpenAI made the same change in their v3.0.0 release two weeks ago. Anthropic… 36 Page 4 of 10 · 500 articles ← Newer Older →