News / #model-release Tag Model releases 500 articles archived under #model-release · RSS Sign in to follow Vercel — AI dev-tools 9d ago GPT-5.6 Sol pricing drops on AI Gateway, and 50% discount still applies OpenAI just lowered list pricing for GPT-5.6 Sol . The 50% AI Gateway discount continues and applies to the new list price now through September 18th. Pricing The 50% discount applies on every tier on the OpenAI provider, against the new list price. Service tier You pay now… 36 Vercel — AI dev-tools 9d ago GPT-5.6 Sol is now 50% off a lower price OpenAI lowered list pricing for GPT-5.6 Sol , and the 50% AI Gateway discount now applies to the new, lower price through September 18. Input drops 20%, output drops a third. The discount applies on every OpenAI service tier: Service tier You pay now (input / output) New list… 32 TechCrunch — AI news-outlet 9d ago OpenAI is gaining on Anthropic with business users, new data indicates Businesses are willing to flop back and forth as each lab releases new models, volatility that should give both companies' investors pause about how "sticky" enterprise AI spending really is. 16 r/LocalLLaMA community 9d ago Been tweaking my Qwen 3.8 setup, up to 45+ steady T/ps at 8bit quant. Realised I'm now top T/ps for this model+ctx across all benchmarked M-series chips. Full args linked below, happy to discuss as this was a pain of trial and error. https://omlx.ai/benchmarks/performance/2pko3m1k - you can expand the raw args, but I have full annotations of what worked and what didn't. I'm now testing the model on acutal coding and haven't seen any issues with performance vs default suggested vals for the vanilla model.… 35 r/MachineLearning community 9d ago Notes on Hamiltonian Monte Carlo from a purely probabilistic perspective [P] I’ve been studying Hamiltonian Monte Carlo and wrote a set of notes explaining HMC without relying on the usual physics-based motivation. The notes develop HMC from a probabilistic/MCMC perspective, starting from introducing an auxiliary variable, constructing the corresponding… 32 r/LocalLLaMA community 9d ago Fine-tuning Cactus Needle 2 can match DeepSeek v4 on the specific task Hey LocalLlama, Henry from Cactus here! When we trained Needle 2, I had a strict rule to not expose the model to any data sample that remotely felt like these benchmarks. It seemed over-the-top, but benchmarks are easy to overfit around, yet struggle in the wild, especially… 9 r/LocalLLaMA community 9d ago I pushed Qwen3.8-27B to 381 tps for a single request on a RTX 3090 Four days ago I released a hyper-optimized Qwen3.8-27B inference engine for an RTX 3090 (82 tps single request, 672 peak). Since then it went to ~114, then ~138 tps single-user with DFlash2 drafting and lookup-augmented drafting. Today it's ~133 tps on real chat prompts, 382 tps… 15 r/LocalLLaMA community 9d ago Qwen3.8-27B scored 29/30 on AIME 2026 with FP8 + xhigh reasoning — BF16 vs FP8 results I benchmarked Qwen3.8-27B on MathArena/aime_2026 dataset, comparing BF16 and FP8 weights at medium and xhigh reasoning effort. Interesting findings are: quantized FP8 xhigh is better than BF 16 medium equally good as 16 BF xhigh with better speed. On problem 7, both BF16 xhigh… 28 r/LocalLLaMA community 9d ago Ladies and gentlemen I present to you Qwen3.8 27b 1bit brain damage quant I wanted to just test the unsloth 1bit quant of qwen 3.8 27b as I have just 8gb vram and ngl it gave me a good laugh   submitted by   /u/Ok-Health-7096 [link]   [comments] 6 LangChain releases dev-tools 9d ago langchain-fireworks==1.6.0 Changes since langchain-fireworks==1.5.2 release(fireworks): 1.6.0 ( #39810 ) feat(fireworks): add document reranking ( #39732 ) fix(fireworks): filter invalid tool calls from v1 content ( #39805 ) feat(core): add standard model exception types ( #39538 ) chore(model-profiles):… 4 llama.cpp releases dev-tools 9d ago b10517 vulkan : dequant q8_0 KV once in coopmat1 ( #25494 ) vulkan : dequant q8_0 KV once in coopmat1 Assisted-by: Claude (Opus 4.8) vulkan : fall back instead of aborting when FA scratch exceeds maxStorageBufferRange vulkan : require KV-cache layout in FA dequant path Assisted-by:… 19 TechCrunch — AI news-outlet 9d ago Grok keeps sending gibberish responses to users Affected users told TechCrunch they were using Grok Lite, and noticed the issues as early as Wednesday morning. 31 r/LocalLLaMA community 9d ago Ling-3.0 released all 6 base checkpoints: 2 sizes × 3 stages AntLing has released the full six-checkpoint matrix for the Ling-3.0 base model. tiny: pretrained, mid-trained, WSM-merged flash: pretrained, mid-trained, WSM-merged The concrete artifact is six separate official repositories, not one endpoint repeated under different names. All… 32 TechCrunch — AI news-outlet 9d ago A third of web pages published since ChatGPT’s launch show signs of AI authorship, study finds ChatGPT and other AI models are now authoring and editing much of the new web. 38 r/LocalLLaMA community 9d ago QwenMix-3.7: Kept seeing posts about Qwen3.8 and 3.6 sharing the same structure.. so I had Qwen3.8 combine them. I chose to do this thing, not because it was hard, but because it was silly. Posts kept discussing how 3.8 and 3.6 were functionally the same, but based on training (3.8 does have seven new tokens!).. so I figured I'd see if they could be merged. They can. I used… 34 TechCrunch — AI news-outlet 9d ago Ramp launches its own AI model router, called Router Ramp has launched its own AI model routing service, dubbed Router, that lets users and companies use and switch between various large language models via an API. 28 Hacker News — AI on Front Page community 9d ago Linux 7.2 Article URL: https://www.igalia.com/2026/08/19/Linux-72-Released.html Comments URL: https://news.ycombinator.com/item?id=49376265 Points: 247 # Comments: 96 29 Simon Willison community 9d ago A shot-scraper-style JSON API on Bun 1.4's new Bun.WebView Research: A shot-scraper-style JSON API on Bun 1.4's new Bun.WebView Today saw the long awaited release of Bun 1.4 , the first stable version since the infamous Rust rewrite a few months ago . Interestingly, the Rust rewrite was downplayed in the release notes, which… 35 Hacker News — AI on Front Page community 10d ago Vomit: Clean up Claude 5's token output with a separate LLM Article URL: https://github.com/zachahn/vomit Comments URL: https://news.ycombinator.com/item?id=49375996 Points: 248 # Comments: 245 10 r/LocalLLaMA community 10d ago [MASSIVE TINY RELEASE] - Supra2-Medium-Base - a tiny 25M parameters model competing heavily with our previous 50M model! Hey guys! Supra2-Medium is finally out! It's a 25M parameters qwen3 architecture model trained entirely from scratch (on our new rig: RTX 5060 Ti 16GB + the new RTX 5060 8GB!). Here's how it competes in benchmarks with Supra-50M-Base (which is double as large!!):… 8 LangChain releases dev-tools 10d ago langchain==1.3.16 Changes since langchain==1.3.15 release(langchain): 1.3.16 ( #39806 ) feat(core): add standard model exception types ( #39538 ) feat(langchain): support custom token_counter in ContextEditingMiddleware ( #39754 ) fix(langchain): re-raise non-retryable exceptions in… 37 LangChain releases dev-tools 10d ago langchain-anthropic==1.6.1 Changes since langchain-anthropic==1.6.0 release(anthropic): 1.6.1 ( #39804 ) fix(anthropic): filter invalid tool calls from v1 content ( #39803 ) 7 r/LocalLLaMA community 10d ago Qwen3.8 27b just exceeded my expectations on svg generation :D https://reddit.com/link/1vtkgdj/video/595yn0ckdjkh1/player I wanted to try out Qwen3.8 27B 's SVG capabilities but with something different than the pelican on a bicycle. Promt was literally just : create a single html file with an embedded svg of a cat riding a zebra, riding an… 9 The Information — AI news-outlet 10d ago Robots Are in Their GPT-2 Era It’s no secret that AI-powered robots aren’t so good yet. They struggle with a broad range of simple tasks, from untangling cables to chopping vegetables. Nonetheless, morale is high among roboticists who are flush with venture cash as they work toward a “ ChatGPT moment ” when… 4 The Information — AI news-outlet 10d ago Anthropic’s Enterprise AI Venture Buys Consultancy A joint venture established by Anthropic and Wall Street firms such as Blackstone has made its first acquisition since its July launch as it looks to boost the growth of Claude, Anthropic’s chatbot, among businesses. Anthropic ’s venture, Ode, plans to announce Thursday that it… 8 Ars Technica — AI news-outlet 10d ago Grok exfiltrates user data when malicious instructions are encrypted Cryptographic Context Injection is only the latest way to break an LLM safety guardrail. 36 r/LocalLLaMA community 10d ago The boring way to run Deepseek V4 Flash-0731 130-150 tks - 16x5060ti 16GB over 2 PLX88096 switches Component Validated configuration Motherboard ASRock Rack SPC621D8U-2T/OVH CPU Xeon Gold 6330 (Get gold/platinum if interested in Optane Pmem gimmicks) GPU fabric Two Broadcom/PLX PEX88096 islands, eight GPUs per island GPUs 16 x RTX 5060 Ti 16 GB OS Ubuntu 22.04.5 LTS Kernel… 15 r/LocalLLaMA community 10d ago Aurora-80K releases! A modern tiny language model. I'm introducing Aurora-80K, a small language model with exactly 80 thousand parameters. It uses a factorized 4,096-token vocabulary despite having only 80K parameters. The benchmarks: Wikitext-2 BPB: 3.2902 BLiMP: 52.31% Arc-Easy: 26.05% More information about the model is… 7 r/LocalLLaMA community 10d ago Tencent begins testing its new flagship model Hunyuan Hy4 From the screenshots: Hy4 is now live, labeled "Expert-Level Model" + "Use Tools to Solve Problems" Hy3 is tagged with "New Upgrade," positioned as a brand-new general-purpose model DeepSeek, focused on reasoning, is listed alongside it From SuSu_酥酥👅on 𝕏:… 18 r/LocalLLaMA community 10d ago I just built a mini Kimi-K3 from Scratch under 250$. Already beats GPT-2 (124M)! I pre-trained a 1.02-billion-parameter on Kimi K3 replica trained on 5.00 billion decontaminated tokens for $250. This model has 1.02 billion parameters, of which 145 million are active per token. It is roughly one two-thousandth of K3 by total size. It saw 5,000,003,584 tokens,… 12 r/LocalLLaMA community 10d ago [Draft - Open PR] AVX2: Speed up large batch size prompt processing of IQ models by bartowski1182 · Pull Request #27402 · ggml-org/llama.cpp IQ quants are particularly slow on CPU at large batch sizes (what you'd see for imatrix and perplexity) Benchmark numbers I ran PPL against master and this PR to get speed and numbers on --chunks 50 for Qwen3.6-27B and Qwen3.6-35B-A3B on EPYC 9654 using 24 threads Created pure… 23 llama.cpp releases dev-tools 10d ago b10514 model : GraniteSWAForCausalLM / GraniteMoeSWAForCausalLM ( #25505 ) feat(convert): Add conversion for GraniteSWAForCausalLM Branch: GraniteSWAForCausalLM AI-usage: full (Bob, OpenCode + Qwen3.6-35b) Signed-off-by: Gabe Goodhart ghart@us.ibm.com feat(llama): Add granite_swa… 38 r/LocalLLaMA community 10d ago AirLLM - Recent Updates - with Qwen3.8-27B, Kimi-K3 too AirLLM dramatically reduces inference memory usage, letting 70B large language models run on a single 4GB GPU card — without quantization, distillation, or pruning. You can even run 405B Llama 3.1 on 8GB , DeepSeek-V3 (671B) on ~12GB , and Kimi K3 (2.8T) — the largest… 12 r/LocalLLaMA community 10d ago G9v3-39A5B on artificialanalysis looks good. Has anyone tested it? I see https://github.com/linuxid10t/llama.cpp/tree/feature/g9v3-support but yeah... Considering that some results place it above Qwen 3.6 27B (the top performer until a few days ago) and that it is an MoE model, I think it could be interesting.… 22 r/LocalLLaMA community 10d ago New benchmark just dropped! The pelican on a bicycle is sooo outdated, so I came up with a new, improved version. Qwen3.8-27b medium (UD-Q4_K_XL) vs. Sol 5.6 high vs. Qwen3.6-35B (UD-Q6_K_XL) Prompt (only real with typo!): "Create a svg of a horse on a blue bycicle in the desert, with a camel in the… 15 TechCrunch — AI news-outlet 10d ago Binance now lets AI agents trade, but keeping them in check is largely up to users Binance's Agent OS works with tools including ChatGPT, Claude Code, and Cursor. 21 r/MachineLearning community 10d ago Discussion thread for EMNLP 2026 Notifications/Results [D] Discussion thread for EMNLP 2026 notifications/results which should be released today. Wishing everybody to be in Budapest.   submitted by   /u/sweetsalt10 [link]   [comments] 25 r/LocalLLaMA community 10d ago Qwen 3.8 27B KV f16 vs q8_0 are not equivalents I'm testing it since release, now with UD 3.0 in my AMD R9700 with ROCm, I always read everywhere that F16 and q8_0 for KV cache are essentially the same... well, I tested it and I can see differences. Some differences are minimal, F16 is more careful and detailed with… 9 OpenAI official-blog 10d ago Introducing AI Futures Introducing AI Futures, a new OpenAI blog exploring how transformative AI could reshape power, governance, the economy, and individual freedom. 4 OpenAI official-blog 10d ago Introducing Intelligence Age Introducing Intelligence Age, a new OpenAI blog exploring how transformative AI could reshape power, governance, the economy, and individual freedom. 37 r/LocalLLaMA community 10d ago 3 days benchmarking most llama.cpp flags on my weird 40gb vram laptop + tb4 egpu setup. Got +70% generation, +40% prefill, 60k more context, and filed a bug in llama around MTP. What I learned. tldr: went from 16~ t/s to 27~ t/s generation. got my usable context up from 220k to the full 262k without sacrificing anything. prefill also increased from 376 to 573 command I ended up with, fwiw: llama-server -m Qwen3.8-27B-UD-Q6_K_XL.gguf -c 262144 -ngl 999 -fa on \ -ctk… 5 arXiv — Machine Learning research 10d ago Mechanistic Interpretability of Structure-Aware Numerical Reasoning in LLaMA 3.1 8B arXiv:2608.18419v1 Announce Type: new Abstract: Recent work has shown that large language models (LLMs) exhibit strong numerical sequence modeling capabilities and show promise in time-series prediction. While LLMs display in-context learning capabilities, the mechanisms with… 8 arXiv — Machine Learning research 10d ago ProxyGuard: Direct Reliability Inference for Randomized Data Release Mechanisms with Shared Targets arXiv:2608.18643v1 Announce Type: new Abstract: Researchers often choose a proxy dataset from many releases, transformations, or seeds. Search can make an invalid release appear adequate, while one adequate release does not establish that its generator is reliable. ProxyGuard… 26 arXiv — NLP / Computation & Language research 10d ago Decomposing Wrong-Consensus Agreement in LLM Self-Consistency: A GPT-4.1 Case Study arXiv:2608.18795v1 Announce Type: new Abstract: Majority voting over multiple LLM samples is widely used to raise answer accuracy, yet its gain varies erratically: on hard questions it can even backfire. This paper gives a quantitative account of this failure. A pluralistic… 26 arXiv — NLP / Computation & Language research 10d ago Introducing the Privacy-HSD Trade-off: Hate Speech Detection, but not at the Cost of Privacy arXiv:2608.19006v1 Announce Type: new Abstract: Hate speech is a real and timely threat that affects a large portion of online users, especially youth and minority groups. While building reliable and robust automatic hate speech detection (HSD) systems is paramount, we argue… 24 arXiv — NLP / Computation & Language research 10d ago Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale arXiv:2608.19026v1 Announce Type: new Abstract: Released in 2025, Institutional Books: Harvard Library (IB-HL) is a collection of 983,004 volumes (242B o200k_base tokens), originally digitized through Harvard Library's participation in the Google Books Library project. As… 17 arXiv — NLP / Computation & Language research 10d ago What is Missing from AI Post-Training AI: An Empirical Analysis arXiv:2608.19072v1 Announce Type: cross Abstract: Large language model (LLM) agents can now post-train an LLM end-to-end. They can write code, launch training, evaluate checkpoints, and improve downstream performance, raising the prospect of AI-for-AI. We argue that this picture… 13 r/LocalLLaMA community 10d ago Qwen3.8-27B took a serious hit to *knowledge* vs 3.6 Like many of you I've spent the last few days throwing Qwen3.8-27B against all of my usual use-cases and personal tasks/harnesses and workflows. It's great, phenomenal sometimes, but that's not what this post is about. One of my little personal benchmarks is a little set of… 10 r/LocalLLaMA community 10d ago Qwen3.8-27b has the highest level of "agency" I've ever seen in a local model Off a single prompt, given my credentials and the name of my university, qwen3.8-27b was able to successfully pull my class schedule from the kinda shitty and convoluted web of university websites. It needed no human intervention, and executed 80 tool calls. Another time, I… 28 OpenAI official-blog 10d ago How ChatGPT Work helps Stampli move ideas to market With a fixed deadline and design resources committed elsewhere, Stampli used Codex and ChatGPT Work to compress weeks of launch production into days. 7 Page 7 of 10 · 500 articles ← Newer Older →