News / #music Tag Music 430 articles archived under #music · RSS Sign in to follow r/LocalLLaMA community 9h ago Nemotron-3.5-Lightning at 11.77 GiB, a 16 GB option for a model that didn't have one TL;DR: Every public low-bit GGUF of this model is secretly ~4.70 bpw. Shim the rows to 256 and it becomes a real 3.07 bpw / 11.77 GiB file that runs 262K context on 16GB. Needs patched llama.cpp — not LM Studio or Ollama. In the AtomicChat HuggingFace repo it says "There is… 4 r/LocalLLaMA community 18h ago Exo labs claiming 4.8 tb/s memory bandwidth through m5u Mac Studio clustering Exo labs making some very exciting and interesting claims. The headline is bandwidth scales linearly on Mac Studio clusters with their solution. There is a thread over at localllm subreddit ( https://www.reddit.com/r/LocalLLM/s/qEYLOFaYwc ) where one of their employees speaks… 17 r/LocalLLaMA community 1d ago AtomicChat/Qwen3.8-Flash-Next-GGUF is Really Good specs hardware: M4 Max 128GB Studio inference engine: llama.cpp (qwen4exp branch) judge: claude-opus-4-6 AtomicChat/Qwen3.8-Flash-Next-GGUF Qwen3.8-Flash-Next is a great model I benched in my previous post , but it is very tight, since all n-grams / PLE are loaded along with the… 17 arXiv — NLP / Computation & Language research 2d ago When Text Misleads: Inconsistent-Aware Reasoning for Audio-Grounded Dialogue arXiv:2608.27176v1 Announce Type: new Abstract: Understanding spoken dialogue requires joint reasoning over lexical content and paralinguistic acoustic signals such as emotion and conversational intent. However, existing evaluations often allow shortcuts based on transcripts or… 32 r/LocalLLaMA community 2d ago 5090 now officially cost 5090 I was planning on another 5090, but then I realize... perhaps I am much better off getting an M5 Ultra Mac Studio with 256gb of ram. We are so genuinely cooked.   submitted by   /u/Sadge404 [link]   [comments] 5 r/LocalLLaMA community 2d ago Let’s be real, memory and gpus price will continue go up next year and the year Models will continue improve for open and closed labs, the demand for compute and memory will continue to increase. Expect to pay double or more for ddr 5 and 6 ram and +60% plus for new consumer gpus . Even a 512 gb mac studio will likely be over 22k .   submitted by  … 8 r/LocalLLaMA community 2d ago [audio.cpp] Release 0.7: 62 audio model families (85+ variants), Arena UI for model comparison, MiniMax Music 3, FireRed TTS3/Audio, ControlFoley, Personaplex, and more audio.cpp 0.7 is out :) This release adds a lot of new audio models and a new way to compare them locally. Audio.cpp is now at 62 model families and 85+ model variants. And it keeps growing! The biggest user-facing change is the new Arena UI . Instead of testing one model at a… 22 Google DeepMind official-blog 2d ago Gemini Omni 1.1 Flash lets you build with more control Gemini Omni 1.1 Flash lets you build with more control Aug 27, 2026 | x.com Facebook LinkedIn Mail Omni now delivers studio-quality video production, including the ability to extend a scene, first and last frame interpolation, crisp 4K upscaling, faster prototyping, and more.… 26 r/LocalLLaMA community 2d ago Qwen3.8-Flash-Next: Time to Update Those Benchmarks specs hardware: M4 Max 128GB Studio inference engine: oMLX & lllama.cpp insights it still very early, so had to disable oMLX K/V caching, qwen4_exp architectureis not yet supported + the obvious n-grams with which the whole 4 bit quant takes ~100G, so pretty tight nevertheless,… 35 Hugging Face Daily Papers research 3d ago Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds Abstract JoyAI-Echo-1.5 unifies long-form video and interactive world generation through cross-shot memory, geometry-aware camera control, and rollout-aware training to maintain identity and coherence over extended sequences. Generated by thinkingmachines/Inkling-Small Video… 36 Hugging Face Daily Papers research 3d ago Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios Abstract A new benchmark evaluates how well multimodal language models follow diverse video-based instructions with visual, audio, and structural constraints. Generated by thinkingmachines/Inkling-Small Multimodal Large Language Models (MLLMs) have shown strong performance in… 14 arXiv — NLP / Computation & Language research 3d ago Can We Read the Mind of an Audio LLM? A Verbalizable, Multilingual Middle-Layer Workspace arXiv:2608.24958v1 Announce Type: cross Abstract: An audio language model is a black box in a specific way: we see what it says, never what it works out on the way there, and chain-of-thought monitoring helps only if the model writes its reasoning down. Reading a base Qwen3-Omni… 8 Stratechery (Ben Thompson) community 3d ago Apple Updates Mini and Studio, AI Computers, OpenAI Jalapeño Apple and OpenAI have two completely different hardware announcements; both represent pressure on Nvidia. 30 Hugging Face Daily Papers research 4d ago LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training Abstract LAION-BVD is a large-scale open video dataset enabling multimodal pre-training across video, audio, and image modalities with synthetic captions and strong benchmark performance. Generated by thinkingmachines/Inkling-Small We present LAION-BVD, a large-scale open video… 22 Vercel — AI dev-tools 4d ago Gemini 3.5 Transcribe now available on AI Gateway Gemini 3.5 Transcribe from Google is now available on AI Gateway. It takes audio and returns text, in two variants: google/gemini-3.5-transcribe transcribes a complete recording in a single request. google/gemini-3.5-transcribe-live transcribes audio over a WebSocket, returning… 18 r/LocalLLaMA community 4d ago M5 Ultra 96GB vs M5 Max 128GB — is 2x bandwidth worth losing 32GB of RAM, with Qwen3.8-Flash-Next dropping tomorrow? I’ve been going back and forth on this for a week and I can’t settle it, so I’m hoping someone here has hands-on numbers. The two configs (German prices, dealer quote, incl. VAT): Config Price Mac Studio M5 Max, 128GB / 512GB SSD €5,859 Mac Studio M5 Max, 128GB / 1TB SSD €6,189… 32 r/LocalLLaMA community 4d ago Mac Studio M5 Max Cost Analysis At $10k, you could get - 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan) - 5.7B tokens with DeepSeek V4 Pro OpenRouter - 100B tokens with DeepSeek V4 Flash OpenRouter As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait… 26 r/LocalLLaMA community 4d ago Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory   submitted by   /u/themixtergames [link]   [comments] 18 Hacker News — AI on Front Page community 4d ago Apple Introduces New Mac Studio with M5 Max and M5 Ultra Article URL: https://www.apple.com/newsroom/2026/08/apple-introduces-new-mac-studio-with-m5-max-and-m5-ultra/ Comments URL: https://news.ycombinator.com/item?id=49433316 Points: 277 # Comments: 162 26 r/LocalLLaMA community 4d ago tencent/WeMM-Embedding 9B/4B/2B WeMM-Embedding-9B is a universal multimodal embedding model built on Qwen3.5. It accepts text, images, videos, visual documents, and interleaved multimodal inputs, and returns a 4,096-dimensional L2-normalized embedding. Audio input is not supported.… 10 arXiv — NLP / Computation & Language research 5d ago PUMA: A Polish Benchmark for Culturally Grounded Multimodal Understanding arXiv:2608.21853v1 Announce Type: new Abstract: Large language models are increasingly moving beyond text processing, adding support for other modalities such as images and audio. While text understanding and generation have been extensively studied, multimodal data processing… 12 arXiv — NLP / Computation & Language research 5d ago WnW: Waxing-and-Waning KV Cache for Long-Form Speech LLMs arXiv:2608.22704v1 Announce Type: new Abstract: Long-form audio inputs make the KV cache the dominant memory cost of speech LLMs. Prefill-only KV compression methods permanently discard audio KV positions once evicted, with no pathway to recover them during decoding. We show… 11 Hugging Face Daily Papers research 5d ago TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming Abstract TLive-Omni is an omni-modal model for live-commerce that unifies image, video, audio, and text via timestamped token grouping, staged supervised training, and reinforcement fine-tuning with verifiable feedback to enable accurate real-time understanding. Generated by… 24 Vercel — AI dev-tools 5d ago Wan 3.0 now available on AI Gateway Wan 3.0 from Alibaba is now available on AI Gateway as alibaba/wan-v3.0-video . One model covers text to video, image to video, first and last frame conditioning, and reference-based generation, and it takes image, video, and audio as references. Clips run up to 30 seconds at… 29 r/LocalLLaMA community 5d ago What's the best local model you've found for 8 GB of VRAM? I'm curious what other people are using for local LLM coding / agentic coding with only 8 GB of VRAM . My current setup is: Intel Core i7-11800H RTX 3070 Laptop , 8 GB VRAM 32 GB DDR4 RAM openSUSE Tumbleweed / KDE Unsloth Studio pi.dev as the coding agent After testing quite a… 33 r/LocalLLaMA community 5d ago iPhone Local TTS EPUB Reading - Audiobookify I've been working on this project (been a developer for a few years) for a few months now, and it has been in active testing for ~2 months. Its an EPUB reader that also offers local offline TTS, so its not just TTS focused, its meant to be a good regular reading app as well. I'd… 38 arXiv — Machine Learning research 6d ago AudioWorldSim: Realistic Binaural Audio Datasets For World Models arXiv:2608.21075v1 Announce Type: cross Abstract: This technical report presents AudioWorldSim, an open-source platform designed to generate realistic binaural audio datasets and advance research in audio-based machine learning, particularly world models. Built as a custom… 37 arXiv — NLP / Computation & Language research 6d ago Building and Evaluating a Synthetic Bengali Speech Resource for Telecom Customer Care arXiv:2608.20346v1 Announce Type: new Abstract: Speech systems used in customer-facing applications often require domain-specific language coverage. We present a synthetic Bengali speech dataset for telecom customer-care scenarios. The dataset contains 10,000 audio-text pairs,… 8 r/LocalLLaMA community 6d ago I trained a game music generator I trained a instrumental game music generator. The 1.2B DiT was trained on 1 cloud H100 from scratch in 8 days; I used the VAE from Stable Audio 3. https://huggingface.co/Localsong/Localsong https://huggingface.co/Localsong/Localsong/tree/main/samples… 23 r/LocalLLaMA community 7d ago Current best model for narrative, chat, prompt creation (so basically everything except agentic coding)? - 5090 Im looking to set up a new local llm (probably on unsloth studio as that seemed to be doing pretty well last time I tested it). This one won't need to do agentic coding or app building or anything (not this time) but instead more 'text' based tasks such as - being given… 10 llama.cpp releases dev-tools 7d ago b10580 mtmd: support dots3-note vision+audio ( #27524 ) text: conversion init impl mtmd: conversion impl mtmd cpp Update gguf-py/gguf/tensor_mapping.py Co-authored-by: Sigbjørn Skjæret sigbjorn.skjaeret@huggingface.co Co-authored-by: Sigbjørn Skjæret sigbjorn.skjaeret@huggingface.co… 22 r/LocalLLaMA community 8d ago I tried to do agenic coding with Qwen 3.8 27B 3bit quant on a macbook air m2 24gb. It took 63 hours, but amazingly, the flight simulator worked. I used LM Studio Bionic with Qwen 3.8 27B Q3_K_S with 57k context. It took a staggering 63 hours to finish coding. After the first prompt "Create a beautiful, relaxing flight simulator in a single HTML page" taking 47.8 hours, it created an html file that showed the title screen… 15 r/LocalLLaMA community 8d ago I feel like I finally graduated. I finally made the move from LM Studio to vLLM thanks to this post https://www.reddit.com/r/LocalLLaMA/s/NmS9CgHvqz . I may not know what it all means yet but I’m going to start diving into the docs to learn as much as I can. I’m running an endpoint on each of my 3090s one for… 4 r/LocalLLaMA community 8d ago model: add dots3-note by ngxson · Pull Request #27060 · ggml-org/llama.cpp dots3-note preview is the first open-weight model in the dots3 family. It is a Mixture-of-Experts model with 280B total parameters, 16B activated parameters, and support for a context length of up to 512K tokens. The model can understand text, images, video, and audio, and… 14 r/LocalLLaMA community 8d ago FireRedAudio & FireRedTTS3 by FireRedTeam - Huggingface FireRedAudio: A General-Purpose Audio Language Model with Decoupled Continuous Representations for Understanding and Generation HuggingFace : https://huggingface.co/FireRedTeam/FireRedAudio GitHub : https://github.com/FireRedTeam/FireRedAudio Demo :… 33 r/LocalLLaMA community 8d ago Ultrafast Qwen3-TTS at 34 ms Time-to-First-Audio, Handling 10 Requests Per Second [OSS] Hey locallama! We recently open sourced a Qwen3-TTS 1.7B implementation that achieves 10 requests per second (RPS) and sub-50 ms p95 time-to-first-audio (TTFA) while maintaining real-time playback on 1 x H100. This extends to 20 RPS at sub-100 ms p95 TTFA. By adjusting settings,… 31 Google DeepMind official-blog 8d ago From Atari to EVE Online: Building on 15 Years of AI Research in Games Google DeepMind partners with game studios to prototype breakthrough AI gameplay. 28 Hugging Face Daily Papers research 8d ago Listening Forward: Next Patch Embedding Prediction Enables Scalable Audio Learners Abstract NAPE uses causal Transformers to predict successive spectrogram patch embeddings for self-supervised audio learning without auxiliary components. Generated by thinkingmachines/Inkling-Small Self-supervised learning (SSL) has driven substantial progress in audio… 6 Hugging Face Daily Papers research 9d ago Towards Quantifying Benchmark Optimization in ASR Models Abstract High-performing speech recognition models reproduce benchmark transcripts despite contradictory audio, revealing benchmark-optimized behaviors that inflate scores without improving real-world transcription. Generated by thinkingmachines/Inkling-Small Public benchmarks… 13 arXiv — NLP / Computation & Language research 9d ago Represented but Ignored: A Causal Account of Prosodic Underuse in Audio-Language Models arXiv:2608.19211v1 Announce Type: new Abstract: Human speech is richly expressive, with prosody carrying linguistic and emotional information beyond the lexical content. A capable large audio-language model (audio-LLM) should therefore support expressive speech understanding,… 4 r/LocalLLaMA community 9d ago Make Jensen Huang Sound Like Anyone. New Streaming Voice Conversion Model MeanVC2 Released! Finally see a new voice conversion model. MeanVC2 supports cross-gender and cross-language voice conversion. 3x realtime on CPU with audio.cpp. Disclaimer: The converted voice quality of MeanVC2 is decent; the noise comes from my rough demo engineering, not the model itself.… 38 Hugging Face Daily Papers research 9d ago VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation Abstract A human-aligned chain-of-thought reward model and preference dataset improve joint video-audio generation by replacing fragmented metrics with coherent, dimension-wise reinforcement learning. Generated by thinkingmachines/Inkling-Small Using reinforcement learning to… 29 Hacker News — AI on Front Page community 9d ago AliExpress runs silent WebAudio fingerprinting that breaks Bluetooth multipoint Article URL: https://blog.laserphile.com/2026/08/aliexpress-webpage-keeping-multipoint.html Comments URL: https://news.ycombinator.com/item?id=49372583 Points: 368 # Comments: 122 23 arXiv — NLP / Computation & Language research 10d ago Alignment Is All You Need: Instruction-Free Training for General Audio-Language Models arXiv:2608.18132v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are typically built through a multi-stage pipeline consisting of cross-modal alignment, supervised fine-tuning (SFT), and preference optimization. This pipeline assumes that adapting an LLM… 30 r/LocalLLaMA community 10d ago Finally found a really solid suno-like minimax music UI!! Been messing with minimax music gen lately. I really like Suno and was basically looking for something that gave me a similar workflow for minimax. I got completely sick of running everything through the CLI. I went digging on github, sorted by recent, and took a gamble on this… 8 Vercel — AI dev-tools 11d ago Fish Audio models now available on Vercel AI Gateway for free Fish Audio 's audio models are now available on AI Gateway. To celebrate the launch, every Fish Audio model is free on AI Gateway for the next 30 days, through September 18. Capability Regular Through September 18 Text-to-speech $15.00 per million characters Free Speech-to-text… 22 Hacker News — AI on Front Page community 11d ago Claude writing a macOS driver for my obscure HP printer built only for Windows https://xcancel.com/kuberwastaken/status/2089377982536388964 https://cdn.kuber.studio/chat/hp-laser-1008a-driver Comments URL: https://news.ycombinator.com/item?id=49344643 Points: 248 # Comments: 185 5 r/LocalLLaMA community 12d ago Qwen 3.8 27B is faster than expected i ran this model on my two 5060 TI 16GB cards at Q4 in unsloth and LM studio ( i downloaded NVFP4 but didn't try it in vLLM ) i think it runs faster than expected it gives me 50 - 60 t/s with MTP. this is surprising because it's a dense model and Qwen 3.6 was giving me 30t/s… 34 Hugging Face Daily Papers research 12d ago Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift Abstract PRISM is a fast, training-free test-time adaptation method that reverses low-rank affine noise distortions in audio-text models using frozen text prototypes and geometric corrections. Generated by thinkingmachines/Inkling-Small Audio-Text Foundation Models (ATMs) fail… 33 arXiv — Machine Learning research 13d ago DeaMoE: Efficient MoE Structure for Fast Small-Batch Decoding arXiv:2608.14385v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models have been widely adopted in real-time interactive applications such as coding assistants, real-time audio-video interaction systems. To meet the extremely low response latency requirements of these… 26 Page 1 of 9 · 430 articles Older →