News / #music Tag Music 430 articles archived under #music · RSS Sign in to follow arXiv — NLP / Computation & Language research 13d ago Measuring Fairness in Large Audio Language Models via Semantic-Aware Bias Estimation arXiv:2608.13624v1 Announce Type: new Abstract: Large Audio Language Models (LALMs) have seen increasing use for audio understanding tasks such as speech recognition and audio question answering, raising concerns about fairness across demographic subgroups. Fairness evaluation… 7 arXiv — NLP / Computation & Language research 13d ago StreamHear: Domain-Adapted Pseudo-Labeling for Semi-Supervised Streaming Speech Recognition arXiv:2608.13717v1 Announce Type: new Abstract: Streaming automatic speech recognition (ASR) underperforms on domain-shifted target audio, where labeled in-domain data is costly to prepare while unlabeled audio is abundant. We present StreamHear, a semi-supervised pipeline that… 19 arXiv — NLP / Computation & Language research 13d ago Lost in Phonation: Voice Quality Variation as an Evaluation Dimension for Speech Foundation Models arXiv:2510.25577v2 Announce Type: replace-cross Abstract: Recent advances in Speech Foundation Models (SFMs) enable direct processing of raw audio, allowing models to respond to subtle paralinguistic variation. However, how these models interpret non-lexical cues remains largely… 6 r/LocalLLaMA community 13d ago [audio.cpp] Release 0.6: dots.tts, MiniMax-H3 text2audio (up to 3x realtime), MiniMax-Music3 (preview), and more new audio models. 5+ demos included. Hi all :) audio.cpp release 0.6 has been out for a little while, so this is more of an update on what landed and what has been improving around it. 0.6 added 5 new model families: dots.tts, NeuTTS-2e, MuScriptor (Music to MIDI), MiniMax-H3, and SenseVoice-Small, bringing… 25 r/LocalLLaMA community 14d ago 5090: Windows or Linux for Qwen3.8.27b I've got a dedicated AI rig sitting here with a RTX 5090 and 96GB RAM and for the past few years have been using Windows 11 and primarily LM Studio, but have also used vLLM, llama.cpp and Ollama. With Qwen3.8.27b I want to get the most out of this model. I get the feeling from… 14 Simon Willison community 14d ago CORS Chat Tool: CORS Chat I built this today ( with GPT-5.6-Sol xhigh ) to help test Qwen 3.8 27B running in LM Studio on both my M5 MacBook Pro and an NVIDIA DGX Spark. It provides a web UI for exercising an OpenAI-Responses-compatible chat endpoint. I've tried it against LM Studio with… 36 r/LocalLLaMA community 14d ago (Newbie) what do you use local models for? I’m a long time lurker and have a newbie rig. i run Qwen 3.6 27B and I use LM studio. I use it to summarize long and terse financial and legal documents that I don’t want to upload to cloud. that isn’t everyday though and I am not really learning anything new. id like for this… 32 r/LocalLLaMA community 15d ago Recommend harness for local coding? I am looking to start coding locally with qwen3.8 27B, but what is a good harness + backend for this? LM studio is not cutting it   submitted by   /u/Desperate-Data-3747 [link]   [comments] 25 r/LocalLLaMA community 15d ago Does anyone else suddenly experience unusually fast model loading times? In the last week or so, in LM Studio, i've had models load into memory very fast for some reason, even when i'm loading them from HDD. I know that if you just had a certain model in memory, ejected it, and then try to reload it right away, it often loads almost instantly,… 13 Hugging Face Daily Papers research 15d ago From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs Abstract Researchers propose a black-box red-teaming method using inaudible low-frequency waveforms to expose vulnerabilities in audio-language models, alongside a defense that detects distribution shifts and requests a second recording to recover accuracy. Generated by… 37 Hugging Face Daily Papers research 16d ago UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Abstract UniSwap enables synchronized appearance and voice replacement in talking videos through a unified streaming audio-visual diffusion transformer with specialized training and inference adaptations. Generated by thinkingmachines/Inkling-Small Talking-video character… 31 arXiv — NLP / Computation & Language research 16d ago EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory arXiv:2608.12627v1 Announce Type: cross Abstract: Long-horizon egocentric memory transforms continuous first-person video and audio into a searchable record of past experiences. We demonstrate two bottlenecks in existing systems: indices built from context-poor captions are… 11 r/LocalLLaMA community 16d ago 1BIT Qwen 3.8 2.4T a95b (unsloth iQ1_S) (MEDIUM Reasoning) Processing img az99qopcg8jh1... So same as my prior post 1bit test... although this 1bit is a bit interesting you can read on unlsoth blog https://unsloth.ai/docs/models/qwen3.8 508 gigs being used I am using unsloth studio on the mac ultra 512. Im getting ~50pp and ~9.6 tgen… 28 r/LocalLLaMA community 16d ago dots-studio/dots3-note-prev · Hugging Face dots3-note preview is the first open-weight model in the dots3 family. It is a Mixture-of-Experts model with 280B total parameters, 16B activated parameters, and support for a context length of up to 512K tokens. The model can understand text, images, video, and audio, and… 27 r/LocalLLaMA community 17d ago Minimax Music 3 open weight release soon? Diffusers has a PR with deets: https://github.com/huggingface/diffusers/pull/14456 Minimax is working on this repository right now and put up a bunch of samples: https://github.com/MiniMax-AI/music3-demo/tree/main/assets/audio/tracks Comfy-Org is teasing about a big release in… 30 arXiv — NLP / Computation & Language research 17d ago Easper: An Accessible ASR Pipeline for Language Documentation arXiv:2608.11629v1 Announce Type: new Abstract: Audio transcription is a critical bottleneck in language documentation. While multilingual Automatic Speech Recognition (ASR) models like Whisper offer solutions, field linguists often lack the expertise to utilise them. We present… 4 arXiv — NLP / Computation & Language research 17d ago Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning arXiv:2608.11587v1 Announce Type: cross Abstract: Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challenging due to limited labeled data, low… 18 arXiv — NLP / Computation & Language research 17d ago Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder arXiv:2608.11650v1 Announce Type: cross Abstract: Recent advances in zero-shot text-to-speech (TTS) have substantially improved speech quality and voice cloning fidelity. However, many zero-shot TTS systems still depend on audio prompt transcripts at inference time. This… 8 Hugging Face official-blog 17d ago Introducing OlmoEarth embeddings: Custom embedding exports from OlmoEarth Studio for downstream analysis Back to Articles a]:hidden"> Introducing OlmoEarth embeddings: Custom embedding exports from OlmoEarth Studio for downstream analysis Enterprise Article Published August 12, 2026 Upvote 1 Kyle Wiggers Ai2Comms allenai 📄 Tech Report: https://allenai.org/papers/olmoearth | 📊… 9 Hacker News — AI on Front Page community 17d ago Show HN: Woxi - Open-source Mathematica / Wolfram Language reimplementation Woxi is an interpreter for the Wolfram Language written in Rust. It comes with Woxi Studio, a Mathematica-like GUI built with iced, but you can also use Woxi through a CLI, Jupyter kernel, Python package, npm package, or WASM module. Compared with wolframscript / Mathematica,… 28 OpenAI Python SDK releases dev-tools 18d ago v2.54.0 2.54.0 (2026-08-11) Features api: Add new Responses model identifiers ( #3595 ) ( 0652787 ) Bug Fixes api: clarify audio upload metadata requirements ( #3596 ) ( 28888f9 ) Chores api: Update generated-file header attribution to Castiron ( #3583 ) ( ea17fda ) 23 r/LocalLLaMA community 18d ago Introducing Unsloth Desktop app Hi LocalLlama, we're super excited to release Unsloth Desktop today! 🦥 It's the first desktop app that enables you to run and train models locally. Open-source. Available on Mac, Windows, and Linux Supports MLX, diffusion image/video models, audio models, and GGUF You can run… 19 r/LocalLLaMA community 19d ago Why have 8B-12B models been dropped? I am a Macbook Pro M4 user with the 16GB of unified ram. The best model I have been able to run on LM Studio is Gemma4 12B QAT, this model is 66 days old. After that the next best thing LM studio suggests is Nemotron 3 Nano 4B and Qwen3.5 9B, which both are 147-161 days old. It… 37 arXiv — NLP / Computation & Language research 19d ago Multilingual Emotion Neurons in Large Audio-Language Models arXiv:2608.08772v1 Announce Type: new Abstract: Emotion is central to human communication, and its expression varies across languages. Large audio-language models (LALMs) achieve strong performance on multilingual speech tasks, yet it remains unclear whether they encode emotion… 23 arXiv — NLP / Computation & Language research 19d ago REFRAMED: Towards Realistic Audio Description Generation for Movies arXiv:2608.09765v1 Announce Type: new Abstract: Audio Description (AD) is a verbal narration of key visual content in videos, enabling access for visually impaired audiences. Unlike standard video captioning, AD is a structured editorial task: descriptions must be inserted into… 35 arXiv — NLP / Computation & Language research 19d ago Structured Phonological Representations for Audio-Articulatory rtMRI Speech Classification arXiv:2608.09767v1 Announce Type: new Abstract: Real-time MRI makes it possible to observe vocal-tract articulation during speech, but mapping these articulatory patterns to phonetic and phonological categories remains challenging. We investigate whether PhonoQ, an audio-based… 35 arXiv — NLP / Computation & Language research 19d ago Comparing British and American Audio Description of Movies arXiv:2608.09792v1 Announce Type: new Abstract: Narrating the visual component of movies is known as audio description. It is a narrative technique designed to enable blind and visually impaired individuals to follow the story. However, it is far more constrained than most… 24 Simon Willison community 19d ago Introducing Muse Glimmer Introducing Muse Glimmer Meta are back in the open weights game! Muse Glimmer is a brand new 30B model under a clean Apache 2.0 license (a step up from the janky Llama licenses of old). Here's a pelican which I generated using LM Studio's 18.16 GB version of the model : I really… 21 Hacker News — AI on Front Page community 19d ago Humanising LLM Outputs Is Dumb Article URL: https://kuber.studio/blog/Reflections/Humanising-LLM-Outputs-is-Actually-Dumb Comments URL: https://news.ycombinator.com/item?id=49243474 Points: 200 # Comments: 131 5 r/LocalLLaMA community 20d ago Chat UIs with native audio input for multimodal models? I've been running Gemma 4 E4B with oMLX and I can't find any chat interfaces that directly send the audio file to the model instead of running the audio through a separate STT layer. I can confirm the audio layers work because I ran a couple of requests through Pydantic AI in… 6 r/LocalLLaMA community 20d ago MiniMax H3: A New Open-Weight Video Model, Live in ComfyUI MiniMax H3 is an open-weight, general-purpose multimodal video generation model that works across text, images, video, and audio. In ComfyUI, you can use H3 for text-to-video, image-to-video, first- and last-frame generation, and reference-driven creation. H3 jointly generates… 10 Hugging Face Daily Papers research 20d ago Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Abstract Audio reasoning is essential for machine understanding of the acoustic world. Reinforcement learning with verifiable rewards can elicit such reasoning, yet existing reward designs are complementary in their limitations: outcome-based rewards supervise only the final… 18 arXiv — NLP / Computation & Language research 20d ago Separating Decision-Rule Misalignment from Readout-Coverage Limitations in Speech Language Models arXiv:2608.06409v1 Announce Type: new Abstract: Speech language models are increasingly evaluated on paralinguistic tasks by the accuracy of prompted answers, but answer accuracy combines failures at different stages of the audio-to-answer computation. We introduce a… 4 arXiv — NLP / Computation & Language research 20d ago Do Audio Language Models Use Paralinguistic Evidence? Counterfactual Audits for Response Evaluation arXiv:2608.06718v1 Announce Type: new Abstract: Audio-language models (ALMs) are increasingly used as judges for speech-to-speech systems, but a judge that receives audio may not actually use paralinguistic evidence. We introduce counterfactual audits for paralinguistic response… 13 Hugging Face Daily Papers research 20d ago StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Abstract Deploying autonomous multimodal agents in continuous, real-world environments requires them to ingest unbounded audio-visual streams and maintain hour-scale memory. However, current evaluations predominantly rely on brief clips and multiple-choice formats. This design… 27 r/LocalLLaMA community 20d ago The Gemma team will host a special event on August 20 Tweet by u/hackerllama Could be copium, but I would love to see Gemma 4.1 there with unified audio input for all model sizes perhaps even up to 120B, much improved tool calling (even with the latest template there are still bugs ), higher precision QAT from the start and… 20 r/LocalLLaMA community 20d ago DeepSeek-V4-Flash-0731 Q8_K_XL sometimes stops mid-task in OpenCode - anyone else seeing this? Hey everyone, I've been experimenting with the new DeepSeek-V4-Flash-0731 release locally using the Unsloth Studio Q8_K_XL GGUF with OpenCode. Overall, it's been working really well, but I've noticed a strange behavior during longer agentic coding sessions. Once the context gets… 38 r/LocalLLaMA community 20d ago Open-weight video gen that actually delivers. Five days with MiniMax H3 on local hardware. H3 weights went live on HuggingFace August 3rd and I started pulling them immediately. An omni-modal video model with native stereo audio in the same forward pass, where audio can actually drive the video generation? On open weights? I had to try it. Five days in, the quality is… 31 Simon Willison community 22d ago Quoting John Gruber Me, I try to get into the mindset of playing live music, not recording a studio album. Except when I’m writing a piece where I really want it to be an album. Those aren’t rare , per se, but they’re occasional . If I tried to make every post a hall-of-famer I’d never get anything… 12 Simon Willison community 22d ago Quoting John Gruber Me, I try to get into the mindset of playing live music, not recording a studio album. Except when I’m writing a piece where I really want it to be an album. Those aren’t rare , per se, but they’re occasional . If I tried to make every post a hall-of-famer I’d never get anything… 15 llama.cpp releases dev-tools 22d ago b10326 tts: account for the vocoder pass in the timings line ( #26733 ) get_output runs the waveform work the pipeline defers to it, from a single trailing window to a full pass depending on the model. Measuring it keeps the reported total and the audio to process ratio honest.… 6 r/LocalLLaMA community 22d ago parakeet.wgsl – Fast, accurate ASR in the browser, via raw WebGPU & SIMD WASM High-performance inference of NVIDIA's Parakeet TDT 0.6B V2 English transcription model, in the browser. Check out the live demo: https://parakeet.narcotic.sh/ A fully custom, dependancy-free implementation with raw WebGPU compute shaders and SIMD WebAssembly audio frontend. 1… 19 Simon Willison community 22d ago The Tokenpocalypse Is Here: Companies Are Scrambling To Stop Spending So Much on AI The Tokenpocalypse Is Here: Companies Are Scrambling To Stop Spending So Much on AI There's a fun anecdote from Accenture (apparently via leaked meeting audio recordings) in this 404 Media piece from June 24th: “We’re seeing from some of the data internally at least that it’s… 5 Simon Willison community 22d ago The Tokenpocalypse Is Here: Companies Are Scrambling To Stop Spending So Much on AI The Tokenpocalypse Is Here: Companies Are Scrambling To Stop Spending So Much on AI There's a fun anecdote from Accenture (apparently via leaked meeting audio recordings) in this 404 Media piece from June 24th: “We’re seeing from some of the data internally at least that it’s… 4 Hugging Face Daily Papers research 22d ago Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval Abstract Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2.0 audio embeddings. Yet their weights do not map onto electrophysiological quantities,… 20 Ars Technica — AI news-outlet 23d ago Suno hopes to go legit with watermarks for AI-generated music Suno plans watermarks and download limits to stop "large-scale abuse." 32 r/LocalLLaMA community 23d ago nvidias nemotron omni only loads its text half on a mac, so i wrote the vision and audio towers in mlx nvidias nemotron omni is open weights and it sees, hears and reasons. theres already a 4bit mlx quant on hugging face but only the text backbone loads with standard mlx tooling. the model card says it plainly, the vision and audio towers need a runtime that implements the… 20 TechCrunch — AI news-outlet 23d ago Amid legal battles, Suno says it will start watermarking songs Suno's watermarking feature comes as the company is fighting legal battles on several fronts. 35 r/MachineLearning community 24d ago What are the biggest challenges in collecting high-quality speech and egocentric video datasets? [D] We're currently involved in collecting two types of datasets that seem to be increasingly important for multimodal AI Studio quality speech/audio datasets (high fidelity recordings) Egocentric household activity video datasets (first person daily task recordings) One thing that… 37 Hugging Face Daily Papers research 24d ago AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities Abstract While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. Existing benchmarks primarily evaluate visual transformations on… 15 Page 2 of 9 · 430 articles ← Newer Older →