News / #music Tag Music 430 articles archived under #music · RSS Sign in to follow arXiv — NLP / Computation & Language research 24d ago Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models arXiv:2608.05126v1 Announce Type: new Abstract: Spoken Language Understanding (SLU) is the core component of task-oriented dialogue systems and a pivotal link in achieving seamless human-agent interaction. While traditional SLU can effectively extract user semantics for… 7 r/LocalLLaMA community 24d ago Scenema Audio Comes to ComfyUI, Runs on 8GB VRAM Hey everyone! Scenema Audio is now a native ComfyUI custom node. Same model that powers scenema.ai now quantized so it fits on 8GB VRAM. When we first released it a few months ago as an API and Docker stack, the full precision transformers were too heavy for most people to… 6 Hugging Face Daily Papers research 24d ago Multi-Task Multi-Frame Visual Piano Transcription Abstract Audio-based piano transcription performs well on onset, pitch, and velocity, but the sustain pedal lets sound persist long after key release, so audio systems predict pedal-extended offsets rather than physical key release. Yet existing Visual Piano Transcription (VPT)… 38 r/LocalLLaMA community 25d ago Building a Fully Local PDF Read-Aloud & PDF-to-Audiobook Desktop App with Kokoro 82M, Qwen, and llama.cpp Hey everyone, I’ve been building Speechfony - a desktop app for reading PDFs (and EPUBs) with offline text-to-speech. Open a document, listen sentence-by-sentence with highlighting, or export selected pages to an MP3. Everything runs locally: Kokoro for speech, and an on-device… 5 arXiv — NLP / Computation & Language research 25d ago string2string Studio: An Interactive, In-Browser Platform for String-to-String Algorithms arXiv:2608.03984v1 Announce Type: new Abstract: We present string2string Studio, an interactive in-browser platform for string-to-string analysis across natural language processing, computational biology, and the digital humanities. The system integrates six main modules… 15 arXiv — NLP / Computation & Language research 25d ago Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv:2608.02831v1 Announce Type: cross Abstract: Audio reasoning is essential for machine understanding of the acoustic world. Reinforcement learning with verifiable rewards can elicit such reasoning, yet existing reward designs are complementary in their limitations:… 14 Hugging Face Daily Papers research 25d ago OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Abstract Omni-modal large language models (Omni-LLMs) have achieved remarkable performance on audio-visual understanding tasks, but processing long and highly redundant visual and audio token sequences incurs substantial computational overhead, demanding aggressive token… 6 Hugging Face Daily Papers research 25d ago AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling Abstract Language remains an outlier in generative modeling: while images, video, and audio are increasingly modeled in continuous latent spaces, text generation still relies predominantly on discrete tokens. Existing continuous language models either inherit embedding spaces… 32 llama.cpp releases dev-tools 25d ago b10274 mtmd: correcting duplicate empty audio chunks for short inputs ( #26536 ) correcting duplicate empty audio chunks for short inputs tests.sh code restored Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED… 30 TechCrunch — AI news-outlet 25d ago Meet Wrinkles, an app that uncovers the hidden stories of the places around you Wrinkles, available on both iOS and Android, essentially acts as an AI-powered audio tour guide that reveals hidden history and local stories. 19 Simon Willison community 25d ago PipeNetwork/minimax-h3-mlx PipeNetwork/minimax-h3-mlx MiniMax released MiniMax-H3 two days ago - they describe it as a "a general-purpose, omni-modal generative system", which in practice means it accepts text, images, audio and video and can use them to generate up to 15 second video clips with audio… 28 Simon Willison community 25d ago PipeNetwork/minimax-h3-mlx PipeNetwork/minimax-h3-mlx MiniMax released MiniMax-H3 two days ago - they describe it as a "a general-purpose, omni-modal generative system", which in practice means it accepts text, images, audio and video and can use them to generate up to 15 second video clips with audio… 26 r/LocalLLaMA community 25d ago Why are Chinese models better* at Frontend than the western top labs? I use A LOT both openAI and Anthropic products. When I need some frontend work (pure web dev) (or answer that feel less verbose and more to the point) I use Anthropic. For multimodality openAI feels better (understanding audio, screenshots, generating images, etc). But openAI… 17 r/LocalLLaMA community 26d ago Time to finally migrate from LM Studio -> llama.cpp, your experience? Has anyone moved from LM Studio to llama.cpp? What was your experience like? What did you have to learn in order to recreate your experience? Which harness/GUI did you switch to? Thanks in advance!   submitted by   /u/CSEliot [link]   [comments] 19 r/LocalLLaMA community 26d ago Is LM Studio abandoning their core product? Some of you may be aware that a few weeks ago, LM Studio announced a new agent, Bionic. This is pretty much an agentic harness for both local models and paid cloud models. But most aren't aware that LM Studio replaced almost every link to the original app that built their brand… 7 arXiv — Machine Learning research 26d ago Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval arXiv:2608.01481v1 Announce Type: new Abstract: Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2.0 audio embeddings. Yet their weights do not map… 38 arXiv — NLP / Computation & Language research 26d ago Discriminative Axis, Not Data Volume: What a Contrastive Corpus Teaches an Audio Embedding arXiv:2608.01560v1 Announce Type: new Abstract: Scaling the corpus is the default remedy when a contrastive representation lacks an attribute. We report a case where it does nothing, and identify what does: adding a lexical-speech round to a frozen-base multimodal embedding… 29 Hugging Face Daily Papers research 26d ago SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Abstract Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices without reference recordings, control speaker styles with natural… 4 Hacker News — AI on Front Page community 26d ago MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video Article URL: https://blog.comfy.org/p/minimax-h3-day-0-support-in-comfyui Comments URL: https://news.ycombinator.com/item?id=49155629 Points: 202 # Comments: 58 18 arXiv — NLP / Computation & Language research 27d ago TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models arXiv:2607.28896v1 Announce Type: cross Abstract: Unified audio models capable of audio understanding, audio generation and, increasingly, audio editing are proliferating rapidly. Yet a basic question about them remains unanswered: do the two heads of a unified model agree about… 26 r/LocalLLaMA community 27d ago MiniMax-H3 now on huggingface MiniMax H3 is a general-purpose, omni-modal generative system. It supports unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds. Thanks… 16 r/LocalLLaMA community 27d ago DeepSeek V4 @ IQ3XXS on M1 Ultra 128GB- 16 tok/s in LM Studio after patch M1 Ultra 128GB, Unsloth UD-IQ3_XXS, wired limit at 120GB. I was at 5-6 tok/s before the patch. Getting 15-16 tok/s now with the patched engine, and the output seems to have improved. Big thanks to this guy.   submitted by   /u/mil_phickelson [link]   [comments] 9 r/LocalLLaMA community 28d ago DeepSeek V4 Flash 0731 IQ2_M benchmark for Dual 3060 and 96GB RAM ≈ 3.5 tok/s. Thanks to the community help I finally launched this llm. LM Studio refused to load weight onto second GPU but Unsloth Studio did so everything was done in there. Not a proper benchmark (used PC in parallel as well) but it gives an idea of the performance from dual 3060 with… 16 r/LocalLLaMA community 28d ago EU AI Act takes effect tomorrow, August 2, 2026. 🤡 Basically you now have to mark all AI generated images, audio, video and text as AI generated. :P   submitted by   /u/xoxaxo [link]   [comments] 32 r/LocalLLaMA community 29d ago Deepseek V4 Flash 0731. LM Studio loading only into RAM. The model refuses to load into VRAM and uses only RAM. What can be an issue? Q2_K_XL from Unsloth if that changes something.   submitted by   /u/esw123 [link]   [comments] 29 r/LocalLLaMA community 29d ago [audio.cpp] Release 0.5: DramaBox expressive TTS, Confucius4 cross-lingual voice transfer, plus 7 more models and ROCm/HIP audio.cpp 0.5 is out :) The most fun new model in 0.5 is DramaBox . It is closer to prompt-directed voice acting. DramaBox is built on the LTX-2.3 audio architecture, and prompts can control emotion, delivery, laughs, sighs, pauses, transitions, and speaker behavior. Example… 15 Hugging Face Daily Papers research 29d ago OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Abstract Existing token compression methods for omnimodal large language models typically rely on one modality to determine what to retain in the other. We show that this assumption often breaks down: for the same query, audio and video relevance often peaks at different… 27 Hacker News — AI on Front Page community 29d ago Getting 25 Gbps Thunderbolt Ethernet on My Mac Studio Article URL: https://www.jeffgeerling.com/blog/2026/getting-25g-ethernet-mac-thunderbolt/ Comments URL: https://news.ycombinator.com/item?id=49125034 Points: 200 # Comments: 100 24 r/LocalLLaMA community 1mo ago Minimax-H3 video model released, open weights coming in the next few days https://x.com/MiniMax_AI/status/2083006198828417501?s=20 Quote from their article: Today, we're launching MiniMax H3, a general-purpose multimodal generation model. H3 understands unified context across text, images, video, and audio, generating video with native stereo sound,… 12 arXiv — Machine Learning research 1mo ago Journey Operators for Structured Multi-Axis Composition arXiv:2607.26775v1 Announce Type: new Abstract: Many kinds of data have structure along one or more axes: words in a sentence, pixels in an image, nodes in a tree, frames in audio, or cells in a 3D volume. Along one axis, order matters: "the dog bit the man" is different from… 10 arXiv — NLP / Computation & Language research 1mo ago Dissecting Sensitivity to Training Language in Self-Supervised Speech Learning Using Neural Audio Codec Tokens arXiv:2607.26350v1 Announce Type: cross Abstract: Neural audio codecs (NACs) have become popular for obtaining speech representations as discrete tokens. Beyond compression, discrete tokens can be used to train self-supervised learning (SSL) models. Such models, referred to as… 6 arXiv — NLP / Computation & Language research 1mo ago Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis arXiv:2607.26541v1 Announce Type: cross Abstract: Audio-capable foundation models enable end-to-end spoken interaction, but they also introduce safety risks beyond transcript content. It remains unclear how much jailbreak capability can arise from matched-text variation in… 29 Vercel — AI dev-tools 1mo ago Inkling Small from Thinking Machines is now available on AI Gateway Inkling Small from Thinking Machines is now available on AI Gateway. Inkling Small reaches performance comparable to the larger Inkling model at about a quarter of the size, using much less compute per task. It is a broad generalist with native reasoning over audio and images,… 30 Hugging Face Daily Papers research 1mo ago OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs Abstract Emerging Omni-modal Large Language Models (OmniLLMs) enable unified understanding of text, audio, and video, but their long audio-video token sequences introduce substantial memory and inference costs. Existing compression methods mainly focus on selecting important… 7 Vercel — AI dev-tools 1mo ago Grok Voice Think Fast 2.0 now available on AI Gateway Grok Voice Think Fast 2.0 from xAI is now available on AI Gateway. It is a speech-to-speech voice model that takes audio in and audio out, improving on the previous Grok Voice model in reasoning, transcription accuracy, and conversation. The model reasons in parallel with… 38 TechCrunch — AI news-outlet 1mo ago Fish Audio raises $50M seed to build AI voice models for creators and enterprises Since launching last year, the startup today has more than 8 million people using the open-source or hosted version of its models, and now generates annual recurring revenue of $21 million. 4 arXiv — NLP / Computation & Language research 1mo ago The JEPA Paradox in Language: The Geometry of Linguistic Alternatives arXiv:2607.23531v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs) are effective for images, video, and audio, yet deterministic JEPA-style latent prediction has not become a standard objective for text encoders. We argue that this gap reflects a… 37 arXiv — NLP / Computation & Language research 1mo ago Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages arXiv:2607.23808v1 Announce Type: new Abstract: In this work, we introduce Indic DiarBench, a speaker diarization and ASR benchmark dataset spanning all 22 scheduled languages of India. This corpus comprises approximately 108 hours of natural multi-speaker audio from near-field… 5 arXiv — NLP / Computation & Language research 1mo ago Open Your Model's Eyes: Video and Context-Aware Multimodal Backchannel Prediction arXiv:2607.22729v1 Announce Type: cross Abstract: Backchannels, which signal listener states like empathy and understanding, are fundamental to natural human interaction. However, current approaches rely solely on audio and text. This omits crucial visual cues, such as facial… 18 Hugging Face Daily Papers research 1mo ago JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents Abstract Creative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent generative models can synthesize high-quality images, videos, audio clips, UI elements, storyboards, slides, and other creative assets, real-world creative… 4 Hugging Face Daily Papers research 1mo ago OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Abstract Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging… 12 r/LocalLLaMA community 1mo ago Viable ways to run K3 locally just curious how would people run it cheap if they really want kimi k3. dgx spark / strix halo clusters optane persistent memory platform + some gpus mac studio clusters orange pi 6 clusters ssd streaming + gpus multiple ddr3 + connectx 5 rdma clients two dgx stations power 10… 37 arXiv — Machine Learning research 1mo ago Phylogenetic signal in marine mammal and bird vocalizations captured by audio foundation models: the limited benefit of domain-specific pretraining arXiv:2607.22458v1 Announce Type: new Abstract: Do learned audio embeddings encode structure that nobody told them to encode? We probe four large pretrained audio models (AST, CLAP, BEATs-bio and BirdNET) with a downstream task none of them saw during training: recovering… 17 arXiv — Machine Learning research 1mo ago Probing Speaker Identity Sensitivity in Audio Deepfake Detectors arXiv:2607.21820v1 Announce Type: cross Abstract: Audio deepfake detectors are trained to distinguish genuine speech from synthetic speech and often perform well on standard benchmarks. Yet the same detector that achieves less than 1% error on one dataset can see its error rate… 31 r/LocalLLaMA community 1mo ago ai-sage/GigaChat3.1-Audio-10B-A1.8B · Hugging Face GigaChat Audio 10B is an audio-native LLM built on top of the GigaChat 3.1 Lightning text model. A Conformer speech encoder and a modality adapter feed audio embeddings directly into a Mixture-of-Experts decoder, so the model keeps the text quality of its base while adding… 36 r/LocalLLaMA community 1mo ago Benchmarks: TensorSharp vs. llama.cpp Cuda and Vulkan Benchmark: TensorSharp vs. llama.cpp I would like to share my latest open source local Unsloth (GGUF) LLM inference engine and applications. It supports many models from Unsloth, like Gemma4, DiffusionGemma, Qwen3.6 with multi-modal (image, vision, audio), Qwen… 38 r/LocalLLaMA community 1mo ago My GX10 died Everything ran fine, I was using UD 3.6 Q6 for 35 and 27B, each 4 concurrent requests at 200K context. I had Dify and Mastra to play around with, Unsloth studio to get around to and vLLM ready for whenever I decided to do some more testing. LLama-swap above lama.cpp and liteLLM… 30 r/LocalLLaMA community 1mo ago OrangePi AI Studio Pro - Qwen3.5-122B-A10B https://preview.redd.it/wbq8ullnbafh1.png?width=1409&format=png&auto=webp&s=e6d2fe2b1c87c724bc64003c25f917dcee53260f I finally got round to tweaking this, with a bit of help from GLM5.2. The trick to getting it running with vLLM (which I couldn't get anything really out of… 24 r/MachineLearning community 1mo ago I built an open-source multi-agent SDLC harness that beats a cold Claude Code run on large repos, by learning the repo once. Real benchmarks (incl. where it loses) inside. [P] Built an open-source AI coding agent that was 7%–75% cheaper than a cold "claude -p" run on 6/6 well-localized tasks across repositories up to ~82k LOC. The biggest difference: Cold agent: $6.83, 207 turns AutoDev Studio: ~$1.70 for the same bug The full benchmark (including… 14 r/LocalLLaMA community 1mo ago FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence Introducing FLUX 3. One multi-modal model for Image, Video, Audio and Action-Prediction. Creations are truer to life in every kind of style. Blog Post : https://bfl.ai/blog/flux-3   submitted by   /u/pmttyji [link]   [comments] 23 Page 3 of 9 · 430 articles ← Newer Older →