News / #voice Tag Voice 500 articles archived under #voice · RSS Sign in to follow r/LocalLLaMA community 13h ago Any current Voice2Voice AI model that runs locally that’s good? You guys remember sesame AI? With their really good AI voice model? Obviously ChatGPT has their voice model that’s also really good. Is there any smaller local variant that runs on like consumer grade gpu‘s (12,16 24gb?) I think NVidia released something but I didn’t really… 25 r/LocalLLaMA community 1d ago Breeze-TTS-2 initial impressions: genuinely 'frontier' TTS You can test it out on breezblue's playground or use it locally, its only ~7GB.   submitted by   /u/Gohab2001 [link]   [comments] 20 llama.cpp releases dev-tools 1d ago b10672 OpenVINO: Update OV to 2026.3.1, whisper.cpp support, Qwen3.5 on NPU, and new ops ( #27843 ) OpenVINO Backend: Fuse IM2COL + MatMul convolution into OpenVINO convolution ci:ggml-ov: Skip recurrent state rollback tests ci:ggml-ov: Skip recurrent state rollback tests Update… 33 r/LocalLLaMA community 1d ago TontaubeV1 - Open TTS model release for local long-form generation Hey everyone, I am the co-founder of Tontaube. My brother and I just released TontaubeV1, a 2.9B-parameter open-weight TTS model focused on expressive speech, long-form generation, and low-latency local inference. The model is primarily aimed at English and German and supports… 36 arXiv — Machine Learning research 2d ago Benchmarking_Fast_Domain_Adaptation_for_Unsupervised_Speech_Units arXiv:2608.26992v1 Announce Type: new Abstract: Representation learning has attracted great atten- tion and managed to reach good performances as a pretraining method for downstream tasks or as a first step towards unsu- pervised speech modeling. Yet, little is known about how… 4 arXiv — Machine Learning research 2d ago Soft Active Electromyography Interface for Machine Learning-Enabled Silent Speech Recognition arXiv:2608.27048v1 Announce Type: new Abstract: Silent speech recognition (SSR) provides an alternative communication pathway in the absence of audible speech. However, conventional approaches are limited by the need for constant facial attachment, privacy concerns, and unstable… 33 arXiv — Machine Learning research 2d ago Refusal Is Not Robustness: Auditing Confident Fabrication in Large Language Models on a Provably Uninformative Clinical Pain Speech Transcript arXiv:2608.26167v1 Announce Type: cross Abstract: Hallucination and abstention benchmarks rarely establish that a model could not have known the correct answer, making it difficult to distinguish appropriate abstention from an unsupported prediction. Seven large language models… 32 arXiv — NLP / Computation & Language research 2d ago Training-Time Explainability for Multilingual Hate Speech Detection: Aligning Model Reasoning with Human Rationales arXiv:2608.26125v1 Announce Type: new Abstract: Online hate against Muslim communities often appears in culturally coded, multilingual forms that evade conventional AI moderation. Such systems, though accurate, remain opaque and risk bias, over-censorship, or under-moderation,… 17 arXiv — NLP / Computation & Language research 2d ago Beyond Accuracy: A Qualitative Analysis of Vision-Language Models for Hate Speech Detection in Memes arXiv:2608.26143v1 Announce Type: new Abstract: Memes have turned out to be a powerful tool through which individuals share their ideas concerning contemporary social and political problems. Their anonymity, as well as their ability to go viral, make them a powerful medium for… 28 arXiv — NLP / Computation & Language research 2d ago Vagdhenu: A Vrutta (Meter) Aware Shloka-to-Chant (TTS) System for Sanskrit arXiv:2608.26146v1 Announce Type: new Abstract: We present Vagdhenu, a vrutta (meter) aware shloka-to-chant system for Sanskrit: a text-to-speech system that maps a metrical verse to its chanted parayana recitation at high fidelity. This is an experience report, not a new… 33 arXiv — NLP / Computation & Language research 2d ago Towards Interpretable Depression Detection: Linking Acoustic Features to DSM-5 Indicators arXiv:2608.26148v1 Announce Type: new Abstract: Depression affects millions worldwide, yet diagnosis relies on subjective self-reports that may miss authentic behavior. This paper presents an approach linking speech acoustics to DSM-5 depressive-behavior indicators through a… 26 arXiv — NLP / Computation & Language research 2d ago A Reranker for Orchestrating Heterogeneous Speech and Text Retrievers arXiv:2608.26194v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) systems have attracted significant interest for their ability to mitigate hallucinations in Large Language Models (LLMs). Although knowledge databases for RAG are increasingly diversifying to… 16 arXiv — NLP / Computation & Language research 2d ago AfriSwitch: A Benchmark for In-the-Wild African Code-Switched Speech Recognition arXiv:2608.26434v1 Announce Type: new Abstract: Code-switching is pervasive in bilingual African conversation, yet most ASR systems assume monolingual input and are evaluated on curated monolingual benchmarks. We present AfriSwitch, a 61.36-hour human-transcribed benchmark of… 22 arXiv — NLP / Computation & Language research 2d ago Scaling phoneme-based TTS augmentation for ASR: A unified pipeline and controlled study arXiv:2608.26697v1 Announce Type: new Abstract: Synthetic speech provides scalable supervision for automatic speech recognition (ASR), but its benefit depends on the selected texts, reference speech, and amount of synthesized data. We present a unified phoneme-based TTS-to-ASR… 14 arXiv — NLP / Computation & Language research 2d ago Mapping Written Words to Spoken Words in a Different Language Using Only Visual Grounding arXiv:2608.26925v1 Announce Type: new Abstract: In many low-resource settings, even just eliciting speech for data collection is difficult. One promising approach has been to ask speakers to describe images. But how do we build models from such visually grounded speech data?… 12 arXiv — NLP / Computation & Language research 2d ago Said Aloud, Read Different: Cross-Modal Instability in Multimodal Models arXiv:2608.27135v1 Announce Type: new Abstract: Multimodal foundation models are increasingly used in speech-first assistants that must interpret spoken queries and produce visually grounded decisions. Yet it remains unclear whether semantically equivalent queries yield… 32 Hugging Face official-blog 2d ago The Open ASR Leaderboard Adds Its First Global South Language Back to Articles a]:hidden"> The Open ASR Leaderboard Adds Its First Global South Language Published August 28, 2026 Update on GitHub Upvote 3 Eric Bezzam bezzam Shobhit Banga Shobhitbanga VoiceArena Manas Dhir manasdhir04 VoiceArena Bhaskar Singh bhaskarJT VoiceArena Manmeet… 34 Hugging Face Daily Papers research 2d ago LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale Abstract LibriBrain100 is a large-scale MEG speech dataset that demonstrates improved decoding through extensive within-subject recordings and multi-subject supervised fine-tuning of pre-trained models. Generated by thinkingmachines/Inkling-Small We introduce LibriBrain100, a… 12 Hugging Face Daily Papers research 2d ago Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans Abstract A real-time framework for online co-speech gesture generation uses a causal multimodal autoregressive model with streaming speech and motion history, supported by synthetic dialogue data and continual user-feedback adaptation. Generated by thinkingmachines/Inkling-Small… 13 Hugging Face Daily Papers research 3d ago VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction Abstract VoiceMem introduces a dual-brain streaming memory architecture for speech language models that improves retrieval accuracy, emotional personalization, and real-time efficiency. Generated by thinkingmachines/Inkling-Small Conversational systems, such as duplex speech… 15 arXiv — NLP / Computation & Language research 3d ago LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale arXiv:2608.25204v1 Announce Type: cross Abstract: We introduce LibriBrain100, a large-scale MEG dataset for speech decoding designed from the ground up for reproducible, standardised evaluation. LibriBrain100 more than doubles the size of the original LibriBrain release,… 15 arXiv — NLP / Computation & Language research 3d ago Leveraging Speech Acts for Low-Data and Cross-Domain Conversation Derailment Forecasting arXiv:2608.25359v1 Announce Type: new Abstract: Conversational derailment forecasting aims to predict when online discussions will escalate into hostility, enabling proactive moderation. Existing approaches often struggle in low-data settings and to generalize across domains.… 30 arXiv — NLP / Computation & Language research 3d ago Generative vs. Encoder Large Language Models for ASR Evaluation: A Comparative Study arXiv:2608.25574v1 Announce Type: new Abstract: Automatic Speech Recognition (ASR) is typically evaluated using Word Error Rate (WER), which poorly reflects semantic similarity. While embedding-based metrics correlate better with human judgments, the respective roles of encoder… 23 arXiv — NLP / Computation & Language research 3d ago From Specialization to Generalization: Instruction-tuned LLMs for Robust Harmful Content Mitigation arXiv:2608.25605v1 Announce Type: new Abstract: Large language models (LLMs) demonstrate impressive performance across a wide range of general NLP tasks; however, their effectiveness in sensitive domains, such as hate speech detection, remains less clear. Prior studies comparing… 25 arXiv — NLP / Computation & Language research 3d ago Lost but not erased: Finding traces of a forgotten language in neural speech models arXiv:2608.25976v1 Announce Type: new Abstract: International adoptees retain phonological traces of a birth language they can no longer speak or comprehend, a persistence typically attributed to a biologically-timed critical period. We asked whether it could instead reflect the… 4 arXiv — NLP / Computation & Language research 3d ago Fine-Tuning Whisper for Automatic Speech Recognition in Baniwa: A Preliminary Study arXiv:2608.26060v1 Announce Type: new Abstract: Automatic Speech Recognition (ASR) technologies have achieved remarkable performance in recent years through the use of large multilingual foundation models. However, most advances remain concentrated on high-resource languages,… 9 Ars Technica — AI news-outlet 3d ago Google announces Gemini 3.5 Transcribe for AI-powered speech-to-text The AI that powers Gboard's Rambler is coming to more Google products, including Chrome. 23 Google DeepMind official-blog 3d ago Intelligent transcription with Gemini 3.5 Transcribe Now you can get more intelligent speech-to-text transcription with Gemini 3.5 Transcribe. 21 VentureBeat — AI news-outlet 3d ago Orchestration is the new challenge for CX in the age of AI agents Presented by Tata Communications Enterprises are deploying AI agents, voice AI, and automation across messaging, voice, and digital channels faster than the architecture meant to support it. Most of that deployment has involved attaching conversational AI to legacy systems never… 5 TechCrunch — AI news-outlet 4d ago India’s Ringg gets backing from Peak XV as it pushes voice AI past the phone call Ringg has raised $10 million from Peak XV as a part of its Series A extension. 25 r/LocalLLaMA community 4d ago Granite Speech 5.0 Turbo CTC: Extremely Fast and Accurate Transcription   submitted by   /u/coder543 [link]   [comments] 4 Hugging Face Daily Papers research 5d ago EchoWM: Open and Enterable Omnimodal World Models Abstract EchoWM is an omnimodal world model that generates synchronized high-resolution video, sound, music, and speech while following continuous 6-DoF navigation trajectories across first- and third-person views. Generated by thinkingmachines/Inkling-Small We present EchoWM,… 17 Hugging Face Daily Papers research 5d ago Better Retrieval, Worse Robustness:How Multi-hop RAG Amplifies Upstream ASR Errors Abstract Retrieval-augmented generation extensions amplify automatic speech recognition errors in spoken multi-hop question answering, primarily through corrupted query entities. Generated by thinkingmachines/Inkling-Small Speech-based applications pass spoken queries through… 12 arXiv — NLP / Computation & Language research 5d ago Bulbul: A Dataset for Dialectal Arabic Speech Recognition arXiv:2608.21950v1 Announce Type: new Abstract: Arabic automatic speech recognition (ASR) faces unique challenges due to diglossia, extensive regional dialect variation, and limited speech resources. Existing speech datasets often focus on single dialects or large-scale… 5 arXiv — NLP / Computation & Language research 5d ago Whitewashing Hate, Smearing Harmless Content: Annotator-Style Rebuttal Attacks on LLM-Based Moderation arXiv:2608.22230v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for hate speech moderation, often within human--AI workflows in which reviewers provide feedback before a final decision. Such feedback introduces two manipulation directions:… 35 arXiv — NLP / Computation & Language research 5d ago WnW: Waxing-and-Waning KV Cache for Long-Form Speech LLMs arXiv:2608.22704v1 Announce Type: new Abstract: Long-form audio inputs make the KV cache the dominant memory cost of speech LLMs. Prefill-only KV compression methods permanently discard audio KV positions once evicted, with no pathway to recover them during decoding. We show… 11 arXiv — NLP / Computation & Language research 5d ago Better Retrieval, Worse Robustness:How Multi-hop RAG Amplifies Upstream ASR Errors arXiv:2608.22872v1 Announce Type: new Abstract: Speech-based applications pass spoken queries through automatic speech recognition (ASR) before any retrieval module, so ASR errors enter the pipeline as a fixed upstream constraint. We empirically test whether two extensions to… 33 arXiv — NLP / Computation & Language research 5d ago Do Spoken Language Models Hear Speech as They Read Text? Bridging Structural Gaps Between Speech and Text arXiv:2608.22908v1 Announce Type: new Abstract: Spoken Language Models (SLMs) generate textual responses directly from speech, offering an alternative to cascaded systems. Despite recent advances, existing SLMs still exhibit weaker instruction-following behavior and limited… 10 r/LocalLLaMA community 5d ago iPhone Local TTS EPUB Reading - Audiobookify I've been working on this project (been a developer for a few years) for a few months now, and it has been in active testing for ~2 months. Its an EPUB reader that also offers local offline TTS, so its not just TTS focused, its meant to be a good regular reading app as well. I'd… 38 r/LocalLLaMA community 6d ago Best AI Voice Cloning in 2026: How to Clone Your Voice With AI   submitted by   /u/NextgenAITrading [link]   [comments] 14 arXiv — Machine Learning research 6d ago Training DeepFilterNet with Accurate Room Acoustic Simulations Improves Single-Channel Speech Enhancement arXiv:2608.20971v1 Announce Type: cross Abstract: We investigate how the realism of synthetic room impulse response (RIR) datasets affects the training of DeepFilterNet3 for single-channel speech enhancement. We compare a DNS4 image-source-method (ISM) RIR dataset with a… 6 arXiv — NLP / Computation & Language research 6d ago Building and Evaluating a Synthetic Bengali Speech Resource for Telecom Customer Care arXiv:2608.20346v1 Announce Type: new Abstract: Speech systems used in customer-facing applications often require domain-specific language coverage. We present a synthetic Bengali speech dataset for telecom customer-care scenarios. The dataset contains 10,000 audio-text pairs,… 8 arXiv — NLP / Computation & Language research 6d ago Trilingual Topic Modeling of Sri Lankan Parliamentary Debates arXiv:2608.20365v1 Announce Type: new Abstract: Sri Lankan parliamentary debates (Hansards) constitute a trilingual corpus of speeches in Sinhala, Tamil, and English, including code-mixed content, yet remain inaccessible to standard NLP pipelines due to layout-complex PDFs,… 27 arXiv — NLP / Computation & Language research 6d ago Poly-InstructTTS: Learning In-the-Wild Expressive Speech Synthesis from Open-Ended Instructions arXiv:2608.20387v1 Announce Type: new Abstract: While recent text-to-speech (TTS) models achieve high naturalness, controlling fine-grained expression via natural-language instructions remains challenging. We introduce Poly- InstructTTS, which learns expressive speech from… 9 arXiv — NLP / Computation & Language research 6d ago Self-Supervised Speech Representations Track Spoken Language Convergence to Adult Models in Infants and Children Who Are Deaf/Hard-of-Hearing arXiv:2608.20396v1 Announce Type: new Abstract: Language development is characterized by a gradual convergence of children's speech toward adult patterns. Measuring this process has traditionally required detailed transcription and language-specific expertise, limiting… 4 arXiv — NLP / Computation & Language research 6d ago A Factorial Ablation of a Speech-to-SFT Pipeline: Differential Effects on Data Quality and Downstream Transfer arXiv:2608.20394v1 Announce Type: cross Abstract: Industry pipelines that turn speech into supervised fine-tuning (SFT) data via multi-stage refinement are increasingly adopted but, to our knowledge, have not been publicly ablated stage-by-stage, leaving each stage's marginal… 8 arXiv — NLP / Computation & Language research 6d ago TurboBias 2.0: Streaming Context-Biasing for Production-Efficient ASR Systems arXiv:2608.21343v1 Announce Type: cross Abstract: Contextualization is essential for production automatic speech recognition (ASR) systems, where user-provided phrases must be recognized accurately under strict latency constraints. Although many context-biasing methods improve… 13 r/LocalLLaMA community 6d ago Has anyone tried agent-lightning?   submitted by   /u/MrMrsPotts [link]   [comments] 14 r/LocalLLaMA community 7d ago Create tts voice from actual animal sound recording In short, I want to create voices for my real chickens that I'm creating generated videos of. Ultimately I would like to create voices to be used in a tts application that are based on their real "voice patterns", as though the voice was being made with their own vocal chords. I… 26 r/LocalLLaMA community 7d ago How to remove trendy speech from llms? For example: Instead of saying: "I created this new ID" It says: "I minted this new ID" Instead of: "This alternative path is available" It says: "this escape hatch is available" This speech is so nonsensical and annoying. Just. Speek. Literally ... OR NORMALLY. Where did LLMs… 19 Page 1 of 10 · 500 articles Older →