News / #voice Tag Voice 500 articles archived under #voice · RSS Sign in to follow r/LocalLLaMA community 26d ago nvidia/NVIDIA-NemotronLabs-VoiceChat-11B · Hugging Face (full duplex)   submitted by   /u/adefa [link]   [comments] 9 OpenAI official-blog 27d ago How we built a realtime system for responsive voice AI in six months GPT-Live enables continuous voice interaction with AI, using a turnless speech model and low-latency architecture for faster, more natural conversations. 16 arXiv — NLP / Computation & Language research 27d ago ZeroR@CHiPSAL 2026: Two-Stage Vision-Language Adaptation with Contrastive Learning for Nepali Meme Classification arXiv:2607.28637v1 Announce Type: new Abstract: This paper presents our system for the CHiPSAL 2026 shared task on multimodal hate speech and sentiment detection in Nepali memes. We address both subtasks: binary hate speech classification and three-class sentiment analysis. Our… 33 r/LocalLLaMA community 27d ago Parlor v2: best-effort fully local GPT-Live clone on an M3 Pro GPT-Live is so good that I use it almost every day. I've been wanting to replicate it since it was released. My first attempt was to fine-tune Gemma 4 12B to behave like a full-duplex model. Something like grafting a decision tick + speech head to the model. It failed after… 24 r/LocalLLaMA community 29d ago [audio.cpp] Release 0.5: DramaBox expressive TTS, Confucius4 cross-lingual voice transfer, plus 7 more models and ROCm/HIP audio.cpp 0.5 is out :) The most fun new model in 0.5 is DramaBox . It is closer to prompt-directed voice acting. DramaBox is built on the LTX-2.3 audio architecture, and prompts can control emotion, delivery, laughs, sighs, pauses, transitions, and speaker behavior. Example… 15 TechCrunch — AI news-outlet 29d ago Smallest.ai raises $13M to build ultra-fast voice AI that sounds genuinely human The startup is building voice models designed to make AI phone calls pass the Turing test. 9 Hugging Face Daily Papers research 1mo ago AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition Abstract On-device speech emotion recognition (SER) is critical for real-time applications, yet large self-supervised models that excel at SER are too costly for edge devices. Multi-teacher knowledge distillation can compress them into a lightweight student, but two challenges… 17 arXiv — NLP / Computation & Language research 1mo ago Digital Harf: A Clinically Integrated Multimodal AI System for Pervasive Arabic Speech and Language Therapy arXiv:2607.27212v1 Announce Type: cross Abstract: Children with Autism Spectrum Disorder in Arabic-speaking countries face compounded barriers to effective speech and language therapy: a shortage of qualified specialists, limited service reach beyond urban centers, and a… 26 ThursdAI news-outlet 1mo ago This Week in AI: Open Weights, Frontier Models, Sandbox Escapes, Voice & AI Detection From CoreWeave: Alex is back to cover a crazy end of July week, with Kimi K3, Opus 5, 3 open letters, one asking for pacing AI progress and 3 guests! Tune in 36 TechCrunch — AI news-outlet 1mo ago Friend, the lonely AI wearable, returns with a new voice and a much bigger price tag Friend, the AI wearable, can now talk to its users — for an enhanced price. 38 Hugging Face Daily Papers research 1mo ago Voice Memory for Agentic Speech Recognition Abstract We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory.md and decides per utterance whether to act on the hypothesis or abstain and keep the 1-best. Asynchronously, a… 12 arXiv — NLP / Computation & Language research 1mo ago Voice Memory for Agentic Speech Recognition arXiv:2607.26410v1 Announce Type: new Abstract: We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory.md and decides per utterance whether to act on the hypothesis or abstain and keep… 29 arXiv — NLP / Computation & Language research 1mo ago Latent-IM: Latent Interaction Management for Speech LLMs arXiv:2607.26928v1 Announce Type: new Abstract: Classical spoken dialogue systems often separated dialogue management from response realization: a policy selected the next dialogue action, and a generation component expressed that action. As dialogue systems shift toward LLMs,… 18 arXiv — NLP / Computation & Language research 1mo ago Dissecting Sensitivity to Training Language in Self-Supervised Speech Learning Using Neural Audio Codec Tokens arXiv:2607.26350v1 Announce Type: cross Abstract: Neural audio codecs (NACs) have become popular for obtaining speech representations as discrete tokens. Beyond compression, discrete tokens can be used to train self-supervised learning (SSL) models. Such models, referred to as… 6 arXiv — NLP / Computation & Language research 1mo ago A Cross-lingual Comparison of Human and Classification Model Entrainment Behavior in Code-switched Speech Settings arXiv:2607.25202v1 Announce Type: new Abstract: Conversational entrainment is well-studied in monolingual and written contexts, but remains underexplored in spoken code-switching (CSW). We present a novel cross-lingual analysis of entrainment in Mandarin-English, Hindi-English,… 32 arXiv — NLP / Computation & Language research 1mo ago Evaluation of forced alignment of code-mixed speech: the case of Hindi-English arXiv:2607.25581v1 Announce Type: new Abstract: Code-mixed speech poses unique challenges to forced alignment: expanded inventories, orthographic errors, and speaker variation. We evaluate forced alignment of Hindi-English code-mixed speech using the Montreal Forced Aligner. We… 35 arXiv — NLP / Computation & Language research 1mo ago MyMentorLLM: A psychotherapy GenAI environment with multimodal voice/text patients, trainees and experts for deliberate practice arXiv:2607.25667v1 Announce Type: new Abstract: Psychotherapists need repeated training and supervision by experts; however, scalability is problematic. Here we present MyMentorLLM, a multimodal voice- and text-based simulation environment for deliberate practice, used to… 20 arXiv — NLP / Computation & Language research 1mo ago SpeechLLM Meets Federated Learning for End-to-End ASR: English and Italian Case Studies arXiv:2607.25716v1 Announce Type: new Abstract: Federated learning (FL) enables privacy-preserving training of automatic speech recognition (ASR) systems across distributed data sources, yet its application to large-scale speech language models (SpeechLLMs) remains unexplored.… 18 arXiv — NLP / Computation & Language research 1mo ago Two Views, One Voice: Evidence-Grounded Conversational Music Recommendation arXiv:2607.24846v1 Announce Type: cross Abstract: Traditional conversational recommenders entangle retrieval and response generation within a single text interface, so exact entity cues fade as the dialogue's intent evolves, which compromises explanation credibility. We address… 23 r/LocalLLaMA community 1mo ago A.X-K2 released https://huggingface.co/skt/A.X-K2 https://huggingface.co/skt/A.X-K2-ALM https://huggingface.co/KRAFTON/A.X-K2-Raon-Speech-21B-A3B 688B-A33B + About South Korea's Soverign AI Foundation Model Project. South Korea's Soverign AI Foundation Model Project (This will not be official… 11 Vercel — AI dev-tools 1mo ago Grok Voice Think Fast 2.0 now available on AI Gateway Grok Voice Think Fast 2.0 from xAI is now available on AI Gateway. It is a speech-to-speech voice model that takes audio in and audio out, improving on the previous Grok Voice model in reasoning, transcription accuracy, and conversation. The model reasons in parallel with… 38 TechCrunch — AI news-outlet 1mo ago Fish Audio raises $50M seed to build AI voice models for creators and enterprises Since launching last year, the startup today has more than 8 million people using the open-source or hosted version of its models, and now generates annual recurring revenue of $21 million. 4 r/LocalLLaMA community 1mo ago microsoft/VibeVoice-ASR-BitNet VibeVoice-ASR-BitNet is a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs — no GPU required. Through heterogeneous quantization, the model is compressed from 4.62 GB to 1.58 GB while achieving 1.6–2.3× faster inference than Whisper.cpp with… 17 arXiv — Machine Learning research 1mo ago Distribution-Specific Curvature Control with Finite-Sample Guarantees for Open-Weight Safety arXiv:2607.22929v1 Announce Type: new Abstract: A short fine-tuning run can undo the safety guards of an open-weight model---retraining a refusal-trained assistant to aid weapons development or produce hate speech. Preventing such harmful fine-tuning while retaining benign… 30 arXiv — NLP / Computation & Language research 1mo ago Simple Language Normalization Wins: Cross-Lingual Speaker Verification for the TidyVoice 2026 Challenge arXiv:2607.22923v1 Announce Type: new Abstract: Cross-lingual mismatch remains a key source of overall degradation in modern speaker verification. The TidyVoice2026 Challenge targets this setting with text-independent verification, comprising 3,666 training and 808 development… 29 arXiv — NLP / Computation & Language research 1mo ago Speech Signals Complement LLMs for Predicting Interpersonal Attraction in Speed Dating arXiv:2607.23037v1 Announce Type: new Abstract: Large language models (LLMs) can predict interpersonal attraction from conversation transcripts, but it remains unclear what a speech predictor can add beyond transcript-only LLM prediction. Using Japanese speed-dating… 19 arXiv — NLP / Computation & Language research 1mo ago Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages arXiv:2607.23808v1 Announce Type: new Abstract: In this work, we introduce Indic DiarBench, a speaker diarization and ASR benchmark dataset spanning all 22 scheduled languages of India. This corpus comprises approximately 108 hours of natural multi-speaker audio from near-field… 5 arXiv — NLP / Computation & Language research 1mo ago Earnings25: A Comprehensive 500-Hour Speech Benchmark for Finance arXiv:2607.23813v1 Announce Type: new Abstract: We introduce Earnings25, a finance-domain benchmark for evaluating automatic speech recognition (ASR) on English-language earnings calls under realistic conditions. Earnings25 comprises two complementary test sets: (i)… 12 arXiv — NLP / Computation & Language research 1mo ago MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition arXiv:2607.24030v1 Announce Type: new Abstract: Massively multilingual automatic speech recognition (ASR) models covering hundreds of languages must maintain robust performance across diverse linguistic and acoustic conditions. However, these models often encounter the curse of… 18 arXiv — NLP / Computation & Language research 1mo ago Looking for Affect in Spontaneous Finnish Speech through Linguistic Interpretability arXiv:2607.24155v1 Announce Type: new Abstract: Existing research on affect in speech has shown how acoustic surface characteristics and content-related linguistic aspects of speech both relate to perceived emotional arousal and valence. However, it is not clear what the… 14 r/LocalLLaMA community 1mo ago You can now fine-tune my 3.96M-parameter TTS on your own voice or language When I released Inflect v2 last week, I thought most people would ask whether a TTS model this small actually sounded decent. Instead, I kept getting two questions: “Can I train it on my own voice?” “Can I move it to another language?” At the time, my answer was basically:… 16 Hugging Face Daily Papers research 1mo ago Multimodal Speaker Verification as a Threat to Speaker Anonymization Abstract Most automatic speaker verification (ASV) systems operate on individual utterances, despite real-world interactions typically consisting of multiple utterances. As speech accumulates, increasingly rich speaker information becomes available through acoustic, prosodic,… 27 arXiv — Machine Learning research 1mo ago Synthetic Speech, Real Signal: Paralinguistic Preservation and Cross-Lingual Augmentation via Voice Cloning arXiv:2607.22304v1 Announce Type: new Abstract: Synthetic data augmentation in speech is common practice for linguistic tasks like ASR, but has seen far less work for paralinguistic ones, especially clinical tasks where labelled data is expensive and some patient groups are… 35 arXiv — Machine Learning research 1mo ago Probing Speaker Identity Sensitivity in Audio Deepfake Detectors arXiv:2607.21820v1 Announce Type: cross Abstract: Audio deepfake detectors are trained to distinguish genuine speech from synthetic speech and often perform well on standard benchmarks. Yet the same detector that achieves less than 1% error on one dataset can see its error rate… 31 arXiv — NLP / Computation & Language research 1mo ago MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond arXiv:2607.22100v1 Announce Type: new Abstract: Lightweight projectors are an established way to connect pre-trained speech encoders with large language models (LLMs), mapping acoustic features into token-level embeddings for tasks like ASR and spoken question answering.… 24 arXiv — NLP / Computation & Language research 1mo ago Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision arXiv:2603.07025v2 Announce Type: replace Abstract: Speech Large Language Models (LLMs) that understand and follow instructions in many languages are useful for real-world interaction, but are difficult to train with supervised fine-tuning, requiring large, task-specific speech… 17 r/LocalLLaMA community 1mo ago ~20s that you'll never get back There is no quality or value to this post, however, I hope you might find humor in this broken output from my local Qwen TTS setup. Evidently I messed something up.   submitted by   /u/Full_Dimension_3495 [link]   [comments] 34 r/LocalLLaMA community 1mo ago ai-sage/GigaChat3.1-Audio-10B-A1.8B · Hugging Face GigaChat Audio 10B is an audio-native LLM built on top of the GigaChat 3.1 Lightning text model. A Conformer speech encoder and a modality adapter feed audio embeddings directly into a Mixture-of-Experts decoder, so the model keeps the text quality of its base while adding… 36 r/LocalLLaMA community 1mo ago I released Inflect v2: two ultra-tiny complete TTS models under 4M and 10M parameters I’ve spent the past month trying to find the point where an extremely small TTS model stops feeling like a size experiment and starts feeling genuinely useful. Today I’m releasing Inflect v2 , with two complete local text-to-speech models: Inflect-Nano-v2: 3.96M parameters,… 9 Ars Technica — AI news-outlet 1mo ago Canadian legislator reads out apparent LLM response in floor speech "Here’s a more natural, flowing version of that section..." 17 TechCrunch — AI news-outlet 1mo ago OpenAI’s new voice mode makes it to the ChatGPT desktop app ChatGPT Voice on desktop can work with both ChatGPT Work and Codex to complete tasks and control agents. 27 arXiv — NLP / Computation & Language research 1mo ago Instruct-FD: Can Your Full-Duplex Speech System Follow Turn-Taking Instructions? arXiv:2607.20460v1 Announce Type: new Abstract: Current full-duplex (FD) spoken dialogue systems can produce fluid interactions, yet it remains unclear whether they can adapt their turn-taking behavior when explicitly instructed. This is critical for real-world deployment, where… 17 arXiv — NLP / Computation & Language research 1mo ago DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages arXiv:2607.21540v1 Announce Type: new Abstract: We present DONDO, a family of open, permissively licensed automatic speech recognition (ASR) base models for African languages, built on the w2v-BERT 2.0 self-supervised speech encoder. DONDO comprises twenty-one monolingual models… 16 r/LocalLLaMA community 1mo ago [audio.cpp] Release 0.4: Higgs Audio v3 TTS 4B (10x real time)+ Fish Audio S2 Pro in C++/GGML, full GGUF loading, Q8 speed and VRAM gains audio.cpp again :) Release 0.4 is out. The headline this time is new high-quality TTS coverage plus GGUF becoming a first-class across the project. What’s new: Added Higgs Audio v3 TTS 4B, Fish Audio S2 Pro, Voxtral Realtime ASR and two community models OuteTTS TTS and… 7 TechCrunch — AI news-outlet 1mo ago Anthropic updates Claude voice mode with more capable models Claude's new voice model will let you reschedule your meeting or draft an email. 30 arXiv — NLP / Computation & Language research 1mo ago Abstraction Induces the Brain Alignment of Language and Speech Models arXiv:2602.04081v2 Announce Type: replace Abstract: Research has repeatedly demonstrated that intermediate hidden states extracted from large language models and speech audio models predict measured brain response to natural language stimuli. Yet, very little is known about the… 13 arXiv — NLP / Computation & Language research 1mo ago Simultaneous Speech-to-Speech Translation Without Aligned Data arXiv:2602.11072v2 Announce Type: replace Abstract: Simultaneous speech translation requires translating source speech into a target language in real-time while handling non-monotonic word dependencies. Traditional approaches rely on supervised training with word-level aligned… 25 arXiv — NLP / Computation & Language research 1mo ago The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation arXiv:2604.26347v2 Announce Type: replace-cross Abstract: Objective metrics for emotional expressiveness are vital for speech generation, particularly in expressive synthesis and voice conversion requiring emotional prosody transfer. To quantify this, the field widely relies on… 20 Interconnects research 1mo ago Open models recap: more on Kimi K3, Qwen 3.8, Xi's WAIC speech, distillation, the open-closed gap, and what's next A podcast with Florian Brand. 21 r/LocalLLaMA community 1mo ago We built NeuTTS-2E, an open-source on-device TTS model with 7 controllable emotions We’re open sourcing an alpha release of NeuTTS-2E : an on-device TTS model with 125M active parameters and 7 controllable emotions. The goal was simple: when you select “angry,” “fearful,” or “happy,” the delivery should follow that instruction rather than whatever emotion the… 14 Page 4 of 10 · 500 articles ← Newer Older →