News / #voice Tag Voice 500 articles archived under #voice · RSS Sign in to follow arXiv — NLP / Computation & Language research 1mo ago Best-of-$N$ TTS Evaluation is Confounded by ASR Family Alignment arXiv:2607.08256v1 Announce Type: new Abstract: Best-of-$N$ (BoN) inference improves content consistency in zero-shot text-to-speech by selecting from $N$ candidates with an automatic speech recognition (ASR) verifier. We identify an underexplored evaluation confound: a… 18 arXiv — NLP / Computation & Language research 1mo ago When Synthetic Speech Is All You Have: Better Call GRPO arXiv:2607.08409v1 Announce Type: new Abstract: LLM-based ASR adapted to regulated domains such as banking is bottlenecked by privacy: real speech is costly and legally constrained to collect, making synthetic text-to-speech (TTS) an attractive substitute. Yet synthetic speech… 6 arXiv — NLP / Computation & Language research 1mo ago UtterTune: LoRA-Based Target-Language Pronunciation Edit and Control in Multilingual Text-to-Speech arXiv:2508.09767v3 Announce Type: replace-cross Abstract: We propose UtterTune, a lightweight method for adapting a multilingual text-to-speech (TTS) system built on a large language model (LLM). It improves control of pronunciation in the target language while preserving… 25 Hugging Face Daily Papers research 1mo ago Vidu S1: A Real-Time Interactive Video Generation Model Abstract Vidu S1 is a real-time interactive video generation model that supports voice-controlled digital character animation with infinite-length output and high frame rate on consumer hardware. Generated by Qwen/Qwen2.5-Coder-32B-Instruct We introduce Vidu S1, a real-time… 15 TechCrunch — AI news-outlet 1mo ago Paris-based AI voice startup Gradium raises $100M seed, backed by Nvidia The Paris-based ElevenLabs competitor, just announced a hefty seed extension round. 11 arXiv — NLP / Computation & Language research 1mo ago Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts arXiv:2607.06611v1 Announce Type: new Abstract: Automatically recognizing the sentiment, positive or negative, from speech is a challenging task, requiring both the analysis of vocal inflections and the interpretation of uttered words. Recent solutions rely on audio foundation… 10 arXiv — NLP / Computation & Language research 1mo ago Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs arXiv:2607.06831v1 Announce Type: new Abstract: Speech-to-text alignment means finding the temporal boundaries of each word in the audio. Some models provide such an alignment directly and others do not. Connectionist temporal classification (CTC) and transducer models have an… 14 arXiv — NLP / Computation & Language research 1mo ago Transformer-based segmentation of prosodic boundaries in Brazilian Portuguese arXiv:2607.07408v1 Announce Type: new Abstract: Automatic prosodic segmentation identifies boundaries between speech units from acoustic and linguistic evidence. Although recent deep learning approaches have produced strong results for English, automatic segmentation for… 34 arXiv — NLP / Computation & Language research 1mo ago Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders arXiv:2607.07294v1 Announce Type: cross Abstract: Turn-taking prediction is a key requirement for social robots involved in human-human interaction, particularly in mediator settings, where the robot must anticipate conversational dynamics rather than merely react to pauses.… 23 arXiv — NLP / Computation & Language research 1mo ago Simulstream: Open-Source Toolkit for Evaluation and Demonstration of Streaming Speech-to-Text Translation Systems arXiv:2512.17648v2 Announce Type: replace Abstract: Streaming Speech-to-Text Translation (StreamST) requires producing translations concurrently with incoming speech under strict latency constraints, demanding models that balance low latency with high translation quality.… 7 r/LocalLLaMA community 1mo ago [audio.cpp] What Does the Fox Say: 4 ASR models (Nemotron 3.5 ASR, Higgs Audio STT, VibeVoice ASR, and Hviske ASR) in native C++/GGML, init streaming support, and 327s of audio transcribed in 2.17s. I just pushed a new audio.cpp update with streaming support and 4 ASR models: Nemotron 3.5 ASR, Higgs Audio STT, VibeVoice ASR, and Hviske ASR (da only). Overall 1.07x to 2.41x faster than Python. I decided to drop Parakeet-TDT since good implementations already exist, and I… 25 Simon Willison community 1mo ago Introducing GPT‑Live Introducing GPT‑Live OpenAI finally upgraded the model used by ChatGPT voice mode! I've had preview access for a few weeks in the iPhone app, and the new model is very impressive. It also has the ability to spin off harder tasks to GPT-5.5: For questions that require web search,… 13 Hugging Face Daily Papers research 1mo ago VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech Abstract Large Audio-Language Models exhibit systematic generative biases in realistic scenarios when evaluated through open-ended tasks using human-recorded speech, with bias magnitude varying significantly by task and triggered by gender and accent cues. Generated by… 21 TechCrunch — AI news-outlet 1mo ago OpenAI releases new voice models for more natural live conversations OpenAI says its new voice mode can speak and listen at the same time, a key ability for live translation. 23 arXiv — NLP / Computation & Language research 1mo ago Revisiting the Relation Between Language Model Perplexity and ASR Word Error Rate for Modern End-to-End Speech Recognition arXiv:2607.05612v1 Announce Type: new Abstract: Language model (LM) perplexity (PPL) has historically been used as a proxy for automatic speech recognition (ASR) word error rate (WER), with prior work reporting an approximately linear relation in log-log space. Modern end-to-end… 7 arXiv — NLP / Computation & Language research 1mo ago NAVER LABS System Re-implementation for the IWSLT 2026 Instruction-Following Task arXiv:2607.05623v1 Announce Type: new Abstract: We re-implement the NAVER LABS IWSLT 2025 instruction-following pipeline for the IWSLT 2026 Shared Task (constrained condition, short audio track), adapting it to the mandated components: SeamlessM4T-v2-large as the speech encoder… 27 arXiv — NLP / Computation & Language research 1mo ago Umm... With Transformers? Insights from Filled Pause Use across Four Slavic Parliaments arXiv:2607.05964v1 Announce Type: new Abstract: Filled pauses (FPs) are a universal feature of spontaneous speech, yet most studies rely on small, single-language corpora, limiting the generalisability of their findings. We analyse ~4,000 hours of parliamentary speech across… 27 arXiv — NLP / Computation & Language research 1mo ago From Sinhala to Dhivehi: Cross-Lingual Transfer Learning for Low-Resource Speech Recognition arXiv:2607.06289v1 Announce Type: new Abstract: Dhivehi, the national language of the Maldives, is currently under-resourced for automatic speech recognition (ASR) and other NLP tasks. This study investigates whether cross-lingual transfer learning from Sinhala, a linguistically… 29 arXiv — NLP / Computation & Language research 1mo ago Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs arXiv:2607.06540v1 Announce Type: new Abstract: Developing seamless, high-performance, native intelligent full-duplex Spoken Language Models (SLMs) remains a critical challenge and long-standing goal for the speech and NLP community. Despite notable progress, recent endeavors… 13 arXiv — NLP / Computation & Language research 1mo ago BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching Speech arXiv:2607.06054v1 Announce Type: cross Abstract: Off-the-shelf TTS systems are poorly adapted to Taiwanese Mandarin. Their accent defaults to other Mandarin variants, their tokenizers over-segment common Taiwanese text, and their pronunciation degrades at code-switching… 33 arXiv — NLP / Computation & Language research 1mo ago WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS arXiv:2607.06461v1 Announce Type: cross Abstract: While recent Large Language Model (LLM)-based Text-to-Speech (TTS) systems have achieved remarkable naturalness, they predominantly rely on implicit end-to-end generation paradigms, resulting in coarse-grained control. In… 18 arXiv — NLP / Computation & Language research 1mo ago Multi-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic SpeechLLMs arXiv:2601.12494v3 Announce Type: replace-cross Abstract: Audio large language models (LLMs) enable unified speech understanding and generation, but adapting them to linguistically complex and dialect-rich settings such as Arabic-English remains challenging. We present a… 24 Hugging Face Daily Papers research 1mo ago Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment Abstract A supervised contrastive alignment framework maps WavLM embeddings from English and Mandarin into a shared clinical space for depression detection, addressing cross-lingual generalization challenges and revealing performance artifacts caused by speaker identity leakage.… 38 Vercel — AI dev-tools 1mo ago Chat SDK adds Dial support Chat SDK now supports Dial with the new vendor-official adapter . Build bots that send and receive SMS, MMS, and iMessage on a real phone number, with bidirectional media and inbound voice-call transcripts. Replies use the standard Chat SDK thread and message APIs, with… 19 OpenAI official-blog 1mo ago Introducing GPT-Live A new generation of voice models for natural human-AI interaction, now powering ChatGPT Voice. 30 Hacker News — AI on Front Page community 1mo ago Local, CPU-Friendly, High-Quality TTS (Text-to-Speech) with Kokoro Article URL: https://ariya.io/2026/03/local-cpu-friendly-high-quality-tts-text-to-speech-with-kokoro/ Comments URL: https://news.ycombinator.com/item?id=48821576 Points: 372 # Comments: 75 7 r/LocalLLaMA community 1mo ago Gepard : 0.6B streaming TTS built for real-time dialogue - 20× realtime factor, ~50ms time-to-first-audio, vLLM-native, Apache 2.0 We just open-sourced Gepard 1.0 , a TTS model built for real-time conversation. It’s streaming-first: audio starts the moment text arrives, generated frame by frame instead of waiting for a full sentence. - ~555M params : Qwen3.5 0.8B backbone (14 layers) + Nemo NanoCodec (FSQ,… 36 Hugging Face Daily Papers research 1mo ago Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study Abstract Temporal aggregation methods for speech-based depression detection show inconsistent performance across different backbones and training runs, highlighting the need for robust benchmarking criteria. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Speech-based depression… 6 Hugging Face Daily Papers research 1mo ago Unified Audio Intelligence Without Regressing on Text Intelligence Abstract A unified audio-text large language model is presented that integrates audio and text processing through a shared transformer decoder, achieving superior performance across multiple audio and speech tasks while maintaining strong text reasoning capabilities. Generated… 33 r/LocalLLaMA community 1mo ago Are there any local ASR models that surpass Whisper right now? Hey everyone, I'm currently using faster-whisper(medium/large turbo) for local speech recognition, running it on an 8GB VRAM GPU. It works great, but I was wondering if there are any new open-source/local models that outright beat Whisper at this point? Here is exactly what I'm… 18 arXiv — Machine Learning research 1mo ago GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech arXiv:2607.02633v1 Announce Type: new Abstract: We present GRAFT, a per-word pronunciation conditioning mechanism for text-to-speech neural codec language modeling. Existing systems reach high intelligibility and naturalness but inherit the ambiguity of text and mispronounce… 19 arXiv — NLP / Computation & Language research 1mo ago Reinforcement Learning for Data-Efficient Code-Switched ASR arXiv:2607.02757v1 Announce Type: new Abstract: Audio-language models can be prompted for code-switched speech, but their decoding is not optimized for code-switching and often fails at language boundaries. We propose a practical reinforcement learning with verifiable rewards… 6 arXiv — NLP / Computation & Language research 1mo ago LuxSQA: Ask Me in Luxembourgish with TTS-Augmented Spoken Question Answering arXiv:2607.02763v1 Announce Type: new Abstract: Spoken Question Answering (SQA) remains largely focused on high-resource languages and carefully recorded speech, limiting the reach of speech-LLM methods in low-resource settings. This paper investigates whether text-to-speech… 20 arXiv — NLP / Computation & Language research 1mo ago Jointly Improving Dialect Identification and ASR in Indian Languages using Multimodal Feature Fusion arXiv:2607.02862v1 Announce Type: new Abstract: Automatic Speech Recognition (ASR) and Dialect Identification (DID) are crucial for Indian languages, many of which are low-resource and exhibit significant dialectal differences. Existing methods often optimize ASR or DID… 24 arXiv — NLP / Computation & Language research 1mo ago S-DiverSe: Spanish Diverse Speech arXiv:2607.03207v1 Announce Type: new Abstract: Automatic speech recognition (ASR) has advanced remarkably for standard speech, yet speech affected by neurological conditions remains a challenge. We present S-DiverSe (Spanish Diverse Speech), a corpus of 3.2 hours of in-the-wild… 20 arXiv — NLP / Computation & Language research 1mo ago Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization arXiv:2607.04064v1 Announce Type: new Abstract: Unsupervised syllabic tokenization aims to learn discrete syllabic tokens that capture latent linguistic content-related structure from raw speech. Recent syllabic tokenization methods employ teacher-student distillation of the… 19 arXiv — NLP / Computation & Language research 1mo ago Towards Digital Preservation of Efik: TTS for a Low-Resource African Language arXiv:2607.04515v1 Announce Type: new Abstract: Efik, a tonal language spoken by about 3 million second language speakers and 1.5 million native speakers in Southeastern Nigeria, remains underrepresented in speech synthesis research. We present the first documented end-to-end… 24 arXiv — NLP / Computation & Language research 1mo ago Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition arXiv:2607.04814v1 Announce Type: new Abstract: Extending automatic speech recognition (ASR) to low-resource African languages is constrained by the prohibitive demands of data collection at scale. A promising direction is to leverage linguistic relatedness to enhance… 28 arXiv — NLP / Computation & Language research 1mo ago DuplexChat: Constructing Speaker-Separated Full-Duplex Dialogue Speech at Scale for Spoken Dialogue Language Modeling arXiv:2607.04941v1 Announce Type: new Abstract: Full-duplex spoken dialogue models are trained on conversational speech in which each speaker is represented as a separate stream, but existing large-scale public speech corpora are mostly monaural, making them unsuited for SDLM… 33 arXiv — NLP / Computation & Language research 1mo ago RABBiT: Rapidly adaptive BOLD foundation model via brain-tuning for accurate zero-shot and few-shot prediction of speech-elicited responses in the brain arXiv:2607.05171v1 Announce Type: new Abstract: Language understanding in the brain is context-dependent, varying across experimental stimuli and individuals, which makes it difficult to build computational models that generalize across both. This calls for a foundation model of… 33 arXiv — NLP / Computation & Language research 1mo ago Unified Audio Intelligence Without Regressing on Text Intelligence arXiv:2607.05196v1 Announce Type: new Abstract: Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we introduce Nemotron-Labs-Audex-30B-A3B (Audex), a unified audio-text LLM built on Nemotron-Cascade-2-30B-A3B, a… 7 Hugging Face Daily Papers research 1mo ago Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization Abstract A speaker-disentangled syllabic tokenizer regresses perturbed student representations toward clean teacher targets to improve syllable boundary detection and speech language modeling performance. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Unsupervised syllabic… 8 r/MachineLearning community 1mo ago CPU TTS benchmark with UTMOS MOS scoring: Kokoro, Supertonic, Inflect-Nano, and Kyutai's new Pocket TTS [P] Sharing a CPU TTS benchmark with objective MOS scores in case it's useful for anyone evaluating small TTS models. Adding this because Kyutai's Pocket TTS is architecturally different from the others in the field and I hadn't seen a head-to-head with it yet. Models: Kokoro 82M… 33 r/LocalLLaMA community 1mo ago Kyutai's Pocket TTS clones a voice from 5 seconds of audio, on CPU, under MIT. Benchmarked against Kokoro, Supertonic, and Inflect-Nano for Eng. TTS Kyutai dropped Pocket TTS a bit ago and I've been sitting on it for a benchmark. Finally ran it head to head against the three CPU TTS models that have been getting attention (Kokoro 82M, Supertonic 3, Inflect-Nano-v1). 180 timed runs, 36 audio samples, objective MOS scores via… 8 r/LocalLLaMA community 1mo ago As promised, here is the GitHub link for my 100% local voice-to-voice assistant I've posted up earlier versions of this project before, promising a GitHub link, but never got around to pushing the code from my local system. Sorry all, I have a very busy life :P Anyway, without further ado: GitHub: https://github.com/igorbarshteyn/athena Athena is a fully… 38 r/LocalLLaMA community 1mo ago Gemma Avatar: Talk to Gemma 4-31B face to face This is a voice chat with Gemma 4 31B where you talk to a 3D avatar. It listens while you speak, answers with a voice and a face (the avatar is exposed to the LLM as function tools: set_mood, make_hand_gesture, make_facial_expression) and Gemma decides the expressions on its… 8 arXiv — Machine Learning research 1mo ago Many Voices, One Reward: Multi-Role Rubric Generation for LLM Judging and Reward Modeling arXiv:2607.01830v1 Announce Type: new Abstract: Reliable reward and preference signals are critical for evaluating and optimizing large language models on open-ended tasks. Rubric-based judges offer a transparent way to decompose such judgments into explicit evaluation criteria,… 34 arXiv — NLP / Computation & Language research 1mo ago SPARCLE: SPeaker-aware Aligned Representations via Contrastive Language Embeddings arXiv:2607.01238v1 Announce Type: new Abstract: Recent advances in speech synthesis have shifted from phoneme representations to direct grapheme modeling. While phonemes address the one-to-many mapping between text and acoustics, they rely on grapheme-to-phoneme (G2P) systems… 13 arXiv — NLP / Computation & Language research 1mo ago From Monolingual to Multilingual: Evaluating Mamba for ASR in South African Languages arXiv:2607.01502v1 Announce Type: new Abstract: Recent advances in automatic speech recognition (ASR) have explored different sequence models, including Conformer-based models and newer state space models such as Mamba. Although prior work has evaluated these architectures in… 37 arXiv — NLP / Computation & Language research 1mo ago Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving arXiv:2607.01733v1 Announce Type: new Abstract: Speech-LLM integration has shown promising results by leveraging extensive textual pretraining, yet its specific benefits for automatic speech recognition (ASR) remain unclear. We observe that as supervised ASR training data… 20 Page 6 of 10 · 500 articles ← Newer Older →