News / #voice Tag Voice 500 articles archived under #voice · RSS Sign in to follow r/LocalLLaMA community 8d ago Ultrafast Qwen3-TTS at 34 ms Time-to-First-Audio, Handling 10 Requests Per Second [OSS] Hey locallama! We recently open sourced a Qwen3-TTS 1.7B implementation that achieves 10 requests per second (RPS) and sub-50 ms p95 time-to-first-audio (TTFA) while maintaining real-time playback on 1 x H100. This extends to 20 RPS at sub-100 ms p95 TTFA. By adjusting settings,… 31 Hugging Face Daily Papers research 9d ago Towards Quantifying Benchmark Optimization in ASR Models Abstract High-performing speech recognition models reproduce benchmark transcripts despite contradictory audio, revealing benchmark-optimized behaviors that inflate scores without improving real-world transcription. Generated by thinkingmachines/Inkling-Small Public benchmarks… 13 arXiv — Machine Learning research 9d ago Decoding silent reading from non-invasive EEG arXiv:2608.20186v1 Announce Type: new Abstract: Non-invasive decoding of inner speech faces a fundamental data problem: a corpus pairing brain activity with a person's spontaneous inner monologue cannot be collected, and the available proxy paradigms (cued repetitive and… 15 arXiv — NLP / Computation & Language research 9d ago Represented but Ignored: A Causal Account of Prosodic Underuse in Audio-Language Models arXiv:2608.19211v1 Announce Type: new Abstract: Human speech is richly expressive, with prosody carrying linguistic and emotional information beyond the lexical content. A capable large audio-language model (audio-LLM) should therefore support expressive speech understanding,… 4 arXiv — NLP / Computation & Language research 9d ago A Speech Corpus for Mizo Automatic Speech Recognition: Whisper and SraVaani 1.0 Fine-Tuning with Morphology-Aware Evaluation arXiv:2608.19361v1 Announce Type: new Abstract: This study reports the development of an Automatic Speech Recognition (ASR) system in Mizo, a low-resource language. The development included collecting 17.62 hours of speech data, curating it, and fine-tuning the Mizo ASR system… 4 r/LocalLLaMA community 9d ago Make Jensen Huang Sound Like Anyone. New Streaming Voice Conversion Model MeanVC2 Released! Finally see a new voice conversion model. MeanVC2 supports cross-gender and cross-language voice conversion. 3x realtime on CPU with audio.cpp. Disclaimer: The converted voice quality of MeanVC2 is decent; the noise comes from my rough demo engineering, not the model itself.… 38 Hugging Face official-blog 9d ago Measuring benchmark optimization in speech recognition Back to Articles a]:hidden"> Measuring benchmark optimization in speech recognition Published August 21, 2026 Update on GitHub Upvote 1 Theo Lebryk tlebryk02 HumeAI Eric Bezzam bezzam Alice aliceebaird HumeAI David Ayllon dayllon HumeAI Jakub Piotr Cłapa jpc HumeAI Jens Madsen… 22 arXiv — NLP / Computation & Language research 10d ago StocksTalk: A Voice-Enabled Conversational Agent for Structured Query Generation over Web Data arXiv:2608.18105v1 Announce Type: new Abstract: StocksTalk is a voice-enabled conversational system for transforming spoken financial screening requests into executable and validated structured queries over real-world market data. The system combines streaming speech… 9 arXiv — NLP / Computation & Language research 10d ago X2Streaming-TTS: Causal Token-Level Text-to-Speech from Streaming Text with Speech-State Inheritance arXiv:2608.18661v1 Announce Type: new Abstract: Streaming text-to-speech is essential for low-latency spoken dialogue systems, yet many systems wait for sentence-level text and are therefore only pseudo-streaming. True token-level synthesis must generate speech from uncertain… 30 arXiv — NLP / Computation & Language research 10d ago Understanding Multilingual Medical ASR Adaptation Through Layer-Wise Analysis arXiv:2608.18825v1 Announce Type: new Abstract: Medical automatic speech recognition (MedASR) requires adaptation to specialised terminology, limited annotated clinical data, and multilingual use cases. Although large-scale pretrained ASR models such as Whisper achieve strong… 19 arXiv — NLP / Computation & Language research 10d ago Introducing the Privacy-HSD Trade-off: Hate Speech Detection, but not at the Cost of Privacy arXiv:2608.19006v1 Announce Type: new Abstract: Hate speech is a real and timely threat that affects a large portion of online users, especially youth and minority groups. While building reliable and robust automatic hate speech detection (HSD) systems is paramount, we argue… 24 arXiv — NLP / Computation & Language research 10d ago Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs arXiv:2608.18131v1 Announce Type: cross Abstract: Current safety alignment training for Large Language Models (LLMs) are heavily English-centric. When such safety filters fail for non-English languages, the consequences are immediate and user-facing: voice assistants and spoken… 32 arXiv — NLP / Computation & Language research 10d ago Efficient Adaptation of LLMs for Hate Speech Detection in Low-Resource Languages: A Comparative Study on Roman Urdu arXiv:2608.18142v1 Announce Type: cross Abstract: It is challenging to detect hate speech in Low Resource Languages (LRLs) because of the absence of annotated data, the informality of its language structure, and the lack of standardized grammar. A good example of such a… 19 arXiv — NLP / Computation & Language research 11d ago Emotion Across Speech and Faces: Shared Affective Mechanisms in Multimodal Foundation Models arXiv:2608.17102v1 Announce Type: new Abstract: Modern multimodal foundation models (MFMs) have made rapid progress on tasks requiring integrated perception across speech, vision, and language, including emotion recognition. However, it remains unclear whether they recognize… 15 arXiv — NLP / Computation & Language research 11d ago SpeechSense: A Paralinguistic-Focused Dataset for Fine-Grained Speech Sentiment Analysis arXiv:2608.17931v1 Announce Type: new Abstract: Recent advances in AI have revolutionized speech processing, yet effective speech understanding requires discerning not just what is said, but how it is said. Speech Sentiment Analysis plays a critical role in decoding these… 36 arXiv — NLP / Computation & Language research 11d ago Language Family Matters: Evaluating LLM-Based ASR Across Linguistic Boundaries arXiv:2601.18899v3 Announce Type: replace Abstract: Large Language Model (LLM)-powered Automatic Speech Recognition (ASR) systems achieve strong performance with limited resources by linking a frozen speech encoder to a pretrained LLM via a lightweight connector. Prior work… 15 arXiv — NLP / Computation & Language research 11d ago Speak in Context: Multilingual ASR with Speech Context Alignment via Contrastive Learning arXiv:2603.06505v2 Announce Type: replace Abstract: Automatic speech recognition (ASR) has benefited from advances in pretrained speech and language models, yet most systems remain constrained to monolingual settings and short, isolated utterances. While recent efforts in… 17 Vercel — AI dev-tools 11d ago Fish Audio models now available on Vercel AI Gateway for free Fish Audio 's audio models are now available on AI Gateway. To celebrate the launch, every Fish Audio model is free on AI Gateway for the next 30 days, through September 18. Capability Regular Through September 18 Text-to-speech $15.00 per million characters Free Speech-to-text… 22 Hugging Face Daily Papers research 12d ago AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model Abstract AnyTalk generates 3D speech animations for arbitrary characters without animation data by adapting video diffusion models via character-specific fine-tuning and optimizing blendshape parameters from synthesized talking-head videos, with a distilled real-time variant.… 9 arXiv — NLP / Computation & Language research 12d ago Multi-Modal Generative Fuzzy System: Fuzzy Inference Guided Large Model Interactive Question Answering Framework arXiv:2608.14584v1 Announce Type: new Abstract: In Multimodal Question Answering (MQA), models are required to jointly encode and integrate heterogeneous information from multiple modalities, including text, images, and speech, to perform complex semantic reasoning and decision… 29 arXiv — NLP / Computation & Language research 12d ago Beyond the pale: Assessing prevalence and contents of extremist speech in LLM training data arXiv:2608.14813v1 Announce Type: new Abstract: Despite a strong interest on the part of the research community in the topic of trustworthy and safe AI, the composition of the text corpora that large language models (LLMs) encounter in pre- and post-training has not yet drawn… 26 arXiv — NLP / Computation & Language research 12d ago Semantic Space of Parts of Speech arXiv:2608.15443v1 Announce Type: new Abstract: Parts of speech categorization is understood in the European linguistic tradition as crisp categorization, which is also reflected in corpus linguistics, where each disambiguated token is assigned exactly one POS. However, the… 11 arXiv — NLP / Computation & Language research 12d ago The Null Token Knows: Reducing Message-Free Hallucination in ASR and NMT arXiv:2608.15940v1 Announce Type: new Abstract: Modern encoder-decoder systems can produce fluent text even when their input contains no recoverable message. We study this failure in ASR and NMT through the models' reserved null tokens, asking whether the score for ending… 22 arXiv — NLP / Computation & Language research 12d ago DuplexGen: Decoupling Content, Timing, and Acoustics for Synthetic Dialogue Speech arXiv:2608.16053v1 Announce Type: new Abstract: Synthetic conversational speech has become an important resource for developing and evaluating conversational speech systems. However, existing dialogue synthesis pipelines typically generate dialogue content first and then insert… 28 arXiv — NLP / Computation & Language research 12d ago Unadapted Multilingual ASR on a Garrusi Kurdish Evaluation Set: A Common-Reference Staged Normalization Analysis arXiv:2608.16379v1 Announce Type: new Abstract: Evaluating speech recognition for a Kurdish variety written in a Latin field orthography, using a model that outputs Arabic script, creates a measurement problem before a modelling one: direct scoring treats writing-system… 38 arXiv — Machine Learning research 13d ago VoiceDesigner: Text-to-Voice Generation and Editing via Unified Diffusion Modeling and Data Augmentation arXiv:2608.13613v1 Announce Type: cross Abstract: Recent breakthroughs in generative models have made text-to-voice generation (TTV) possible, enabling the synthesis of speech directly from textual voice descriptions. However, existing systems face two key challenges. First,… 21 arXiv — NLP / Computation & Language research 13d ago Measuring Fairness in Large Audio Language Models via Semantic-Aware Bias Estimation arXiv:2608.13624v1 Announce Type: new Abstract: Large Audio Language Models (LALMs) have seen increasing use for audio understanding tasks such as speech recognition and audio question answering, raising concerns about fairness across demographic subgroups. Fairness evaluation… 7 arXiv — NLP / Computation & Language research 13d ago StreamHear: Domain-Adapted Pseudo-Labeling for Semi-Supervised Streaming Speech Recognition arXiv:2608.13717v1 Announce Type: new Abstract: Streaming automatic speech recognition (ASR) underperforms on domain-shifted target audio, where labeled in-domain data is costly to prepare while unlabeled audio is abundant. We present StreamHear, a semi-supervised pipeline that… 19 arXiv — NLP / Computation & Language research 13d ago Leading-Silence Augmentation and Multi-Stage Synthetic Supervision for the Second MLC-SLM Challenge arXiv:2608.14150v1 Announce Type: new Abstract: The second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge evaluates two tasks over complete, unsegmented multilingual conversations: speaker diarization and recognition (Task 1) and conversational speech… 16 arXiv — NLP / Computation & Language research 13d ago VoiceChat-TTS: A Low-Latency Continuous Speech Synthesis Model for Interactive Agents arXiv:2608.13831v1 Announce Type: cross Abstract: Spoken dialogue is a natural form of human--computer interaction, yet most speech language models remain limited to turn-based operation and lack real-time adaptability, such as user barge-in. Recent duplex speech-to-speech and… 30 arXiv — NLP / Computation & Language research 13d ago Lost in Phonation: Voice Quality Variation as an Evaluation Dimension for Speech Foundation Models arXiv:2510.25577v2 Announce Type: replace-cross Abstract: Recent advances in Speech Foundation Models (SFMs) enable direct processing of raw audio, allowing models to respond to subtle paralinguistic variation. However, how these models interpret non-lexical cues remains largely… 6 r/LocalLLaMA community 13d ago [audio.cpp] Release 0.6: dots.tts, MiniMax-H3 text2audio (up to 3x realtime), MiniMax-Music3 (preview), and more new audio models. 5+ demos included. Hi all :) audio.cpp release 0.6 has been out for a little while, so this is more of an update on what landed and what has been improving around it. 0.6 added 5 new model families: dots.tts, NeuTTS-2e, MuScriptor (Music to MIDI), MiniMax-H3, and SenseVoice-Small, bringing… 25 r/LocalLLaMA community 14d ago A nice local vision test What is the meter reading? It should be 37461. What does your favourite vision model give? (Typo fixed)   submitted by   /u/MrMrsPotts [link]   [comments] 8 Hugging Face Daily Papers research 16d ago UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Abstract UniSwap enables synchronized appearance and voice replacement in talking videos through a unified streaming audio-visual diffusion transformer with specialized training and inference adaptations. Generated by thinkingmachines/Inkling-Small Talking-video character… 31 arXiv — NLP / Computation & Language research 16d ago Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition arXiv:2608.12327v1 Announce Type: new Abstract: Multilingual pretrained models nominally support Nepali, yet no controlled benchmark has compared them under a single fine-tuning protocol. We fine-tune six pretrained models (XLSR-53, IndicWav2Vec, MMS-1B, Whisper-Medium,… 34 arXiv — NLP / Computation & Language research 16d ago FastThaiG2P: Lightning-fast Thai Grapheme-to-phoneme Conversion for Voice Agent Pipelines arXiv:2608.12814v1 Announce Type: new Abstract: FastThaiG2P provides sub-millisecond Thai grapheme-to-phoneme conversion for text-to-speech pipelines (International Phonetic Alphabet and Kokoro-TTS conventions) using a PyThaiNLP-tokenized, extensible dictionary and normalization… 31 arXiv — NLP / Computation & Language research 16d ago CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model arXiv:2608.13101v1 Announce Type: new Abstract: Research on automatic speaking assessment (ASA) has increasingly adopted multimodal speech large language models to assess learners' speaking performance. However, existing studies provide limited analysis of how acoustic and… 36 arXiv — NLP / Computation & Language research 16d ago Motor, Cognitive, or Corpus? What Survives Cross-Lingual Transfer in Speech-Based Parkinsons Disease Detection arXiv:2608.13425v1 Announce Type: new Abstract: Self-supervised learning (SSL) speech representations achieve strong performance for Parkinson's disease (PD) detection within individual corpora. However, it remains unclear whether these models capture disease-related… 28 arXiv — NLP / Computation & Language research 17d ago DonorRank: Donor Language Selection for Low-Resource Cross-Lingual Speech Recognition arXiv:2608.11441v1 Announce Type: new Abstract: Low-resource automatic speech recognition (ASR) commonly relies on cross-lingual transfer, where models are adapted from higher-resource donor languages. However, selecting donors remains challenging for spontaneous speech from… 18 arXiv — NLP / Computation & Language research 17d ago Easper: An Accessible ASR Pipeline for Language Documentation arXiv:2608.11629v1 Announce Type: new Abstract: Audio transcription is a critical bottleneck in language documentation. While multilingual Automatic Speech Recognition (ASR) models like Whisper offer solutions, field linguists often lack the expertise to utilise them. We present… 4 arXiv — NLP / Computation & Language research 17d ago Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning arXiv:2608.11587v1 Announce Type: cross Abstract: Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challenging due to limited labeled data, low… 18 arXiv — NLP / Computation & Language research 17d ago Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder arXiv:2608.11650v1 Announce Type: cross Abstract: Recent advances in zero-shot text-to-speech (TTS) have substantially improved speech quality and voice cloning fidelity. However, many zero-shot TTS systems still depend on audio prompt transcripts at inference time. This… 8 arXiv — NLP / Computation & Language research 17d ago RT-SEMamba: Real-Time Speech Enhancement Mamba via Progressive Knowledge Distillation arXiv:2608.12099v1 Announce Type: cross Abstract: We present RT-SEMamba, a fully causal speech enhancement (SE) model built upon causal time-frequency Mamba blocks. Unlike Transformer-based architectures that rely on a growing key-value cache, Mamba propagates a fixed-size… 21 arXiv — NLP / Computation & Language research 17d ago Marco-Voice Technical Report arXiv:2508.02038v5 Announce Type: replace Abstract: This paper presents a multifunctional speech synthesis system that integrates voice cloning and emotion control speech synthesis within a unified framework. The goal of this work is to address longstanding challenges in… 13 TechCrunch — AI news-outlet 17d ago Why Stream ring-maker Sandbar says the future of AI wearables is voice AI notetaking hardware has taken off over the past couple of years, with credit-card-sized devices, pendants, pins, and even transcribing earbuds all promising to capture your meetings and turn them into summaries and action items. Now, a whole wave of… 4 TechCrunch — AI news-outlet 17d ago Why Sandbar thinks it’s voice-enabled ring can avoid the AI hardware graveyard AI notetaking hardware has taken off over the past couple of years, with credit-card-sized devices, pendants, pins, and even transcribing earbuds all promising to capture your meetings and turn them into summaries and action items. Now, a whole wave of… 19 llama.cpp releases dev-tools 18d ago b10369 mtmd: support pocket-tts ( #26871 ) adapt the api text model ok working impl, need verify and clean up mtmd: build the pocket-tts transposed convolutions as GEMM + col2im ggml_conv_transpose_1d has no grouped mode, so the depthwise upsample was built as one convolution and one… 16 arXiv — NLP / Computation & Language research 18d ago ASR-Roundtrip Evaluation Can Mask Context- and Convention-Dependent Reading Errors in Chinese News TTS arXiv:2608.10606v1 Announce Type: new Abstract: ASR-roundtrip evaluation is widely used as a scalable proxy for text-to-speech (TTS) intelligibility, but it can produce false negatives for reading errors perceived by listeners. We study Chinese news TTS spans whose correct… 5 arXiv — NLP / Computation & Language research 18d ago Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASR arXiv:2608.10670v1 Announce Type: new Abstract: At corpus sizes typical of low-resource dialects, single-run comparisons can yield gains that do not replicate. We show this for Garhwali, an under-resourced Indo-Aryan language of the central Himalaya, building the first… 28 arXiv — NLP / Computation & Language research 18d ago X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction arXiv:2608.10878v1 Announce Type: new Abstract: Accurate and responsive turn-taking is essential for spoken dialogue systems, which must distinguish in real time between user interruptions, backchannels that should be ignored, and the completion of an utterance. Prior modular… 27 Page 2 of 10 · 500 articles ← Newer Older →