News / #voice Tag Voice 500 articles archived under #voice · RSS Sign in to follow OpenAI official-blog 1mo ago Introducing OpenAI Presence Introducing OpenAI Presence, a proven enterprise AI agent platform that helps organizations deploy trusted voice and chat agents for customer and internal workflows. 10 arXiv — NLP / Computation & Language research 1mo ago From a Multilingual Streaming ASR Backbone to Kenyan-Language Systems: Data-Centric Adaptation of Nemotron 3.5 for Kikuyu, Dholuo, and Kalenjin arXiv:2607.18912v1 Announce Type: new Abstract: Automatic speech recognition (ASR) for African languages is constrained by orthographic inconsistency, annotation artifacts, missing audio, speaker and domain imbalance, and evaluation procedures that differ from deployment. We… 24 arXiv — NLP / Computation & Language research 1mo ago Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing arXiv:2607.18934v1 Announce Type: new Abstract: Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an uncontrolled latent variable, causing measurable decoding instability, evaluation confounding (up to 60% of… 22 arXiv — NLP / Computation & Language research 1mo ago Constrained CTC Decoding for Efficient Diacritic Restoration arXiv:2607.18946v1 Announce Type: new Abstract: In this work, we address diacritic restoration for Arabic speech transcripts. Most speech data are undiacritized, limiting the ability of modeling fine-grained phonological distinctions. The speech modality has recently been… 34 arXiv — NLP / Computation & Language research 1mo ago Content is What Remains: Invariant Speech Tokenization from Parallel Utterances arXiv:2607.19033v1 Announce Type: new Abstract: Discrete speech tokenizers aim to disentangle semantic from acoustic information, yet targets from self-supervised learning (SSL) models like HuBERT retain non-linguistic variation: speaker identity, prosody, and channel conditions… 26 arXiv — NLP / Computation & Language research 1mo ago Benchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial Results arXiv:2607.19049v1 Announce Type: new Abstract: Humans are often considered to be the best listeners and seen as the upper-bound performance of automatic speech recognition (ASR) systems. We present a preliminary comparison of the performances of state-of-the-art ASR systems and… 8 arXiv — NLP / Computation & Language research 1mo ago A Situational Speech Synthesizer for Yoruba: System Design, Phonological Rule Architecture, and Orthographic Extensions for Contour arXiv:2607.18317v1 Announce Type: cross Abstract: We present TTSYoruba, a rule-based concatenative diphone speech synthesizer for Yoruba, deployed at online as part of the YorubaName.com open dictionary of Yoruba personal names. The system takes tone-marked Yoruba text as input… 32 arXiv — NLP / Computation & Language research 1mo ago Staged Depth-Pruning Distillation of a Flow-Matching Text-to-Speech Teacher: A Compact Hindi Speech Synthesizer arXiv:2607.18662v1 Announce Type: cross Abstract: We present a practical recipe for building a compact Hindi text-to-speech (TTS) model by distilling a large flow-matching teacher (IndicF5, 337M-parameter DiT) under a severe data budget (~17.6 hours). Training a small model from… 12 arXiv — Machine Learning research 1mo ago Robust Assamese Speech Recognition through Controlled Fine-Tuning of Whisper Models arXiv:2607.17164v1 Announce Type: new Abstract: Developing Automatic Speech Recognition (ASR) for morphologically rich, low-resource languages such as Assamese is challenging due to insufficient annotated speech data. The pretrained Whisper model performs poorly on Assamese… 16 arXiv — NLP / Computation & Language research 1mo ago AI_LectureNote: A Retrospective Pilot Study of a Post-ASR Workflow for English-Script Rendering and Semantic Drift in Korean-English Medical Lectures arXiv:2607.17237v1 Announce Type: new Abstract: AI_LectureNote is a historical, readability-oriented post-ASR workflow for Korean-English medical lectures. It rewrites speech-to-text output into study transcripts while restoring Latin-script medical terms rather than Korean… 5 arXiv — NLP / Computation & Language research 1mo ago When to Use Extra Context: Evidence-Grounded Terminology Adaptation for Simultaneous Speech Translation arXiv:2607.17766v1 Announce Type: new Abstract: Extra context is valuable for simultaneous speech translation of technical talks, but injecting the entire document context into every streaming segment is often too coarse. Through diagnostic experiments, we find that context… 37 arXiv — NLP / Computation & Language research 1mo ago ESCUCHA: A Spanish Speech Benchmark for Heterogeneous Acoustic Conditions arXiv:2607.17812v1 Announce Type: new Abstract: As large audio language models (LALMs) advance, robust evaluation frameworks have become essential. In this context, Spanish speech understanding under realistic acoustic conditions has received particularly little attention. We… 6 Hugging Face Daily Papers research 1mo ago FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications Abstract Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose efficient deployment requires application-specific decisions about placement, streaming, and intra-model parallelism. Existing… 10 r/LocalLLaMA community 1mo ago Running a 13M ASR conformer on a microcontroller Hello everyone, I wanted to share a recent project of mine, which brings a 13.1 million parameter convolution transformer model to a < $10 microcontroller (more specifically, the ESP32-S3). It's a distilled and quantized version of nvidias small conformer model from huggingface.… 38 r/LocalLLaMA community 1mo ago Introducing Scylla's Band, a new TTS model + inference framework with Android sample! Hey all! https://github.com/lowkeytea/scyllasband -> inference code https://huggingface.co/spybyscript/scyllasband -> model, LiteRT, ONNX, and voices https://lowkeytea.github.io/scyllasband/ -> sample audio for the voices, emotions, and languages. The tldr: 10 voices, 7… 13 Smol AI News news-outlet 1mo ago not much happened today **US policy debates** are moving toward restricting Chinese open models like **Kimi**, with potential **procurement restrictions** and **Entity List designations**. Technical voices including **@APompliano**, **@ClementDelangue**, and **@mmitchell_ai** warn this could harm… 18 r/LocalLLaMA community 1mo ago Good ASR and TTS models? Hey everyone, Something I don't see discussed often here are ASR and TTS models. I've been using Whisper and Kokoro (old models, I know!) with koboldcpp for a while now but wondered if there are now solid replacements available. Know of Qwen3-ASR and Qwen3-TTS, but haven't found… 12 arXiv — Machine Learning research 1mo ago SpeechGuard: Online Defense against Backdoor Attacks on Speech Recognition Models arXiv:2607.15697v1 Announce Type: cross Abstract: Backdoor attacks pose a critical threat to neural network models, allowing attackers to implant a backdoor during the training phase by manipulating a small portion of the training data. In security-sensitive applications such as… 34 arXiv — NLP / Computation & Language research 1mo ago Contextual Semantic Relevance Tracks fMRI BOLD Responses During Naturalistic Speech Comprehension arXiv:2607.15856v1 Announce Type: new Abstract: Naturalistic language comprehension requires listeners to process both local probabilistic expectations and contextual semantic relations. Surprisal has been widely used to quantify local word unexpectedness, but evidence that it… 4 arXiv — NLP / Computation & Language research 1mo ago Controlling Implicit Shortcut Reliance in L2 Spoken English Auto-markers arXiv:2607.16085v1 Announce Type: new Abstract: Increasingly, speech and language processing tasks take either audio or text directly rather than extracting features from these as the input to the classifier or regressor. Often these systems make use of complex, for example… 17 Hacker News — AI on Front Page community 1mo ago EEG shows brain can simultaneous encode two speech streams Article URL: https://journals.plos.org/plosbiology/article?id=10.1371/journal.pbio.3003876 Comments URL: https://news.ycombinator.com/item?id=48943745 Points: 209 # Comments: 131 5 arXiv — NLP / Computation & Language research 1mo ago PERL: Pinyin Enhanced Rephrasing Language Model for Chinese ASR N-best Error Correction arXiv:2412.03230v3 Announce Type: replace Abstract: Chinese ASR correction is challenging because errors are often \emph{phonetic} (many characters share similar Pinyin) while the correction model must also obey a \emph{length constraint} under noisy N-best hypotheses. Existing… 18 OpenAI official-blog 1mo ago How Cars24 scales conversations and builds faster with OpenAI Cars24 uses OpenAI-powered voice and chat agents to handle 1M+ monthly conversation minutes, recover 12% of lost leads, and bring agentic workflows to teams across the company. 13 r/LocalLLaMA community 1mo ago Hermes on Android (Graphene OS) https://youtu.be/oxpGq5FITgA?si=nkHWLReGCDYe7QfL I got Hermes running in the native Debian Terminal in Graphene OS and its really slick. Voice dictation works amazingly. Im using a remote Hermes gateway running on my laptop as the backend, with Llama.cpp and Qwen 3.6 35b. Paired… 14 arXiv — NLP / Computation & Language research 1mo ago Hybrid Continual Learning for Low-Resource Australian Aboriginal Language Identification arXiv:2607.11946v1 Announce Type: new Abstract: Language identification is an important step toward integrating endangered Australian Aboriginal languages (AALs) into speech technologies supporting language revitalisation and digital inclusion. However, extreme data scarcity… 9 arXiv — NLP / Computation & Language research 1mo ago Toward Metaphor-Fluid Conversation Design for Voice User Interfaces arXiv:2502.11554v3 Announce Type: replace-cross Abstract: Metaphors play a critical role in shaping user experiences with Voice User Interfaces (VUIs), yet existing designs often rely on static, human-centric metaphors that fail to adapt to diverse contexts and user needs. This… 7 Vercel — AI dev-tools 1mo ago How Speechify serves 500,000 dynamic pages to 60 million users on Vercel Speechify on Vercel 500,000+ pages served across 40+ languages Cut costs 50% by auto-scaling with Fluid compute Zero user impact on bad deploys with Instant Rollbacks Speechify started as a tool for people with dyslexia. Cliff Weitzman, Founder & CEO, built it because reading… 31 r/LocalLLaMA community 1mo ago [audio.cpp] 10 hours of audio generated in 3 minutes on RTX 5090 (demo included)! C++/GGML based Supertonic 3, MOSS-TTS, IndexTTS2, and Irodori-TTS released audio.cpp again. Hopefully you are not sick of it yet :) Release 0.3 adds five new models: Supertonic 3, MOSS-TTS-Local, MOSS-TTS-Nano, IndexTTS2, and Irodori-TTS. The highlight is Supertonic 3. It can hit 200 ×+ real time on CUDA (RTX 5090), 6×+ on CPU, and around 47 ms TTFT in… 37 Hugging Face official-blog 1mo ago Introducing Real World VoiceEQ: Measuring the human quality of voice AI Back to Articles a]:hidden"> Introducing Real World VoiceEQ: Measuring the human quality of voice AI Published July 15, 2026 Update on GitHub Upvote 6 David Ayllon dayllon HumeAI Alice aliceebaird HumeAI Jeff Brooks jeffbrooks HumeAI Franc Camps Febrer francamps HumeAI Jakub… 25 TechCrunch — AI news-outlet 1mo ago The founder of Hinge raised $18M to build a new AI dating service, Overtone Overtone describes itself as "a voice- and audio-forward service, enabled by AI, that provides highly curated introductions." 26 Hacker News — AI on Front Page community 1mo ago Speech Recognition and TTS in less than 500kb Article URL: https://github.com/moonshine-ai/moonshine/tree/main/micro Comments URL: https://news.ycombinator.com/item?id=48911793 Points: 299 # Comments: 33 21 arXiv — NLP / Computation & Language research 1mo ago Efficiently Adapting Spoken Language Models for the Singaporean Context arXiv:2607.10092v1 Announce Type: new Abstract: Spoken language models (SLMs) unify speech perception and reasoning, but adapting them to sensitive domains is underexplored, especially when the original training data is inaccessible and the use case demands multilingual,… 19 arXiv — NLP / Computation & Language research 1mo ago Which Languages Transfer Best to Warlpiri? A Similarity-Based Study for Low-Resource ASR arXiv:2607.10256v1 Announce Type: new Abstract: This paper investigates how language similarity can improve cross-lingual transfer for automatic speech recognition (ASR) in extremely low-resource settings. Warlpiri, an Australian Aboriginal language, has very limited transcribed… 5 arXiv — NLP / Computation & Language research 1mo ago Quantifying the Sources of Instability in LLM-Based Stance Analysis of Public Discourse arXiv:2607.10846v1 Announce Type: new Abstract: Computational social science increasingly relies on automated preprocessing pipelines -- speaker diarization, ASR transcript cleaning, sentence segmentation -- to convert raw media into analyzable text. When these pipelines produce… 38 arXiv — NLP / Computation & Language research 1mo ago Unified Gradient Projection: Language-Balanced Continual Learning for Multilingual Low-Resource ASR arXiv:2607.11163v1 Announce Type: new Abstract: Large-scale pretrained ASR models such as Whisper exhibit strong multilingual capabilities. However, fine-tuning on low-resource languages often causes catastrophic forgetting. Although continual learning mitigates this issue,… 36 arXiv — NLP / Computation & Language research 1mo ago FAD-SA-GRU: Enhancing Hate Speech Detection in Algerian Dialect Through Feature-Augmented Self-Attention GRU Networks arXiv:2607.11279v1 Announce Type: new Abstract: The widespread adoption of social media platforms has transformed online communication by enabling users to exchange information and opinions instantly. However, these platforms have also facilitated the dissemination of abusive… 34 arXiv — NLP / Computation & Language research 1mo ago Beyond Benchmarks: Exposing the Hidden Crisis in Bangla Hate Speech Detection arXiv:2607.11597v1 Announce Type: new Abstract: The spread of hate speech (HS) across different social media platforms (SMPs) poses a major concern for online safety and ethical moderation. Automatic detection of HS remains a challenging task, especially in under-resourced… 10 r/LocalLLaMA community 1mo ago Self-hosted voice for any agent/harness of your choice (open-source) For a while now I've been maintaining tts-bench ( https://github.com/5uck1ess/tts-bench ) and a blind voting arena ( https://5uck1ess-tts-arena.hf.space ) where people A/B test open text-to-speech (TTS) models without knowing which is which. One problem I've always wanted to… 31 Hacker News — AI on Front Page community 1mo ago Apple's new SpeechAnalyzer API, benchmarked against Whisper and its predecessor Article URL: https://get-inscribe.com/blog/apple-speech-api-benchmark.html Comments URL: https://news.ycombinator.com/item?id=48894752 Points: 206 # Comments: 104 10 arXiv — NLP / Computation & Language research 1mo ago Phone Segmentation and Recognition through Phonological Activation Mapping arXiv:2607.09020v1 Announce Type: cross Abstract: Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately. We argue that phonetic structure is already latent in the representations of self-supervised speech models… 32 arXiv — NLP / Computation & Language research 1mo ago FreyaTTS Technical Report arXiv:2607.09530v1 Announce Type: new Abstract: We introduce Freya-TTS, a compact, tokenizer-free, Turkish-first text-to-speech model designed for highly reliable and efficient conversational synthesis. Freya-TTS is a 183.2M-parameter non-autoregressive conditional flow-matching… 33 arXiv — NLP / Computation & Language research 1mo ago Tokenizer Transplantation: Mitigating Autoregressive Collapse in Edge-Efficient Bengali ASR arXiv:2607.09598v1 Announce Type: new Abstract: Lightweight speech recognition models are critical for edge deployment, yet highly optimized architectures like Moonshine often fail on morphologically rich, non-Latin languages such as Bengali. This study identifies the root cause… 26 arXiv — NLP / Computation & Language research 1mo ago Point of Order: Action-Aware LLM Persona Modeling for Data-Grounded Civic Deliberation arXiv:2511.17813v3 Announce Type: replace Abstract: LLM-based simulations can enable controlled studies of civic deliberation, but current systems lack speaker-attributed data and methods for evaluating long-form institutional behavior. ASR transcripts typically use anonymous… 5 Hugging Face Daily Papers research 1mo ago Phone Segmentation and Recognition through Phonological Activation Mapping Abstract Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately. We argue that phonetic structure is already latent in the representations of self-supervised speech models (S3Ms), and one only needs to steer them to… 16 r/LocalLLaMA community 1mo ago Current state of Voice-To-Voice models Hi, Has there been any improvement to Voice models (like RVC) in the last two years, or has nothing changed? Thanks!   submitted by   /u/Iwishlife [link]   [comments] 36 r/LocalLLaMA community 1mo ago How fast can I get a voice assistant to respond without a GPU? Qwen3-ASR and Kokoro-TTS ONNX on CPU. Been testing out the ONNX models to see how far I can push the CPU to take on ASR and TTS, so the GPU is completely free for running the LLM. The video attached shows me testing latency on a 2022 Macbook M2 and an AMD Ryzen 9 7900. This is just running the regex fast commands,… 17 OpenAI official-blog 1mo ago How Deutsche Telekom is rewiring telecommunications with AI How Deutsche Telekom is becoming an AI-native telco with OpenAI-transforming customer service, employee workflows, network operations, and the future of voice. 37 arXiv — NLP / Computation & Language research 1mo ago A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents arXiv:2607.07985v1 Announce Type: new Abstract: We report the empirical reliability of Gemini models as audio judges that score full-duplex agent conversations directly from the raw stereo waveform, tested across three models in the Gemini family: 2.5 Flash, 3.5 Flash, and 3.1… 5 arXiv — NLP / Computation & Language research 1mo ago COALA: Robust Contextualized Speech-augmented Language Modeling for ASR via Contrastive Regularizer and Biasing Score Estimation arXiv:2607.08117v1 Announce Type: new Abstract: Contextual biasing seeks to integrate external knowledge into automatic speech recognition (ASR) systems to accurately recognize domain-specific entities. In this paper, we propose COALA (Contextualized ASR Leveraging Biasing… 7 arXiv — NLP / Computation & Language research 1mo ago Diarization-Guided Qwen-ASR Adaptation for Multilingual Two-Speaker Conversational Speech arXiv:2607.08208v1 Announce Type: new Abstract: This paper describes our self-designed system for Task 1 of the MLC-SLM 2026 Challenge for multilingual two-speaker conversational speech. The system combines a modular speaker diarization front end with a challenge-adapted… 12 Page 5 of 10 · 500 articles ← Newer Older →