News / #paper Tag Research papers 500 articles archived under #paper · RSS Sign in to follow arXiv — Machine Learning research 2d ago Classical and Hybrid Quantum Machine Learning for Trigger-Like Event Selection on CMS Open Data: An Eight-Qubit, PCA-Constrained Benchmark arXiv:2608.26224v1 Announce Type: cross Abstract: Event triggering sits at the heart of high-energy physics, where the rare events of interest must be retained while an overwhelming background is discarded under tight latency and bandwidth budgets. This work compares four… 23 arXiv — NLP / Computation & Language research 2d ago TreeGraft: Adaptive Multi-Drafter Grafting for Tree-Based Speculative Decoding arXiv:2608.26112v1 Announce Type: new Abstract: Speculative decoding accelerates large language model inference through a draft-then-verify paradigm. Building on this, tree-structured methods improve inference by organizing proposals into multiple candidate paths, increasing the… 27 arXiv — NLP / Computation & Language research 2d ago ElementCheck: Complexity-Aware Long-Form Text Factuality Evaluation via Sentence Elements arXiv:2608.26118v1 Announce Type: new Abstract: Existing long-form factuality evaluation relies on the decompose-retrieve-verify pipeline. However, the pipeline suffers from noise from claim decomposition and fixed verification granularity, resulting in unreliable results. We… 26 arXiv — NLP / Computation & Language research 2d ago DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs arXiv:2608.26119v1 Announce Type: new Abstract: Whether large language models can be prompted to generate rhetorical fallacies on demand, and whether current safety post-training constrains this behavior, has received less attention than the related question of detecting… 30 arXiv — NLP / Computation & Language research 2d ago Which India Survives Translation? Narrative Homogenisation Across Indian Oral Traditions in LLMs arXiv:2608.26123v1 Announce Type: new Abstract: Large language models (LLMs) are trained predominantly on English-language internet text that over-represents certain cultural narratives, raising concerns that models flatten the diversity of non-Western storytelling traditions… 10 arXiv — NLP / Computation & Language research 2d ago Natural-Language Policies to Executable Decisions: An Interpretable Large Language Model Framework arXiv:2608.26124v1 Announce Type: new Abstract: Pricing automation in large-scale tourism is challenging because travel orders are highly unstructured, while pricing policies are complex, rapidly evolving, and inherently open-ended. Traditional rule engines are brittle and… 33 arXiv — NLP / Computation & Language research 2d ago Training-Time Explainability for Multilingual Hate Speech Detection: Aligning Model Reasoning with Human Rationales arXiv:2608.26125v1 Announce Type: new Abstract: Online hate against Muslim communities often appears in culturally coded, multilingual forms that evade conventional AI moderation. Such systems, though accurate, remain opaque and risk bias, over-censorship, or under-moderation,… 17 arXiv — NLP / Computation & Language research 2d ago TelecomGPT-R1: A Unified Open-Source Reasoner for the Telecom Stack arXiv:2608.26126v1 Announce Type: new Abstract: Telecommunications is a high-leverage domain for large language model (LLM)-based reasoning because routine engineering workflows require joint grounding in normative specifications, operational telemetry, vendor-specific fault… 5 arXiv — NLP / Computation & Language research 2d ago Agents Don't Paginate: First-Chunk Selection for LLM Tool Responses arXiv:2608.26130v1 Announce Type: new Abstract: Coding agents built on large language models (LLMs), such as Claude Code, Cursor, OpenAI Codex, GitHub Copilot, and Aider, receive tool responses that routinely exceed the agent's per-turn token budget. The standard remedy,… 32 arXiv — NLP / Computation & Language research 2d ago Evaluating Language Models in Realistic Conversational Contexts arXiv:2608.26131v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed to serve open-ended, multi-turn interactions, evaluating conversational quality at human scale has become a central challenge. Existing evaluation frameworks built for… 20 arXiv — NLP / Computation & Language research 2d ago Agent Seer: Synthesizing Scenarios from Specification Understanding arXiv:2608.26133v1 Announce Type: new Abstract: Evaluating AI agents that use external tools requires realistic test scenarios that capture how practitioners compose tools and iterate across conversation turns. Constructing such scenarios by hand demands deep domain expertise,… 18 arXiv — NLP / Computation & Language research 2d ago Data Science Approaches to Evaluating Honours Candidates arXiv:2608.26135v1 Announce Type: new Abstract: We present a modular data-science pipeline for estimating public sentiment towards individuals from fragmented, unstructured open-source intelligence (OSINT). The method chains web search, text extraction, relevance filtering,… 11 arXiv — NLP / Computation & Language research 2d ago Reward-Informed Sparse Autoencoders and the Solution-Completeness Confound arXiv:2608.26136v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) decompose language-model activations into sparse, interpretable features, and an appealing way to aim them at reasoning is to curate their data with a signal reinforcement learning already produces: the… 26 arXiv — NLP / Computation & Language research 2d ago Syntax vs. Semantics: How Transformers Learn Deep Dependencies arXiv:2608.26139v1 Announce Type: new Abstract: Large Language Models demonstrate remarkable syntactic fluency, yet the optimization dynamics governing their acquisition of deep semantic dependencies remain poorly understood. We propose a mechanistic framework that models this… 4 arXiv — NLP / Computation & Language research 2d ago Position Is All You Need: A Free Lunch Token Compression Strategy for MLLM-based Referring Expression Segmentation arXiv:2608.26142v1 Announce Type: new Abstract: Referring Expression Segmentation (RES) aims to generate pixel-wise segmentation masks from complex and implicit textual queries. While recent advances in Multimodal Large Language Models (MLLMs) have substantially boosted RES… 15 arXiv — NLP / Computation & Language research 2d ago Beyond Accuracy: A Qualitative Analysis of Vision-Language Models for Hate Speech Detection in Memes arXiv:2608.26143v1 Announce Type: new Abstract: Memes have turned out to be a powerful tool through which individuals share their ideas concerning contemporary social and political problems. Their anonymity, as well as their ability to go viral, make them a powerful medium for… 28 arXiv — NLP / Computation & Language research 2d ago Why Current XAI Is Not Enough for Arabic NLP: A Critical Survey of the Explainability Gap arXiv:2608.26144v1 Announce Type: new Abstract: Explainable AI (XAI) is now a major theme in NLP; however, Arabic NLP remains under-explained in three connected senses. First, there is a method gap: Arabic XAI relies heavily on a small set of post-hoc techniques such as LIME,… 21 arXiv — NLP / Computation & Language research 2d ago Vagdhenu: A Vrutta (Meter) Aware Shloka-to-Chant (TTS) System for Sanskrit arXiv:2608.26146v1 Announce Type: new Abstract: We present Vagdhenu, a vrutta (meter) aware shloka-to-chant system for Sanskrit: a text-to-speech system that maps a metrical verse to its chanted parayana recitation at high fidelity. This is an experience report, not a new… 33 arXiv — NLP / Computation & Language research 2d ago CARE: Causally-Aligned Reasoning Exploration for Medical Large Language Models arXiv:2608.26147v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown strong potential for medical reasoning, yet the scarcity and cost of expert-annotated data constrain their progress. While reinforcement learning offers a scalable alternative, standard… 21 arXiv — NLP / Computation & Language research 2d ago Towards Interpretable Depression Detection: Linking Acoustic Features to DSM-5 Indicators arXiv:2608.26148v1 Announce Type: new Abstract: Depression affects millions worldwide, yet diagnosis relies on subjective self-reports that may miss authentic behavior. This paper presents an approach linking speech acoustics to DSM-5 depressive-behavior indicators through a… 26 arXiv — NLP / Computation & Language research 2d ago Artificial Intelligence Models Can Predict and Collaboratively Modulate Human Memory Search arXiv:2608.26152v1 Announce Type: new Abstract: Large language models (LLMs) exhibit unprecedented natural language generation and many text-based problem-solving capabilities. Indeed, in many language-based tasks, for example routine coding, these artificial intelligence models… 28 arXiv — NLP / Computation & Language research 2d ago Evaluating AI Generated Summaries for Cancer Patients arXiv:2608.26154v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly being integrated into digital health platforms to generate summaries of complex medical data. Although these models can improve patient engagement and communication, these systems also… 7 arXiv — NLP / Computation & Language research 2d ago VFA: Empowering Multilingual MLLMs via Vision-Free Adaptation arXiv:2608.26155v1 Announce Type: new Abstract: Multimodal large language models have advanced rapidly, yet most remain English-centric, as scaling multilingual multimodal instruction tuning is limited by the scarcity and high cost of high-quality non-English image-text… 6 arXiv — NLP / Computation & Language research 2d ago Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation arXiv:2608.26159v1 Announce Type: new Abstract: Self-Generated Text Recognition (SGTR)--the ability of an LLM to identify its own outputs--poses risks to AI safeguards that rely on LLMs as evaluators or monitors. Specifically, an LLM may recognize outputs from other copies of… 34 arXiv — NLP / Computation & Language research 2d ago Mutual Debiasing via Dual-Seed Comparison for Probabilistic Sampling in Large Language Models arXiv:2608.26161v1 Announce Type: new Abstract: Although Large Language Models (LLMs) demonstrate remarkable capabilities in reasoning and decision-making, high-fidelity probabilistic sampling remains a persistent challenge. When generating random variables, LLMs consistently… 14 arXiv — NLP / Computation & Language research 2d ago From Sound to Symptom: Real-Time Respiratory Signal Understanding for Conversational Healthcare Agents arXiv:2608.26163v1 Announce Type: new Abstract: Cough events during live spoken conversations carry clinically valuable respiratory signals, yet existing dialogue systems treat them as acoustic noise to be discarded. We present HealthCUES (Clinical Understanding from Embodied… 31 arXiv — NLP / Computation & Language research 2d ago Using Poly-Encoders for Computationally Efficient Automated Creativity Assessment arXiv:2608.26165v1 Announce Type: new Abstract: Automated creativity assessment has been a long standing challenge, with traditional methods often being resource intensive or lacking practical accuracy. We introduce a novel approach by using Poly-Encoder for computationally… 8 arXiv — NLP / Computation & Language research 2d ago Hallucinations in LLMs: A Lifecycle-Based Survey of Causes, Detection, Mitigation, and Prevention arXiv:2608.26168v1 Announce Type: new Abstract: The lifecycle of hallucination in LLMs is a concept that enables building solid frameworks on the control and reliability of LLMs in high-stakes environments, including health, legal, and scientific research. Although previous… 16 arXiv — NLP / Computation & Language research 2d ago Lost in Compression: A Controlled Cross-Lingual Audit of Extractive Prompt Compressors arXiv:2608.26175v1 Announce Type: new Abstract: Extractive prompt compression promises to cut LLM inference costs by removing low-information tokens, and learned compressors such as LLMLingua-2 report strong results on English benchmarks. Most other languages already pay a token… 4 arXiv — NLP / Computation & Language research 2d ago A Multi-Framework Comparison of Outline Stages in Long-Form Generation with LLMs arXiv:2608.26177v1 Announce Type: new Abstract: Long-form generation exposes fundamental limitations of large language models. Even 70B-parameter models exhibit length collapse at 16k-token outputs, and multi-chapter stories frequently trigger the attribute drift characteristic… 11 arXiv — NLP / Computation & Language research 2d ago PACEShop: Evaluating Personalized, Actionable, Compositional, and Evidence-grounded Shopping Assistants arXiv:2608.26180v1 Announce Type: new Abstract: Shopping assistants are shifting from ranked product lists toward structured decision support, where systems must synthesize shopper context, product evidence, and next-step guidance into a coherent recommendation experience. This… 26 arXiv — NLP / Computation & Language research 2d ago Investigating the Influence of Prompt and Response Languages on LLM Content Generation arXiv:2608.26186v1 Announce Type: new Abstract: This study examines how prompt and response language influence the behavior of large language models. Using five models, we evaluated answers to 68 non translation questions across four language conditions: English to English,… 18 arXiv — NLP / Computation & Language research 2d ago Comparing Chunking and Embedding Strategies for Turkish RAG Systems arXiv:2608.26192v1 Announce Type: new Abstract: How documents are segmented into retrievable chunks and how those chunks are embedded strongly affect Retrieval-Augmented Generation (RAG) quality, yet neither has been systematically studied for morphologically rich languages such… 8 arXiv — NLP / Computation & Language research 2d ago A Reranker for Orchestrating Heterogeneous Speech and Text Retrievers arXiv:2608.26194v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) systems have attracted significant interest for their ability to mitigate hallucinations in Large Language Models (LLMs). Although knowledge databases for RAG are increasingly diversifying to… 16 arXiv — NLP / Computation & Language research 2d ago On Scope Classification and Current Knowledge-Editing Benchmarks: A Negative Result, with INLAY as a Gradient-Free Case Study arXiv:2608.26292v1 Announce Type: new Abstract: Every memory-based knowledge editor in the SERAC lineage depends on a scope decision: given a query, does a stored edit apply? We report that current knowledge-editing benchmarks cannot measure this decision at all. Using INLAY, a… 27 arXiv — NLP / Computation & Language research 2d ago MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models arXiv:2608.26295v1 Announce Type: new Abstract: Tool-augmented LLMs must arbitrate between two fallible sources when a tool return conflicts with their parametric memory, yet existing evaluations measure source preference without establishing source correctness. We introduce… 25 arXiv — NLP / Computation & Language research 2d ago When Is Noise Response Universal? Tokenization as the Hidden Variable in Language Models arXiv:2608.26319v1 Announce Type: new Abstract: The performance of textual neural models often degrades when their inputs are corrupted by noise such as typos, OCR errors, or dropped words. We study the degradation rate across neural models, both sentence embeddings and… 18 arXiv — NLP / Computation & Language research 2d ago How Unlikely Is "Unlikely"? Assessing Verbal Probability Perception Across Large Language Models arXiv:2608.26327v1 Announce Type: new Abstract: Large language models increasingly produce and interpret verbal probability expressions, yet whether these expressions carry consistent meaning across models (or match human perceptions of uncertainty) remains unknown. We present a… 33 arXiv — NLP / Computation & Language research 2d ago Neuro-symbolic PRM: Enhancing Scientific Reasoning via Structured Traces and Symbolic Verification arXiv:2608.26329v1 Announce Type: new Abstract: While tool-augmented Large Language Models have significantly improved multi-step reasoning in quantitative STEM tasks, a critical residual failure mode remains: intermediate reasoning steps that are syntactically well-formed,… 17 arXiv — NLP / Computation & Language research 2d ago MoganColBERT-TR: A Late-Interaction Multi-Vector Retrieval Model for Turkish arXiv:2608.26344v1 Announce Type: new Abstract: We previously reported a ModernBERT encoder trained from scratch for Turkish (MoganBERT-TR) and a single-vector embedding model built on top of it (MoganBERT-embed). This work introduces the third model in that lineage:… 37 arXiv — NLP / Computation & Language research 2d ago Cross-lingual Representation Learning via Centroid Intervention Fusion arXiv:2608.26357v1 Announce Type: new Abstract: Large language models (LLMs) exhibit uneven multilingual performance, especially when dealing with low-resource languages. Inference-time intervention offers a lightweight way to improve cross-lingual transfer by modifying the… 32 arXiv — NLP / Computation & Language research 2d ago Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives arXiv:2608.26372v1 Announce Type: new Abstract: Large language models are increasingly deployed as autonomous agents serving users on behalf of companies, placing them in settings where user and deployer interests can conflict. When an agent knows that a user is owed something… 23 arXiv — NLP / Computation & Language research 2d ago Survival-Guided Length Control for Efficient Diffusion Language Models arXiv:2608.26374v1 Announce Type: new Abstract: Diffusion language models (DLMs) generate text by iteratively denoising masked sequences, but standard decoding either fixes the sequence length or relies on ad hoc stopping rules, often leading to unnecessary denoising steps. We… 5 arXiv — NLP / Computation & Language research 2d ago Why RAGs Hallucinate: Penalty-Aware Evaluation of Retrieval-Augmented Generation Systems with Knowledge-Gap Canaries arXiv:2608.26385v1 Announce Type: new Abstract: Volume-based accuracy rewards retrieval-augmented generation (RAG) systems for guessing: a system that answers everything outscores one that declines when its knowledge base cannot support an answer. Building on the… 6 arXiv — NLP / Computation & Language research 2d ago Co-Evolving Structured Knowledge and Reasoning in Language Models arXiv:2608.26386v1 Announce Type: new Abstract: Retrieval-augmented methods improve factual accuracy by grounding language models in external knowledge, but retrieving over unstructured text often introduces irrelevant context and offers limited control over the retrieved… 23 arXiv — NLP / Computation & Language research 2d ago LowRankArena: A Standardized Evaluation Platform for SVD-Based LLM Compression arXiv:2608.26389v1 Announce Type: new Abstract: SVD-based low-rank compression has become a fast-growing direction for reducing the memory and computational cost of large language models (LLMs). However, meaningful comparison across existing studies remains difficult as prior… 30 arXiv — NLP / Computation & Language research 2d ago Case2Flow: Bridging Patient Cases and Guideline Flowcharts through Multimodal Retrieval arXiv:2608.26414v1 Announce Type: new Abstract: Medical guidelines encode rich, evidence-based decision logic, yet the specific decision artifact a clinician needs is hard to locate within a guideline, let alone across guidelines covering plausible diseases and treatments. While… 26 arXiv — NLP / Computation & Language research 2d ago AfriSwitch: A Benchmark for In-the-Wild African Code-Switched Speech Recognition arXiv:2608.26434v1 Announce Type: new Abstract: Code-switching is pervasive in bilingual African conversation, yet most ASR systems assume monolingual input and are evaluated on curated monolingual benchmarks. We present AfriSwitch, a 61.36-hour human-transcribed benchmark of… 22 arXiv — NLP / Computation & Language research 2d ago Vowel Signs Are Not Letters: A Pre-tokenization Ceiling on Multilingual Tokenizer Fertility arXiv:2608.26449v1 Announce Type: new Abstract: Byte-level BPE tokenizers that use the HuggingFace ByteLevel pre-tokenizer inherit GPT-2's word regex, where a word is defined as \p{L}+, one or more Unicode letters. In abugida scripts, vowels are written as combining marks; this… 35 arXiv — NLP / Computation & Language research 2d ago Compositional Generalization via Structural Identification in a Category-Theoretic Framework arXiv:2608.26465v1 Announce Type: new Abstract: Compositional generalization is usually evaluated through model accuracy. We instead ask which structural or lexical identifications make held-out COGS examples admissible from the structures observed in training. Sentences are… 27 Page 3 of 10 · 500 articles ← Newer Older →