arXiv — NLP / Computation & Language
500 articles archived · Visit source ↗ · RSS
-
arXiv — NLP / Computation & Language research 5d ago
Length-Adaptive Decoding for Masked Diffusion Machine Translation
arXiv:2608.22274v1 Announce Type: new Abstract: Machine translation tests masked diffusion language models (dLLMs) because every source token must be rendered faithfully, while fixed canvas decoding must choose target length before denoising. Existing masked diffusion decoding…
5 -
arXiv — NLP / Computation & Language research 5d ago
LLM Evaluation on Unseen Questions: Contextual Multidimensional IRT Model
arXiv:2608.22295v1 Announce Type: new Abstract: Evaluation of large language models (LLMs) increasingly requires predicting how a model will perform on new questions or tasks before collecting large amounts of new annotations. This problem is challenging because question…
11 -
arXiv — NLP / Computation & Language research 5d ago
Text-Anchored Semantic Perturbations for Transferable Jailbreak Attacks on Multimodal Large Language Models
arXiv:2608.22312v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress in vision-language interaction, yet their safety alignment remains vulnerable to jailbreak attacks. A key challenge is that safety behavior learned in the…
13 -
arXiv — NLP / Computation & Language research 5d ago
Semantics or Structure? Auditing Text Sensitivity in Multimodal Time-Series Forecasting
arXiv:2608.22321v1 Announce Type: new Abstract: Multimodal time-series forecasting has emerged as a promising paradigm in which natural-language context is expected to improve predictive performance. Recent multimodal foundation models, including Aurora, as well as early- and…
35 -
arXiv — NLP / Computation & Language research 5d ago
Noise Floor Audit for Agent Benchmarks
arXiv:2608.22331v1 Announce Type: new Abstract: We audit measurement variability for 3 native tool-calling endpoints across 2 providers on the official BFCL multiple and parallel categories, using matched AST grading. At temperature 0, reruns are nearly deterministic across Groq…
16 -
arXiv — NLP / Computation & Language research 5d ago
Mechanistic Interpretability of Chain-of-Thought Reasoning via Sequential Activation Patching
arXiv:2608.22332v1 Announce Type: new Abstract: Large Language Models (LLMs) demonstrate remarkable problem-solving capabilities when guided by Chain-of-Thought (CoT) prompting, yet the internal mechanisms underlying these improvements remain poorly understood. In this work, we…
15 -
arXiv — NLP / Computation & Language research 5d ago
Register Shifts Break LLM Safety: A Bengali Benchmark with Culturally Grounded Harms
arXiv:2608.22335v1 Announce Type: new Abstract: Bengali is the seventh-most-spoken language globally, yet LLM safety evaluation remains overwhelmingly English-centric. We introduce BanglaSafe, a benchmark of 879 Bengali prompts combining 309 natively authored prompts with 570…
12 -
arXiv — NLP / Computation & Language research 5d ago
When Not to Imitate: Boundary-Aware Skill Memory for Reliable Tool-Use LLM Agents
arXiv:2608.22339v1 Announce Type: new Abstract: Extracting skills from past successes is critical for the efficient evolution of Large Language Model (LLM) agents. Prevailing agent self-evolution paradigms typically rely on a core assumption: equipping LLMs with skill memories…
12 -
arXiv — NLP / Computation & Language research 5d ago
Context-Aware Cluster Decoding: Semantic Anchor-Driven Coherence in dMLLMs
arXiv:2608.22367v1 Announce Type: new Abstract: Diffusion multimodal large language models (dMLLMs) frequently produce long-form outputs marred by semantic drift and repetition, with quality generally degrading as output length increases. We identify two structural deficiencies…
4 -
arXiv — NLP / Computation & Language research 5d ago
Can Large Language Models "Hyper-Thread"?
arXiv:2608.22376v1 Announce Type: new Abstract: Large language models generate tokens sequentially, but can they execute multiple tasks concurrently while forming each token? Broader attention allocation may provide a mechanism for such task concurrency. Existing approaches to…
31 -
arXiv — NLP / Computation & Language research 5d ago
ProBel: Propaganda Detection with Techniques, Spans, and Explanations
arXiv:2608.22388v1 Announce Type: new Abstract: Propaganda detection includes several related prediction levels, ranging from sentence-level decisions to technique classification and span identification. However, it remains unclear how supervision at these levels interacts when…
10 -
arXiv — NLP / Computation & Language research 5d ago
SchemaGUI: A Schema-Driven Benchmark for Controllable GUI Generation Evaluation
arXiv:2608.22390v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong potential in graphical user interface (GUI) generation, but reliable evaluation remains challenging due to uncontrolled data distributions, noisy annotations, and limited layout…
35 -
arXiv — NLP / Computation & Language research 5d ago
Don' t Box Me In: Dynamic Cultural Adaptation and Cognitive Tracking for Social Understanding
arXiv:2608.22411v1 Announce Type: new Abstract: Social interaction increasingly takes place in multicultural settings, where individuals may draw on multiple cultural influences and adapt their communicative behavior across contexts. Despite recent advances in equipping Large…
27 -
arXiv — NLP / Computation & Language research 5d ago
Rank Reversal in Multilingual LLM Judges: A Label-Free Double-Centering Calibrator
arXiv:2608.22432v1 Announce Type: new Abstract: Multilingual LLM judges produce different evaluator-backbone rankings depending on the prompt language: on an eight-language Agent-as-a-Judge benchmark, the top-ranked backbone alternates across English, Arabic, Chinese, Hindi,…
33 -
arXiv — NLP / Computation & Language research 5d ago
Aligned Alone, Misaligned Together: Forecasting Adversarial Capture in LLM Agent Populations
arXiv:2608.22444v1 Announce Type: new Abstract: The unit of AI safety evaluation is still the individual model, yet language-model agents are increasingly deployed in interacting populations that read and write one another's decisions. This raises a question no single-agent…
33 -
arXiv — NLP / Computation & Language research 5d ago
Figurative Justice: Detecting metaphors in Hindi judgements with qualitative assessment and transformers
arXiv:2608.22446v1 Announce Type: new Abstract: Metaphors are figurative use of words for conceptual mapping. Metaphor detection in the legal context has been crucial as metaphors are persuasive juridical means of creating legal meaning and concepts resulting in significant…
6 -
arXiv — NLP / Computation & Language research 5d ago
From Exposure to Expectation: Frequency, Surprisal, and Language Across Development in Spanish
arXiv:2608.22452v1 Announce Type: new Abstract: Surprisal, the negative log-probability a language model assigns to a word given its preceding context, reliably predicts adult reading times. Does it contribute as much to explaining when children acquire individual words?…
25 -
arXiv — NLP / Computation & Language research 5d ago
GTA-RAG: Graph-Trajectory-Augmented Reinforcement Learning for Multi-Turn Retrieval-Augmented Reasoning
arXiv:2608.22479v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) enables LLMs to access external knowledge for answering knowledge-intensive questions. For complex multi-hop questions, multi-turn retrieval-augmented reasoning extends RAG into an iterative…
30 -
arXiv — NLP / Computation & Language research 5d ago
Claim-Level Confidence Calibration for Reliable Decision Making with Large Language Models
arXiv:2608.22483v1 Announce Type: new Abstract: Large Language Models (LLMs) increasingly support decision-making in high-stakes domains, but they often hallucinate and express confidence that is misaligned with factual correctness. Response-level confidence is a coarse signal:…
30 -
arXiv — NLP / Computation & Language research 5d ago
Who Pays More for Safety? Measuring the Disparate Cost of Safety Alignment across Languages
arXiv:2608.22490v1 Announce Type: new Abstract: Safety alignment helps models adhere to human values, but it often reduces response utility. We ask a critical but understudied question: Does safety alignment impose the cost equally across language groups? To answer this, we…
26 -
arXiv — NLP / Computation & Language research 5d ago
Kernel Token Contradiction: a Fast and Principled Approach for LLM Claim Uncertainty Quantification
arXiv:2608.22506v1 Announce Type: new Abstract: Claim-level Uncertainty Quantification (UQ) aims to mitigate the lack of reliability of Large Language Models (LLMs) by evaluating the factuality of each claim in their outputs. We introduce Kernel Token Contradiction (KTC), a…
38 -
arXiv — NLP / Computation & Language research 5d ago
From Diagnosis to Redesign: Using Quantitative Ethnography to Improve Multi-Agent LLM Reasoning
arXiv:2608.22566v1 Announce Type: new Abstract: Multi-agent large language model (LLM) systems are designed to improve reasoning by decomposing tasks across multiple agents with specialized functions, but the presence of multiple agents does not inherently guarantee coherent…
20 -
arXiv — NLP / Computation & Language research 5d ago
Hybrid Panels: Toward Human-AI Collaboration in Survey Research
arXiv:2608.22582v1 Announce Type: new Abstract: Large-scale population surveys are essential for generating robust social and scientific insights, yet they face significant challenges, including declining response rates, increasing data collection costs, long delays between data…
27 -
arXiv — NLP / Computation & Language research 5d ago
Teaching LLMs How ICU Physicians Approach Clinical Reasoning Through OMOP-Aligned Retrieval Improves Reasoning Across Clinical Domains
arXiv:2608.22622v1 Announce Type: new Abstract: Clinical decision-making relies on identifying relevant patient information to guide diagnosis and treatment, a challenge that is especially difficult in the data-dense and rapidly changing intensive care unit (ICU). Large language…
11 -
arXiv — NLP / Computation & Language research 5d ago
GeoRisk-RAG: A Hierarchy-Aware Risk Framework for Improving RAG Reliability through Selective Answering
arXiv:2608.22634v1 Announce Type: new Abstract: Current work on improving reliability in large language model (LLM)- generated answers has primarily leveraged Retrieval-Augmented Generation (RAG), knowledge-graph augmentation, and reinforcement learning. While these methods are…
37 -
arXiv — NLP / Computation & Language research 5d ago
Iteration Without Elaboration: A Simple ReAct Architecture Suffices for Text-to-SQL Generation
arXiv:2608.22651v1 Announce Type: new Abstract: Modern text-to-SQL systems have become increasingly elaborate, relying on schema-linking modules, retrieval-augmented prompting, candidate generation, and multi-stage refinement pipelines. While effective, these additions introduce…
27 -
arXiv — NLP / Computation & Language research 5d ago
Enrich-Retrieve-Rank: Scaling Capability Discovery Beyond In-Context Routing
arXiv:2608.22695v1 Announce Type: new Abstract: Agent ecosystems now include thousands of MATS components (Models, Agents, Tools, and Skills), yet their discovery still relies on in-context routing. These systems read a registry (names, hints, or descriptions, as context budget…
26 -
arXiv — NLP / Computation & Language research 5d ago
WnW: Waxing-and-Waning KV Cache for Long-Form Speech LLMs
arXiv:2608.22704v1 Announce Type: new Abstract: Long-form audio inputs make the KV cache the dominant memory cost of speech LLMs. Prefill-only KV compression methods permanently discard audio KV positions once evicted, with no pathway to recover them during decoding. We show…
11 -
arXiv — NLP / Computation & Language research 5d ago
A Source-Grounded Framework for Constructing and Evaluating Progressive Multimodal Diagnostic Dialogues from Clinical Case Reports
arXiv:2608.22713v1 Announce Type: new Abstract: Clinical diagnosis requires progressive integration of patient history, physical examination, laboratory findings, medical images, and diagnostic-informative tests. However, most multimodal medical benchmarks evaluate fixed inputs…
38 -
arXiv — NLP / Computation & Language research 5d ago
DiaRelay: Relaying Dialogue Context with a Constant-Size Memory for Emotion Recognition in Conversation
arXiv:2608.22745v1 Announce Type: new Abstract: Emotion Recognition in Conversation (ERC) requires models to identify subtle emotional cues that are often distributed across distant dialogue turns. Existing methods typically incorporate dialogue history through a fixed context…
20 -
arXiv — NLP / Computation & Language research 5d ago
Beyond Factual Knowledge: Benchmarking and Learning Step-Level Procedural Rule Reasoning in Large Language Models
arXiv:2608.22753v1 Announce Type: new Abstract: Large language models (LLMs) excel at text understanding and generation, yet still struggle to reliably understand and apply externally provided procedural rules at scale. To evaluate this capability, we introduce RuleWorld, a…
24 -
arXiv — NLP / Computation & Language research 5d ago
XTC: Head-Aware Sampling by Excluding Top Choices
arXiv:2608.22758v1 Announce Type: new Abstract: Standard decoding rules for autoregressive language models promote diversity by rescaling the full next-token distribution or truncating its low-probability tail. These strategies overlook a common regime of open-ended generation…
22 -
arXiv — NLP / Computation & Language research 5d ago
Don't Repeat Yourself: Stopping Verbatim Loops at Sampling Time
arXiv:2608.22761v1 Announce Type: new Abstract: Large Language Models generate text autoregressively, but open-ended generation is prone to verbatim looping, in which models repeat spans already present in context. Standard defenses such as repetition, presence, and frequency…
26 -
arXiv — NLP / Computation & Language research 5d ago
DelistBench: Evaluating Search-Enabled LLMs for Auditable Corporate-Event Database Completion
arXiv:2608.22770v1 Announce Type: new Abstract: Financial institutions need an independent way to detect missing, stale, and misclassified corporate-event records in vendor databases. We introduce Search-to-Record, a database-assurance task in which search-enabled large language…
21 -
arXiv — NLP / Computation & Language research 5d ago
SPOC-SQL: Stage-wise Preference Optimization for Controllable Text-to-SQL
arXiv:2608.22772v1 Announce Type: new Abstract: Text-to-SQL aims to translate natural language questions into executable SQL queries over relational databases, requiring multi-stage structured reasoning over database schemas and query constraints. However, existing methods treat…
34 -
arXiv — NLP / Computation & Language research 5d ago
TRACE: A Self-Evolving Skill Bank for Consistent, Limit-Aware LLM Agents
arXiv:2608.22793v1 Announce Type: new Abstract: Reliable deployment of LLM agents in user-facing products depends not on raw task-solving ability but on consistency and limit-awareness: behaving the same way across repeated trials, and recognizing when a request cannot, or…
18 -
arXiv — NLP / Computation & Language research 5d ago
SDoH-Aware Narrative Anchoring Bias in Medical LLMs for Trustworthy Clinical Decision Support
arXiv:2608.22802v1 Announce Type: new Abstract: Medical large language models are often judged by how many clinical questions they answer correctly. That view is useful, but it misses a practical risk. A model may know the right answer and still change its response when the same…
15 -
arXiv — NLP / Computation & Language research 5d ago
DIAG: Diagnostic Iterative Alignment and Generation for Data-Efficient Mathematical Preference Distillation
arXiv:2608.22806v1 Announce Type: new Abstract: Iterative preference optimization is essential for aligning Large Language Models on mathematical reasoning tasks, yet its efficiency is often throttled by signal scarcity: as the model improves, static problem sets become…
28 -
arXiv — NLP / Computation & Language research 5d ago
Industrial-Instruction: An End-to-End Framework for Building Instruction-Tuning and Benchmark Datasets from Industrial Technical Reports
arXiv:2608.22817v1 Announce Type: new Abstract: Industrial technical reports contain high-value knowledge for maintenance, troubleshooting, and product engineering, but their heterogeneous structure (dense prose, specifications, tables) makes them difficult to index and reason…
30 -
arXiv — NLP / Computation & Language research 5d ago
SAVER: Selective Auditing of Verbal Evidence for Error Recovery in VLM Change Reasoning
arXiv:2608.22857v1 Announce Type: new Abstract: Vision-language models (VLMs) frequently fail at visual change reasoning, even when their vision encoders contain sufficient information. We observe that correct VLM outputs tend to contain explicit verbal evidence (object names,…
20 -
arXiv — NLP / Computation & Language research 5d ago
Better Retrieval, Worse Robustness:How Multi-hop RAG Amplifies Upstream ASR Errors
arXiv:2608.22872v1 Announce Type: new Abstract: Speech-based applications pass spoken queries through automatic speech recognition (ASR) before any retrieval module, so ASR errors enter the pipeline as a fixed upstream constraint. We empirically test whether two extensions to…
33 -
arXiv — NLP / Computation & Language research 5d ago
AraDetox: A Multi-Dialect Arabic Detoxification Dataset
arXiv:2608.22894v1 Announce Type: new Abstract: Arabic harmful-language detection has received considerable attention, yet Arabic text detoxification remains underexplored. We introduce AraDetox, a multi-dialect Arabic detoxification dataset comprising 10,500 harmful…
16 -
arXiv — NLP / Computation & Language research 5d ago
SelFusion: Self-distillation for Diffusion Language Models
arXiv:2608.22898v1 Announce Type: new Abstract: Diffusion language models (DLMs) alleviate the inherent latency bottleneck of autoregressive (AR) large language models (LLMs), but their degraded generation quality limits practical applicability. Although knowledge distillation…
38 -
arXiv — NLP / Computation & Language research 5d ago
Do Spoken Language Models Hear Speech as They Read Text? Bridging Structural Gaps Between Speech and Text
arXiv:2608.22908v1 Announce Type: new Abstract: Spoken Language Models (SLMs) generate textual responses directly from speech, offering an alternative to cascaded systems. Despite recent advances, existing SLMs still exhibit weaker instruction-following behavior and limited…
10 -
arXiv — NLP / Computation & Language research 5d ago
Exploring Dowker Homology for Sentence Similarity
arXiv:2608.22909v1 Announce Type: new Abstract: Dowker homology is a topological tool that may be used to analyze the relative position of two point clouds living in a common space. We investigate whether Dowker homology captures sentence similarity information by treating the…
22 -
arXiv — NLP / Computation & Language research 5d ago
Knowing Isn't Always Saying: When Do Spatial Encodings Reach Answers in Vision-Language Models?
arXiv:2608.22916v1 Announce Type: new Abstract: Vision-language models are known to encode spatial information in their hidden states, yet often fail to use it when answering. However, it remains unclear when and where this encoded information reaches the answer. We address this…
33 -
arXiv — NLP / Computation & Language research 5d ago
TSWAP: A Multilingual Retrieval-Augmented Thai Wellness Advisor
arXiv:2608.22917v1 Announce Type: new Abstract: We present TSWAP, a deployed eight-language conversational wellness advisor grounded, via retrieval-augmented generation, in a verified knowledge base of Thai traditional medicine and certified wellness providers. An unmodified…
30 -
arXiv — NLP / Computation & Language research 5d ago
HelaBERT: Enhancing Sinhala Language Understanding with Dual Pooling Classification Head
arXiv:2608.22922v1 Announce Type: new Abstract: We present HelaBERT, a family of two BERT-based masked language models pre-trained from scratch on approximately 1 billion tokens of Sinhala text sourced from MADLAD-400, CulturaX, and a custom corpus comprising news articles,…
20 -
arXiv — NLP / Computation & Language research 5d ago
What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation
arXiv:2608.22948v1 Announce Type: new Abstract: Large language models are increasingly used to propose research ideas, yet the prevailing ways of judging such ideas supply no shared decision rule: free-form judging sways with style and position, and scoring against a later paper…
27 -
arXiv — NLP / Computation & Language research 5d ago
The Illusion of Control: Why Bare Classifier Inversion Silently Fails in Concept-Bottleneck Text Generation
arXiv:2608.22956v1 Announce Type: new Abstract: Concept-bottleneck controllable generation routes multi-attribute control through a low-dimensional concept code that, at deployment, must be synthesised from a target attribute configuration. We study this problem in…
33