arXiv — NLP / Computation & Language
500 articles archived · Visit source ↗ · RSS
-
arXiv — NLP / Computation & Language research 9d ago
Transformer Models for Text Summarization: A Comparative Study of BART, BERT, and RoBERTa
arXiv:2608.19200v1 Announce Type: new Abstract: Text summarization refers to the task of condensing a document into a shorter version while preserving its key information. Automatic text summarization (ATS), driven by advancements in natural language processing (NLP), has…
10 -
arXiv — NLP / Computation & Language research 9d ago
Automatic bioinformatic software named entity recognition from literature
arXiv:2608.19201v1 Announce Type: new Abstract: Bioinformatics software and databases are essential components of modern life science research, yet their mentions in the scientific literature are often inconsistent and difficult to systematically identify at scale. The lack of a…
27 -
arXiv — NLP / Computation & Language research 9d ago
Asymmetric Attention Heads: Structured Head-Wise Context Allocation for Transformer Attention
arXiv:2608.19203v1 Announce Type: new Abstract: Standard multi-head attention (MHA) gives every head the same full causal context span, although heads can serve different contextual roles. Some heads may rely mainly on nearby lexical or syntactic context, while others may depend…
10 -
arXiv — NLP / Computation & Language research 9d ago
Hallucination as a Feature, not a Defect: Evaluating a multi-agent architecture to transform speculative language-model outputs into testable scientific hypotheses
arXiv:2608.19206v1 Announce Type: new Abstract: Contemporary Large Language Models (LLMs) are increasingly aligned to suppress hallucinations, prioritizing factual retrieval over combinatorial creativity. While crucial for mitigating misinformation, this alignment may also…
29 -
arXiv — NLP / Computation & Language research 9d ago
Compliance, Capability, and Conflict: Benchmarking Multimodal LLMs under System Messages
arXiv:2608.19207v1 Announce Type: new Abstract: Production deployments of Multimodal Large Language Models (MLLMs) increasingly rely on system messages to govern model behavior. Yet existing benchmarks either evaluate constraints in text only or embed them into the user turn,…
30 -
arXiv — NLP / Computation & Language research 9d ago
When Irrelevant Text Matters: Affine Margin Shifts in Multimodal Large Language Models
arXiv:2608.19208v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are frequently exposed to auxiliary textual context, the impact of which on visually grounded tasks remains underexplored. In this paper, we investigate the influence of task-irrelevant…
31 -
arXiv — NLP / Computation & Language research 9d ago
Represented but Ignored: A Causal Account of Prosodic Underuse in Audio-Language Models
arXiv:2608.19211v1 Announce Type: new Abstract: Human speech is richly expressive, with prosody carrying linguistic and emotional information beyond the lexical content. A capable large audio-language model (audio-LLM) should therefore support expressive speech understanding,…
4 -
arXiv — NLP / Computation & Language research 9d ago
NepOOC-M: Bilingual Nepali-English Benchmark and Comparative Analysis of Multimodal Architectures for OOC Detection
arXiv:2608.19212v1 Announce Type: new Abstract: Out-of-context (OOC) misinformation pairs authentic images with misleading captions to construct false narratives without image manipulation, making detection a problem of multimodal alignment rather than image forensics. Despite…
20 -
arXiv — NLP / Computation & Language research 9d ago
Time-Series Retrieval for Grounding Multimodal Language Models in Remaining Useful Life
arXiv:2608.19218v1 Announce Type: new Abstract: Large language models (LLMs) and agentic AI systems are increasingly being explored for domain-specific maintenance and prognostics tasks, raising the question of whether they can effectively support prognostics and health…
31 -
arXiv — NLP / Computation & Language research 9d ago
Can Conversational AI loosen Us-Versus-Them Boundaries? The Effects of Common, Dual, and Separate Identity Framings on Pro-Immigrant Intergroup Helping
arXiv:2608.19220v1 Announce Type: new Abstract: Rising immigration has intensified intergroup tensions in many countries. Traditional bias-reduction programs remain difficult to scale and increasingly constrained by U.S. policy. This preregistered experiment tested whether…
9 -
arXiv — NLP / Computation & Language research 9d ago
A Speech Corpus for Mizo Automatic Speech Recognition: Whisper and SraVaani 1.0 Fine-Tuning with Morphology-Aware Evaluation
arXiv:2608.19361v1 Announce Type: new Abstract: This study reports the development of an Automatic Speech Recognition (ASR) system in Mizo, a low-resource language. The development included collecting 17.62 hours of speech data, curating it, and fine-tuning the Mizo ASR system…
4 -
arXiv — NLP / Computation & Language research 9d ago
Linguistic Holonomy and Statistical Watermarks: Inner Geometry of Meaning-Preserving Transformations
arXiv:2608.19369v1 Announce Type: new Abstract: Statistical watermarks for language models live in the freedom of the signifier: they choose among tokens that are nearly equivalent in meaning, and they are therefore eroded by exactly those transformations which move the form of…
31 -
arXiv — NLP / Computation & Language research 9d ago
Are LLMs becoming similarly creative? Evidence from three years of models
arXiv:2608.19437v1 Announce Type: new Abstract: Many benchmarks track Large Language Model (LLM) performance on tasks with verifiable answers, but less is known about how LLM performance is evolving on open-ended tasks, where creativity, originality and diversity may matter as…
8 -
arXiv — NLP / Computation & Language research 9d ago
SynFlow: A Multidimensional Diachronic Semantic Analysis Toolkit
arXiv:2608.19472v1 Announce Type: new Abstract: Lexical semantic change (LSC) is commonly modelled through vector-space representations, but these approaches often provide limited insight into which aspects of usage are changing. Diachronic corpus research instead examines…
11 -
arXiv — NLP / Computation & Language research 9d ago
Hear2Act: Benchmarking When Prosody Should Change What an Assistant Does
arXiv:2608.19515v1 Announce Type: new Abstract: Prosodic cues can convey task-relevant information that alters the trajectory and outcome of a task-oriented dialogue, even when the words themselves remain unchanged. Yet existing benchmarks typically evaluate prosodic perception,…
36 -
arXiv — NLP / Computation & Language research 9d ago
Automated Summarization of Financial News Using Large Language Models and Retrieval-Augmented Generation: An Early Empirical Study (Fall 2023)
arXiv:2608.19526v1 Announce Type: new Abstract: Stock market analysts and investors face a daily challenge: too much financial news, too little time. Manually reading and synthesizing hundreds of company-specific articles is impractical, yet missing key information can directly…
38 -
arXiv — NLP / Computation & Language research 9d ago
When Machines Speak: A Unified Generative Framework for Integrating Machine-Native Symbols into Pretrained Large Language Models
arXiv:2608.19529v1 Announce Type: new Abstract: Many real-world AI systems represent entities, behaviors, and structured information using discrete machine-native symbols rather than natural language. While these representations are compact and preserve task-relevant structure,…
37 -
arXiv — NLP / Computation & Language research 9d ago
Generating Diverse Personas for User Simulators to Test Interview Dialogue Systems
arXiv:2608.19549v1 Announce Type: new Abstract: This paper addresses the issue of the significant labor required to test interview dialogue systems. While interview dialogue systems are expected to be useful in various scenarios, like other dialogue systems, testing them with…
35 -
arXiv — NLP / Computation & Language research 9d ago
Reliable Financial Named Entity Recognition under Domain Shift
arXiv:2608.19558v1 Announce Type: new Abstract: Financial AI systems often train information extractors on one textual register and deploy them across filings, news, and user-generated content, while standard F1 scores do not indicate which predictions remain safe to automate…
13 -
arXiv — NLP / Computation & Language research 9d ago
Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM Agents
arXiv:2608.19564v1 Announce Type: new Abstract: Persistent memory can personalize an LLM agent, but an incorrect durable update can silently distort future behavior. We study the memory-clarification boundary: whether interaction-derived information should be persisted, used…
17 -
arXiv — NLP / Computation & Language research 9d ago
Forking Fast: Efficiently Estimating Uncertainty Dynamics in Text Generation
arXiv:2608.19611v1 Announce Type: new Abstract: LLM reasoning is stochastic, and so understanding a model requires grappling with the distribution of reasoning chains that it might produce for a given question, i.e., its uncertainty. Resampling-based analyses characterize this…
13 -
arXiv — NLP / Computation & Language research 9d ago
Mitigating Identity Essentialism in LLM Agents with Longitudinal Life Trajectories
arXiv:2608.19621v1 Announce Type: new Abstract: Large language models (LLMs) offer a scalable approach to social simulation, but their credibility depends on how agents are constructed. Existing methods can partially reproduce population-level patterns, yet often fail to capture…
30 -
arXiv — NLP / Computation & Language research 9d ago
ReCache: Efficient KV Cache Reuse and Compression for Tool-Augmented LLM Agents
arXiv:2608.19662v1 Announce Type: new Abstract: Agentic language models repeatedly encode tool and skill schemas that recur across requests in different combinations and orders, preventing standard prefix caching from reusing their key--value (KV) states. We introduce…
36 -
arXiv — NLP / Computation & Language research 9d ago
The Asymmetric Harms of LLM Compression
arXiv:2608.19670v1 Announce Type: new Abstract: Large language models (LLMs) compression reduces deployment costs, but standard aggregate metrics like perplexity and accuracy often mask underlying behavioral shifts. In this work, we systematically evaluate 3 LLMs across 11…
6 -
arXiv — NLP / Computation & Language research 9d ago
Projector Is All You Train
arXiv:2608.19726v1 Announce Type: new Abstract: The typical training process of a multimodal large language model (MLLM) involves adapting both the language model backbone and the projector between the backbone and a modality-specific encoder. We ask whether fine-tuning the…
9 -
arXiv — NLP / Computation & Language research 9d ago
One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows
arXiv:2608.19741v1 Announce Type: new Abstract: Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation, app APIs, and function calling. Yet completing consequential work beyond code requires more than producing a…
23 -
arXiv — NLP / Computation & Language research 9d ago
PersonalBench: Measuring the Authorship Gap in LLM Personalization
arXiv:2608.19746v1 Announce Type: new Abstract: Personalized text generation aims to make LLMs write in a specific individual's style, yet existing benchmarks measure task accuracy or preference alignment rather than whether the model's output actually resembles the target…
19 -
arXiv — NLP / Computation & Language research 9d ago
FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving
arXiv:2608.19758v1 Announce Type: new Abstract: Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic complexity of attention remains a critical bottleneck, particularly during the compute-intensive prefilling phase. Our previous work,…
33 -
arXiv — NLP / Computation & Language research 9d ago
SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?
arXiv:2608.19799v1 Announce Type: new Abstract: Software increasingly functions as part of the scientific instrument itself, making failures in scientific code capable of compromising not only program behavior but also the evidence underlying scientific conclusions. Yet existing…
11 -
arXiv — NLP / Computation & Language research 9d ago
LoRA-GA$^2$: Low Rank Adaptation with Multi-step Gradient Adaptive Alignment
arXiv:2608.19800v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) is a prominent fine-tuning method for large models, achieving competitive performance with reduced memory overhead. However, a persistent performance gap remains between LoRA and full fine-tuning. Recent…
11 -
arXiv — NLP / Computation & Language research 9d ago
Stopping and Routing LLM Judge Panels
arXiv:2608.19802v1 Announce Type: new Abstract: LLM evaluation pipelines often have many candidate judges: general LLM-as-a-judge prompts, reward models, safety classifiers, confidence variants, and task-specific verifiers. The deployment question is not only which judge is…
20 -
arXiv — NLP / Computation & Language research 9d ago
A knowledge-guided agentic framework for mitigating patient-context ambiguity in health queries
arXiv:2608.19875v1 Announce Type: new Abstract: Patients often submit short, underspecified queries to healthcare chatbots that lack the patient-specific information needed to determine an appropriate response. Although these queries may be linguistically clear, they can support…
37 -
arXiv — NLP / Computation & Language research 9d ago
Interrupting the Loop: Periodic Subject Changes Raise Judged Surprise and Connection in Base Language Models
arXiv:2608.19893v1 Announce Type: new Abstract: Where does the novelty a base language model produces with no task come from, and what can an LLM judge of a long stream actually see? We dismantle a cognitively inspired generation loop over 24 conditions on three base models.…
19 -
arXiv — NLP / Computation & Language research 9d ago
Learning how to Forget: Fine-tuning for Long-Context Sparse Attention
arXiv:2608.19920v1 Announce Type: new Abstract: A lot of prior work addressed key-value (KV) cache selection and compression by sparse attention to enable long-context inference for transformer language models without excessive hardware budgets. We provide a new method for…
18 -
arXiv — NLP / Computation & Language research 9d ago
Dynamic Gated Cross-Modal Fusion with Sarcastic-aware Contrastive Regularization for Multimodal Sarcasm Detection
arXiv:2608.19942v1 Announce Type: new Abstract: Multimodal sarcasm detection aims to identify sarcastic intent from multimodal content, where inconsistencies between literal meaning and contextual cues often signal irony. This task has attracted increasing research attention.…
31 -
arXiv — NLP / Computation & Language research 9d ago
Natural Language Code Retrieval for 1C:Enterprise: An Open Benchmark and Efficient Bi-Encoder
arXiv:2608.19957v1 Announce Type: new Abstract: Natural language code retrieval is a rapidly evolving task in computer science. However, the 1C:Enterprise ecosystem combines Russian syntax with highly domain-specific terminology, for which open datasets and specialized models…
12 -
arXiv — NLP / Computation & Language research 9d ago
Robust Incomplete Multimodal Sentiment Analysis via Iterative Proxy Correction
arXiv:2608.19971v1 Announce Type: new Abstract: Multimodal sentiment analysis aims to infer affective states by integrating language, visual, and acoustic cues. However, real-world multimodal inputs are often incomplete or corrupted, which can weaken cross-modal complementarity…
30 -
arXiv — NLP / Computation & Language research 9d ago
HealMed: Multilingual Evaluation of Large Language Models in Medicine
arXiv:2608.19981v1 Announce Type: new Abstract: We present HealMed, an expert-reviewed benchmark for multilingual evaluation of large language models in medicine. HealMed contains 1,000 examples in each of nine languages, drawn from nine datasets and covering three task formats:…
27 -
arXiv — NLP / Computation & Language research 9d ago
Auditing Cross-Lingual Fairness in Language Model Watermarking
arXiv:2608.20047v1 Announce Type: new Abstract: Watermarking schemes for large language model output are evaluated almost exclusively on English text using each scheme's detection threshold and a narrow set of quality measurements. Multilingual deployment exposes…
9 -
arXiv — NLP / Computation & Language research 9d ago
SABET-QA: Temporal Knowledge Graph Question Answering
arXiv:2608.20083v1 Announce Type: new Abstract: Question Answering over Temporal Knowledge Graphs (TKGQA) requires reasoning over time-sensitive facts, yet existing embedding-based methods struggle with multi-step queries due to single-pass reasoning pipelines. We propose…
21 -
arXiv — NLP / Computation & Language research 9d ago
OenoBench: A Wine-Domain Benchmark for Knowledge-Grounded Evaluation of Large Language Models
arXiv:2608.20106v1 Announce Type: new Abstract: We introduce OenoBench, a wine-domain knowledge benchmark of 3,266 multiple-choice questions across six pillars (regions, grape varieties, viticulture, winemaking, producers, business) and four difficulty tiers. The corpus is built…
12 -
arXiv — NLP / Computation & Language research 9d ago
When Text and Numbers Disagree: Evidence Arbitration in Large Language Models
arXiv:2608.20116v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in settings where textual summaries, numerical observations, and external tool outputs may provide conflicting evidence. We study how LLMs arbitrate between such sources when they…
19 -
arXiv — NLP / Computation & Language research 9d ago
FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models
arXiv:2608.20153v1 Announce Type: new Abstract: Large language models (LLMs) have shown growing potential for automated theoretical computer science (TCS) research, yet existing benchmarks remain far from realistic research settings. We introduce \ourbenchmark, an…
38 -
arXiv — NLP / Computation & Language research 9d ago
Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection
arXiv:2608.20169v1 Announce Type: new Abstract: We present a novel approach to efficient LLM agent harness optimization through adaptive validation task selection. Harness optimization iteratively rewrites the harness code based on validation performance, enabling substantial…
24 -
arXiv — NLP / Computation & Language research 9d ago
Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization
arXiv:2608.20281v1 Announce Type: new Abstract: Large language models often fail to answer questions about a bounded document collection when the source documents are not retrieved at inference time. We study this setting as document knowledge internalization: converting a fixed…
21 -
arXiv — NLP / Computation & Language research 9d ago
Inducing Task Models from Computer-Use Traces
arXiv:2608.20319v1 Announce Type: new Abstract: Naturalistic computer-use traces, passively recorded screenshots and mouse or keyboard actions, are a valuable resource for deriving symbolic, auditable, and reusable models of how everyday work is done. Such models matter as…
8 -
arXiv — NLP / Computation & Language research 9d ago
G-CARL: Grounded Checklist-Aligned Reward Learning for Patient-Oriented Medical Report Interpretation
arXiv:2608.20331v1 Announce Type: new Abstract: Personalized interpretation of medical reports has emerged as an increasingly important need among patients. Addressing this need requires both evidence-grounded medical factuality and context-dependent patient communication, yet…
34 -
arXiv — NLP / Computation & Language research 9d ago
ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models
arXiv:2608.20338v1 Announce Type: new Abstract: Large Language Models (LLMs) increasingly require selective removal of harmful or sensitive knowledge, called unlearning, yet existing methods and benchmarks fail to evaluate this capability completely. Current approaches rely on…
20 -
arXiv — NLP / Computation & Language research 9d ago
Active Inference as Context Acquisition for AI Agents
arXiv:2608.19202v1 Announce Type: cross Abstract: Interactive AI agents must acquire the right context as efficiently as possible. When a user omits a constraint, preference, file, or task variable, an agent can proceed with a default assumption or spend tokens on a clarifying…
4 -
arXiv — NLP / Computation & Language research 9d ago
Outcome Monitors: Recovery Affordances for Silent Tool Failures
arXiv:2608.19303v1 Announce Type: cross Abstract: When a tool call times out, the agent sees the failure and can route around it. A cached error page or negative price can instead arrive in the expected format and be consumed as fact. We introduce Outcome Monitors, which detect…
38