arXiv — NLP / Computation & Language
500 articles archived · Visit source ↗ · RSS
-
arXiv — NLP / Computation & Language research 6d ago
Beyond Raw Transcripts: Structured Persona Extraction for LLM-Based Digital Twins
arXiv:2608.20344v1 Announce Type: new Abstract: LLM-based "digital twins" aim to simulate how an individual would behavein new environments or respond to novel questions, given some representation of that individual's prior responses. A common approach constructs this…
27 -
arXiv — NLP / Computation & Language research 6d ago
When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha
arXiv:2608.20345v1 Announce Type: new Abstract: Conversational AI systems have become informal mental health support resources for Generation Alpha (Gen Alpha, born 2010-2024), with 13.1% of U.S. adolescents (5.4 million) using generative AI for mental health advice. While these…
5 -
arXiv — NLP / Computation & Language research 6d ago
Building and Evaluating a Synthetic Bengali Speech Resource for Telecom Customer Care
arXiv:2608.20346v1 Announce Type: new Abstract: Speech systems used in customer-facing applications often require domain-specific language coverage. We present a synthetic Bengali speech dataset for telecom customer-care scenarios. The dataset contains 10,000 audio-text pairs,…
8 -
arXiv — NLP / Computation & Language research 6d ago
Who Do Language Models Think Is Competent? A Mechanistic Analysis of Occupational Bias
arXiv:2608.20347v1 Announce Type: new Abstract: Language models (LMs) often pass behavioral bias evaluations, but it remains unclear whether they no longer represent the underlying associations that give rise to biases, or have merely learned not to express them. In this study,…
36 -
arXiv — NLP / Computation & Language research 6d ago
Inhibitory Attention for Clinical Long-Context Reasoning: Characterizing and Mitigating Lost-in-the-Middle Effects in EHR Processing
arXiv:2608.20348v1 Announce Type: new Abstract: Electronic health records now routinely exceed 100,000 tokens per patient. Yet large language models exhibit the lost-in-the-middle (LitM) effect: information near the center of a long context is retrieved less reliably than…
19 -
arXiv — NLP / Computation & Language research 6d ago
Beyond Prompt Engineering: A Systematic Analysis of Prompt Lexical Sensitivity and Its Impacts on Quality
arXiv:2608.20349v1 Announce Type: new Abstract: Large Language Models (LLMs) exhibit extreme sensitivity to surface-level prompt variations, in which minor lexical changes can trigger disproportionate performance fluctuations. Moving beyond black-box optimization and…
10 -
arXiv — NLP / Computation & Language research 6d ago
How to Train a Real-World Silicon Concierge? Internalizing Complex Business Workflow to Only OneModel
arXiv:2608.20350v1 Announce Type: new Abstract: Traditional industrial agents rely on modular pipelines, including Router, Retriever, Planner, Executor, Responder, Reviewer, and other components. These systems often fracture into a labyrinth of ad-hoc patches, leading to…
8 -
arXiv — NLP / Computation & Language research 6d ago
Exploratory As-Analyzed No-Detection of Culturally-Marked Predicate-Triggered PII Amplification in a Synthetic-English RAG Probe: A Predicate-Resource-Confounded Audit
arXiv:2608.20351v1 Announce Type: new Abstract: We ask whether stereotype-loaded queries about culturally marked people leak more personal information from a retrieval-augmented generation (RAG) system than otherwise-equivalent neutral queries. We pre-register a four-culture…
10 -
arXiv — NLP / Computation & Language research 6d ago
The Divergence Hypothesis: Unmasking Lexical Interference and Label Bias in Mental Health NLP
arXiv:2608.20353v1 Announce Type: new Abstract: Computational mental health (CMH) classifiers often degrade under distribution shift because human annotators and distant-supervision pipelines reward different linguistic signals. We introduce TSS (Triple-Stream Stress probe), a…
24 -
arXiv — NLP / Computation & Language research 6d ago
ExpertIVS: Sociological Expert Driven Individual Value Simulation in Large Language Models
arXiv:2608.20355v1 Announce Type: new Abstract: Large Language Model (LLM) agents have demonstrated considerable potential for social simulation, yet struggle to accurately model individual value systems. Most existing methods mechanically stitch survey responses into prompts,…
27 -
arXiv — NLP / Computation & Language research 6d ago
Self-Speculation for Faster Reasoning Models
arXiv:2608.20359v1 Announce Type: new Abstract: Large language models (LLMs) are deployed for increasingly complex tasks involving planning and multi-step decision making, but high-quality performance on these tasks often requires generating long reasoning traces. This is a poor…
9 -
arXiv — NLP / Computation & Language research 6d ago
TriPLU: Bypassing the Gate with Direct Trilinear Product FFNs in Tiny Language Models
arXiv:2608.20360v1 Announce Type: new Abstract: We study whether tiny decoder-only language models benefit from feed-forward layers that directly multiply learned feature projections. TriPLU, a Trilinear Product Linear Unit, replaces the usual gated FFN branch with a…
32 -
arXiv — NLP / Computation & Language research 6d ago
Toward Auto-Research: Mining Falsifiable Research Ideas from Paper Knowledge Graphs with Categorical Structure
arXiv:2608.20361v1 Announce Type: new Abstract: Automated research-idea generation systems built on large language models (LLMs) share a structural weakness: they reduce ideation to free-text recombination, random paper pairing, or embedding-similarity retrieval. The three…
25 -
arXiv — NLP / Computation & Language research 6d ago
Multilingual Verifier Bias in RLVR: Benchmark, Rollout Diagnosis, and the Cross-Lingual Selection Bottleneck
arXiv:2608.20362v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) is a standard recipe for training large language models on mathematical reasoning, where an answer verifier serves as a language-neutral reward function. We show that this…
33 -
arXiv — NLP / Computation & Language research 6d ago
Hadith computational science in the age of large language models: a critical narrative review
arXiv:2608.20364v1 Announce Type: new Abstract: We examine how hadith computational science is being reshaped by transformer models, retrieval-grounded pipelines, and large language models (LLMs). Recent reviews document growth in the literature, but they do not yet provide a…
12 -
arXiv — NLP / Computation & Language research 6d ago
Trilingual Topic Modeling of Sri Lankan Parliamentary Debates
arXiv:2608.20365v1 Announce Type: new Abstract: Sri Lankan parliamentary debates (Hansards) constitute a trilingual corpus of speeches in Sinhala, Tamil, and English, including code-mixed content, yet remain inaccessible to standard NLP pipelines due to layout-complex PDFs,…
27 -
arXiv — NLP / Computation & Language research 6d ago
Research Paper Quality Recognition Through Textual Feature Analysis
arXiv:2608.20368v1 Announce Type: new Abstract: Knowledge and innovations are shaped by using the quality and credibility of the scientific research. Yet, distinguishing between impactful, high-quality work and flawed studies remains a challenge. This paper introduces a…
26 -
arXiv — NLP / Computation & Language research 6d ago
ASTAR: Automated induction of STAndardized radiology Reporting templates from large-scale clinical free-text corpora
arXiv:2608.20369v1 Announce Type: new Abstract: Structured reporting converts free-text radiology narratives into queryable data keys, facilitating cohort assembly, longitudinal tracking, and training label generation for medical AI. The prevailing paradigm follows a two-stage…
17 -
arXiv — NLP / Computation & Language research 6d ago
When Do LLMs Replace Fine-Tuned NLU? A Decision Framework for Intent Detection in Production Conversational Systems
arXiv:2608.20371v1 Announce Type: new Abstract: A common claim is that zero-shot large language models (LLMs) can replace fine-tuned NLU classifiers for intent detection. We test this claim head-to-head and find that the honest answer is: it depends on the intent space. On full…
33 -
arXiv — NLP / Computation & Language research 6d ago
An ambiguity taxonomy for evaluating large language model performance on clinical registry abstraction: a multi-site prospective study
arXiv:2608.20373v1 Announce Type: new Abstract: Objective: To evaluate large language model (LLM) performance on unprocessed electronic medical record (EMR) data for clinical registry abstraction. Methods: We evaluated LLM performance answering registry questions for the…
18 -
arXiv — NLP / Computation & Language research 6d ago
VA-DPO: Valence-Arousal Direct Preference Optimization for Controllable Emotion Generation in Language Models
arXiv:2608.20374v1 Announce Type: new Abstract: How precisely can we tell a language model how to feel? Most work on emotional generation answers with a discrete label - happy, angry, sad - which cannot express a target like "mildly downcast but calm." We instead specify the…
17 -
arXiv — NLP / Computation & Language research 6d ago
GRAFT: Adaptive DLM-Based Draft Tree Construction with Target-Distilled Edge Scoring
arXiv:2608.20375v1 Announce Type: new Abstract: Tree-based speculative decoding raises the mean accepted tokens of standard speculative decoding by verifying multiple draft paths, and existing tree builders typically construct these paths through parent-conditioned expansion,…
13 -
arXiv — NLP / Computation & Language research 6d ago
TH-GNN: Heterogeneous Temporal Graph Neural Networks for LLM-Agent Shilling Attack Detection
arXiv:2608.20376v1 Announce Type: new Abstract: LLM agents can now generate realistic shilling profiles, fluent reviews, and coherent ratings at scale, systematically defeating recommender-system defenses. Text-only detectors that flag semantic drift in review embeddings are…
27 -
arXiv — NLP / Computation & Language research 6d ago
EditPPT: Faithful Long-Deck Slide Editing via Structured Tool-Using Multi-Agent with Dual-Modal Validators
arXiv:2608.20381v1 Announce Type: new Abstract: Automating slide editing requires simultaneously satisfying modification accuracy, preservation fidelity, and robustness to deck length. Existing LLM-based systems often fail on real-world presentation files because they rely on…
30 -
arXiv — NLP / Computation & Language research 6d ago
Decoupled Vision-Language System for Multimodal Understanding and Generation
arXiv:2608.20382v1 Announce Type: new Abstract: We introduce a new architecture design for multimodal large language models (MLLMs), Libra, capable of both multimodal understanding and generation. Libra architecture contains one vision system and one language system, connected…
23 -
arXiv — NLP / Computation & Language research 6d ago
Using Human-LLM Disagreement to Improve Checklist-Based Quality Appraisal
arXiv:2608.20385v1 Announce Type: new Abstract: Systematic reviews rely on quality appraisal of included studies, a process that is time-consuming and sensitive to ambiguity in checklist criteria. Although large language models (LLMs) offer opportunities to support these tasks,…
20 -
arXiv — NLP / Computation & Language research 6d ago
Poly-InstructTTS: Learning In-the-Wild Expressive Speech Synthesis from Open-Ended Instructions
arXiv:2608.20387v1 Announce Type: new Abstract: While recent text-to-speech (TTS) models achieve high naturalness, controlling fine-grained expression via natural-language instructions remains challenging. We introduce Poly- InstructTTS, which learns expressive speech from…
9 -
arXiv — NLP / Computation & Language research 6d ago
Intent Engine: Natural-Language Intent Translation for Intent-Driven Orchestration in the Compute Continuum
arXiv:2608.20388v1 Announce Type: new Abstract: Microservice placement in the compute continuum is driven by low-level Service-level Objectives (SLOs), but requiring users to specify metric-level constraints creates an adoption barrier and increases misconfiguration risk.…
37 -
arXiv — NLP / Computation & Language research 6d ago
Ansari: A Retrieval-Grounded Islamic AI Assistant -- Architecture, Deployment, and Lessons from 140,000 Conversations
arXiv:2608.20390v1 Announce Type: new Abstract: General-purpose large language models (LLMs) are increasingly used to answer religious questions, but for Islamic content they carry two serious risks: factual fabrication (inventing Qur'anic verses or hadith) and subtle value…
13 -
arXiv — NLP / Computation & Language research 6d ago
ImmigrationReason: A Structured Dataset of U.S. Immigration Appeals for Legal Reasoning Research
arXiv:2608.20391v1 Announce Type: new Abstract: Most legal NLP resources draw from federal case law and focus on coarse classification, leaving administrative adjudication, where the vast majority of government decisions occur, essentially unaddressed. We introduce…
10 -
arXiv — NLP / Computation & Language research 6d ago
Evaluation-as-Search: Adaptive Discovery of Grounding Failures in Meeting Assistants
arXiv:2608.20392v1 Announce Type: new Abstract: LLM-powered meeting assistants are deployed at scale, yet systematic evaluation of their grounding fidelity remains limited to static benchmarks that miss failure modes tied to specific discourse structures or reasoning demands. We…
14 -
arXiv — NLP / Computation & Language research 6d ago
Knowledge-Graph-Gated Defactualization for Style-Controllable and Fact-Preserving Generation in Agentic Conversational AI
arXiv:2608.20393v1 Announce Type: new Abstract: Agentic large language models (LLMs) deployed in fact-sensitive applications such as customer support must simultaneously preserve factual correctness and generate responses in a controllable stylistic register. Activation steering…
7 -
arXiv — NLP / Computation & Language research 6d ago
Self-Supervised Speech Representations Track Spoken Language Convergence to Adult Models in Infants and Children Who Are Deaf/Hard-of-Hearing
arXiv:2608.20396v1 Announce Type: new Abstract: Language development is characterized by a gradual convergence of children's speech toward adult patterns. Measuring this process has traditionally required detailed transcription and language-specific expertise, limiting…
4 -
arXiv — NLP / Computation & Language research 6d ago
LingShu: A Large-Scale Symptom-Centric Contextualized Knowledge Graph Bridging Traditional Chinese Medicine and Modern Biomedicine
arXiv:2608.20402v1 Announce Type: new Abstract: Biomedical knowledge graphs (KGs) are pivotal for knowledge organization, yet traditional binary relations often struggle to represent the conditional nature of biomedical knowledge. Symptoms provide a shared phenotypic layer for…
35 -
arXiv — NLP / Computation & Language research 6d ago
ARGUS: Theory-of-Mind Guided Argument Generation with Strategy-Aware Planning and Knowledge Grounding
arXiv:2608.20405v1 Announce Type: new Abstract: Persuasive argument generation requires modeling audience beliefs, rhetorical strategies, and factual grounding. Despite recent advancements, existing methods remain largely audience-agnostic and fail to integrate strategy…
21 -
arXiv — NLP / Computation & Language research 6d ago
LiLiCorr: Lightweight Likelihood Correlation of Parallel Drafts for Speculative Decoding
arXiv:2608.20530v1 Announce Type: new Abstract: Speculative decoding accelerates language-model inference by drafting future tokens that the target model verifies in parallel. A diffusion-style block head such as DFlash is an attractive drafter, predicting an entire block of…
10 -
arXiv — NLP / Computation & Language research 6d ago
JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification
arXiv:2608.20607v1 Announce Type: new Abstract: Panels of inexpensive LLM judges increasingly make accept-or-escalate decisions. In factuality settings, accepting a claim because several reference-free judges agree can create a hidden risk: agreement may reflect shared…
34 -
arXiv — NLP / Computation & Language research 6d ago
When Failures Propagate: Causal Failure Attribution in Agentic Retrieval-Augmented Generation
arXiv:2608.20627v1 Announce Type: new Abstract: Agentic retrieval-augmented generation (RAG) interleaves retrieval, reasoning, and answer generation across multiple hops. A retrieval error at hop 1 can surface only as a wrong answer at hop 3, while later retrieval can also…
20 -
arXiv — NLP / Computation & Language research 6d ago
Sparse Token Routing in Efficient Transformers
arXiv:2608.20632v1 Announce Type: new Abstract: Efficient-transformer research often motivates token pruning and adaptive computation with the claim that not all tokens require equal computational effort. We test this claim end to end using SEWN, a two-stream Transformer that…
24 -
arXiv — NLP / Computation & Language research 6d ago
AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale
arXiv:2608.20634v1 Announce Type: new Abstract: Agents learn to act through interaction with environments, yet the environments used for training are often manually constructed or synthesized around predefined tasks and benchmarks. This task-centric paradigm makes it difficult…
15 -
arXiv — NLP / Computation & Language research 6d ago
MIL-BERT: Classification of Arbitrarily Large Text with Performance and Explanatory Guarantees
arXiv:2608.20636v1 Announce Type: new Abstract: Many text classification decisions are viable based on constituent excerpts alone. Taking inspiration from the field of multiple instance learning, we present an algorithm for training a neural network to classify text by selecting…
10 -
arXiv — NLP / Computation & Language research 6d ago
Directional Contextual Representations for Dependency Relations: Why Cross-Direction Pairing Fails
arXiv:2608.20647v1 Announce Type: new Abstract: Splitting a bidirectional LSTM's contextual representation into a forward-only $F_i$ (strictly a function of tokens $1..i$) and a backward-only $B_i$ (strictly a function of tokens $i..n$) beats either alone and beats a fused…
5 -
arXiv — NLP / Computation & Language research 6d ago
AsmEvo: Agentic Assembly-Level Optimization of AMD GPU Kernels with Functional Equivalence Verification
arXiv:2608.20711v1 Announce Type: new Abstract: High-performance ML systems increasingly rely on GPU kernels whose editable source is unavailable, generated, or too distant from final machine code to expose remaining optimizations. Existing LLM kernel optimizers and autotuners…
31 -
arXiv — NLP / Computation & Language research 6d ago
PSK at WMT 2026 MIST: Task-Specialized QLoRA Adapters for Multilingual Summarization and Question Answering
arXiv:2608.20757v1 Announce Type: new Abstract: We describe the PSK submission to the WMT 2026 Multilingual Instruction Shared Task. Our system uses the 3.35B-parameter Tiny Aya Global model with three QLoRA adapters, one for each task. The adapters are trained on multilingual…
26 -
arXiv — NLP / Computation & Language research 6d ago
Tree-of-Concerns: Hierarchical Multi-Agent Debate for Unstated-Limitation Extraction in Scientific Critique
arXiv:2608.20777v1 Announce Type: new Abstract: As scientific literature grows and papers increasingly under-report limitations, multi-agent LLMs offer a promising approach to systematically uncover these hidden failure modes. Here, we introduce Tree-of-Concerns, a multi-agent…
12 -
arXiv — NLP / Computation & Language research 6d ago
Denoising the Future: Context-Aware Spectral Diffusion for Temporal Knowledge Graph Extrapolation
arXiv:2608.20804v1 Announce Type: new Abstract: Temporal Knowledge Graph (TKG) extrapolation seeks to infer future facts from time-varying relational histories. Recent diffusion-based approaches improve uncertainty modeling through generative denoising, but their aggregated…
19 -
arXiv — NLP / Computation & Language research 6d ago
STAR-OPD: Structured Aspect-Cascade-Aware On-Policy Reward Distillation for ABSA Quadruple Extraction
arXiv:2608.20831v1 Announce Type: new Abstract: Aspect-based sentiment analysis (ABSA) quadruple extraction requires jointly predicting target, aspect, opinion, and sentiment over reviews that often contain multiple fine-grained sentiment tuples. While large chain-of-thought…
34 -
arXiv — NLP / Computation & Language research 6d ago
SAC-Copula: Quality-Preserving Watermarking for Diffusion Language Models via Smooth Correlated Gumbel Fields
arXiv:2608.20839v1 Announce Type: new Abstract: Watermarking diffusion language models (DLMs) requires mechanisms compatible with iterative parallel unmasking rather than autoregressive decoding. Existing sampling-based watermarking methods typically inject position-wise i.i.d.…
14 -
arXiv — NLP / Computation & Language research 6d ago
Ontology-Driven Structural Regularization for Document-Level Relation Extraction
arXiv:2608.20856v1 Announce Type: new Abstract: Document-Level Relation Extraction (DocRE) relies heavily on costly manually annotated datasets, while large distant supervision resources such as DocRED distant remain underexploited due to noise. We show that a critical yet…
12 -
arXiv — NLP / Computation & Language research 6d ago
KREL: Automatic Medical Coding via Knowledge-Guided Reasoning over Clinical Evidence with LLMs
arXiv:2608.20887v1 Announce Type: new Abstract: Automatic Medical Coding (AMC), which assigns standardized International Classification of Diseases (ICD) codes to clinical notes, is essential for medical reimbursement, quality reporting, and clinical research. Existing…
20