arXiv — NLP / Computation & Language
500 articles archived · Visit source ↗ · RSS
-
arXiv — NLP / Computation & Language research 6d ago
ForeDreamer: A Self-Evolving Dual-Agent Memory Architecture for Future Event Prediction
arXiv:2608.20920v1 Announce Type: new Abstract: Open-web future event prediction requires agents to distill reliable signals from noisy, redundant, and incomplete evidence. Existing retrieval/memory mechanisms directly feed retrieved information to agents or rely on simple…
14 -
arXiv — NLP / Computation & Language research 6d ago
Source-Free MT Evaluation Is Not MT Evaluation
arXiv:2608.20925v1 Announce Type: new Abstract: Reference-based metrics remain the standard choice in machine translation evaluation, partly because quality estimation methods often correlate less well with human judgments. As a result, source-free, reference-based evaluation…
7 -
arXiv — NLP / Computation & Language research 6d ago
MentorPulse: Refreshing Cross-Model Latent Guidance for Long-Form Generation
arXiv:2608.20927v1 Announce Type: new Abstract: Cross-model latent guidance lets a frozen large mentor encode an input once and a frozen small student generate from the resulting signal. Existing methods keep this signal fixed, assuming it stays useful as the output grows; we…
35 -
arXiv — NLP / Computation & Language research 6d ago
Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs
arXiv:2608.20953v1 Announce Type: new Abstract: Serving large language models cheaply increasingly means shipping models that are both structurally compressed to a fraction of their parameters and quantized to 4 bits. Together these steps degrade reasoning, mathematics, coding,…
5 -
arXiv — NLP / Computation & Language research 6d ago
Extractive Summarization for Arabic Documents Using SAraBERT with a Semantic Siamese Similarity Evaluation Metric
arXiv:2608.20964v1 Announce Type: new Abstract: In this research, we introduce SAraBERT, an enhanced version of AraBERT which proposes inter-sentence transformer layers for extractive summarization tasks. To ensure that the summaries generated by SAraBERT achieve a high coverage…
21 -
arXiv — NLP / Computation & Language research 6d ago
Target-Aware Calibration Data Selection for Preserving Uncertainty in Quantized Language Models
arXiv:2608.21019v1 Announce Type: new Abstract: Quantization is widely used to deploy large language models, but its effect on uncertainty behavior, such as confidence, margins, and abstention, is rarely treated as a primary objective. We frame calibration-data selection for…
29 -
arXiv — NLP / Computation & Language research 6d ago
Free-Text Evaluation of LLMs for 5G Domain Knowledge and Fault Analysis using LLM-as-Judge
arXiv:2608.21021v1 Announce Type: new Abstract: Real-world fault analysis in 5G and emerging 6G networks demands domain expertise to analyze free-text diagnostics, including root-cause explanations and recommended actions. LLMs have emerged as a promising approach to automating…
16 -
arXiv — NLP / Computation & Language research 6d ago
Scaling Unsupervised Word Alignment to Documents via Structural Constraints
arXiv:2608.21023v1 Announce Type: new Abstract: Word alignment has traditionally been studied between sentences, but many cross-lingual tasks increasingly require correspondences across full documents. While recent multilingual embedding models can encode long inputs, we show…
27 -
arXiv — NLP / Computation & Language research 6d ago
Evidence-Consistent Generative Detection under Scenario-Level Distribution Shift
arXiv:2608.21043v1 Announce Type: new Abstract: Conventional in-distribution evaluation can overestimate robustness when training and test data share recurring task-specific patterns or surface cues. This risk is especially relevant in social-engineering fraud detection, where…
6 -
arXiv — NLP / Computation & Language research 6d ago
PromptResponse: Optimizing Prompts for LLM Coding Tasks
arXiv:2608.21074v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in research workflows and software development pipelines, yet their output remains sensitive to input prompt variations. This paper presents…
35 -
arXiv — NLP / Computation & Language research 6d ago
Jokes Aside: Measuring the Semantic Distance of Double Meanings
arXiv:2608.21087v1 Announce Type: new Abstract: Large language models have significantly enriched the toolkit for computational humor research, particularly in the automated generation of jokes and puns. A key innovation, contextual embedding vectors, offers new opportunities to…
10 -
arXiv — NLP / Computation & Language research 6d ago
When the Feature Pool Goes Algorithmic: Extending Mufwene's Ecology of Language Evolution to LLM-Mediated Exposure
arXiv:2608.21088v1 Announce Type: new Abstract: Mufwene's ecological model locates language evolution in competition among variants contributed by individual idiolects and in speakers' selection from linguistic material made available through interaction. Large language models…
38 -
arXiv — NLP / Computation & Language research 6d ago
No PUN Intended: Plausible Unknown Names for Person-Centred LLM Evaluation
arXiv:2608.21206v1 Announce Type: new Abstract: Person names are widely used as prompt variables in LLM evaluations of factuality, privacy leakage, bias and abstention, but when a name's evidential status is uncontrolled, measurements may conflate memorisation, retrieval, name…
26 -
arXiv — NLP / Computation & Language research 6d ago
RARE: Decoupling Representation Steering from Expert Routing in Mixture-of-Experts Language Models
arXiv:2608.21236v1 Announce Type: new Abstract: Representation engineering offers a lightweight means of controlling language-model behavior by modifying intermediate hidden states, but its direct application to Mixture-of-Experts (MoE) models introduces a structural mismatch.…
21 -
arXiv — NLP / Computation & Language research 6d ago
Affective Context Amplifies Sycophancy in LLM Responses
arXiv:2608.21242v1 Announce Type: new Abstract: As conversational companions, large language models (LLMs) often have access to users' emotional states. We study how this affective context modulates LLM sycophancy in subjective, evaluative interactions, where users share actions…
32 -
arXiv — NLP / Computation & Language research 6d ago
Benchmarking Patent Drafting from Inventor-Style Disclosures
arXiv:2608.21249v1 Announce Type: new Abstract: While recent large language models (LLMs) have achieved promising results on individual patent drafting tasks, they fundamentally fail to investigate the core challenge of real-world patent drafting: generating a complete and…
37 -
arXiv — NLP / Computation & Language research 6d ago
EnSI-RAG: Entity-Structure-Indexed Retrieval-Augmented Generation for Long-Document Question Answering
arXiv:2608.21252v1 Announce Type: new Abstract: Question answering (QA) over long, connected documents remains challenging because relevant evidence may span multiple entities and their relationships. Existing retrieval-augmented generation (RAG) methods typically index…
23 -
arXiv — NLP / Computation & Language research 6d ago
Memory Augmentation Unlocks Efficient Chain-of-Thought Reasoning
arXiv:2608.21265v1 Announce Type: new Abstract: Large language models often rely on Chain-of-Thought (CoT) reasoning to solve complex tasks, but verbose reasoning traces introduce substantial inference overhead. CoT compression shortens generation, yet aggressive compression may…
25 -
arXiv — NLP / Computation & Language research 6d ago
Prompt-Model Interaction Reaches the Fixed Points: A deterministic, task-free structural readout -- and the factorizations of it that failed
arXiv:2608.21315v1 Announce Type: new Abstract: That a prompt's effect is not a property of the prompt is established: prompts optimised for one model degrade on another, and rankings reorder under neutral reformatting. That evidence is about task accuracy, which cannot say…
37 -
arXiv — NLP / Computation & Language research 6d ago
Move by Move: Measuring and Steering How LLMs Conduct Psychotherapy
arXiv:2608.21325v1 Announce Type: new Abstract: Users increasingly turn to large language models for emotional support, yet little is known about how these models actually conduct a psychotherapy interaction. We introduce an ontology of ten therapeutic moves: compact,…
35 -
arXiv — NLP / Computation & Language research 6d ago
A Factorial Ablation of a Speech-to-SFT Pipeline: Differential Effects on Data Quality and Downstream Transfer
arXiv:2608.20394v1 Announce Type: cross Abstract: Industry pipelines that turn speech into supervised fine-tuning (SFT) data via multi-stage refinement are increasingly adopted but, to our knowledge, have not been publicly ablated stage-by-stage, leaving each stage's marginal…
8 -
arXiv — NLP / Computation & Language research 6d ago
When Retrieval Fails Before It Begins: Structurally Indirect Prerequisite Eviction as a Retention Failure in Agentic Memory
arXiv:2608.20400v1 Announce Type: cross Abstract: Agentic memory under a fixed budget involves two stages: retention and retrieval. Existing retrieval-centered paradigms implicitly assume necessary evidence survives eviction, but we challenge this by isolating a pre-retrieval…
8 -
arXiv — NLP / Computation & Language research 6d ago
ProofJudge: Tool-Grounded LLM Evaluation of Formal Proof Quality in Mathlib
arXiv:2608.20432v1 Announce Type: cross Abstract: Formal proofs in Lean 4 that pass the kernel's type checker can nonetheless vary widely in quality. We introduce ProofJudge, an agentic LLM-as-judge system that scores formal proof quality along five dimensions beyond…
7 -
arXiv — NLP / Computation & Language research 6d ago
Open-Weight Masked Introspection: Measuring What Language Models Can Report About Their Own Computation
arXiv:2608.20569v1 Announce Type: cross Abstract: Are frontier models able to introspect about their internal states? Recent work suggests that under certain conditions a complex enough model can audit its own internals, call out what changed, and report back confidently about…
17 -
arXiv — NLP / Computation & Language research 6d ago
Auditable by Construction: An Ontology-Driven Framework for Trustworthy LLM Analytics in Enterprise Finance
arXiv:2608.20661v1 Announce Type: cross Abstract: Enterprise adoption of large language models in finance is constrained less by fluency than by trust: in Financial Planning and Analysis (FP&A) and other regulated workflows, an answer is usable only if it is traceable to…
37 -
arXiv — NLP / Computation & Language research 6d ago
Why2Speak: Faithful Reasoning for Abstaining Action Policies
arXiv:2608.20670v1 Announce Type: cross Abstract: Many agentic systems must repeatedly choose between acting and abstaining, making faithful reasoning important for oversight: an explanation is useful only if it reflects the computation that produced the action. We study this…
25 -
arXiv — NLP / Computation & Language research 6d ago
Temporal Validity on Real Software Histories: Eliminating Stale-Fact Errors in Code-Assistant Memory over GitHub Fixes
arXiv:2608.20685v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) has no model of time: when a fact changes across a coding session - a function is renamed, an endpoint moves, a dependency is bumped - RAG retrieves both the old and new value with…
17 -
arXiv — NLP / Computation & Language research 6d ago
Calibrating Criterion Revision in LLM Agents: Failure Modes and a Trace-Anchored Protocol
arXiv:2608.20729v1 Announce Type: cross Abstract: Language-model agents can improve after failure or carry text across episodes without revising what counts as success. We study the narrower attribution problem of criterion revision: when criterion K0 accepts an outcome…
26 -
arXiv — NLP / Computation & Language research 6d ago
Profiling What Matters: Context-Aware Item Profiles from Large-Scale Metadata for LLM Recommenders
arXiv:2608.20801v1 Announce Type: cross Abstract: While Large Language Models (LLMs) have significantly advanced reranking in recommendation, effectively leveraging item-side information remains challenging. Real-world items are described by vast, heterogeneous, and unstructured…
18 -
arXiv — NLP / Computation & Language research 6d ago
Identify, Locate, Link: End-to-End Key-Value Extraction from Document Images
arXiv:2608.20868v1 Announce Type: cross Abstract: Document processing pipelines traditionally cascade optical character recognition (OCR) engines with downstream models for structured information extraction, leading to multi-stage error propagation. We fine-tune SmolDocling, a…
5 -
arXiv — NLP / Computation & Language research 6d ago
TreeWY: Speculative Verification for Gated DeltaNet Hybrids
arXiv:2608.20961v1 Announce Type: cross Abstract: Modern open models are hybrids: most layers are linear-attention (Gated DeltaNet, GDN) layers carrying a small fixed-size recurrent state instead of a growing key-value (KV) cache. This makes ordinary decoding memory-efficient,…
33 -
arXiv — NLP / Computation & Language research 6d ago
MigrationNarrate: A Dataset for Detection of Migration Narratives in YouTube Videos
arXiv:2608.20984v1 Announce Type: cross Abstract: Narratives are central to how social communication is framed, making their detection critical for understanding and analysing public discourse. Prior work has explored narrative detection and extraction across diverse domains;…
23 -
arXiv — NLP / Computation & Language research 6d ago
COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models
arXiv:2608.21030v1 Announce Type: cross Abstract: Video multimodal large language models have advanced significantly, yet fine-grained motion-temporal understanding remains fragile. The core bottleneck is not only sparse frame sampling, but also the lack of a complete temporal…
37 -
arXiv — NLP / Computation & Language research 6d ago
Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems
arXiv:2608.21095v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) grounds Large Language Model (LLM) outputs in external knowledge, but RAG systems usually trust whatever they retrieve, creating a Security-Reliability Gap: high semantic relevance does not…
12 -
arXiv — NLP / Computation & Language research 6d ago
Personalized Privacy Control in LLMs via Attention Head Intervention
arXiv:2608.21209v1 Announce Type: cross Abstract: The rise of agentic AI enables LLMs to access diverse user data, raising critical privacy concerns. Prior work on contextual privacy studies whether LLMs regulate information disclosure according to context-dependent norms.…
20 -
arXiv — NLP / Computation & Language research 6d ago
Enhancing LLMs in Predictive Political QA with Semi-Structured Data
arXiv:2608.21218v1 Announce Type: cross Abstract: Predictive political question answering (QA), such as predicting how a political actor will vote, goes beyond factual lookup. External political resources offer rich historical evidence, but rarely contain the answer itself.…
28 -
arXiv — NLP / Computation & Language research 6d ago
TurboBias 2.0: Streaming Context-Biasing for Production-Efficient ASR Systems
arXiv:2608.21343v1 Announce Type: cross Abstract: Contextualization is essential for production automatic speech recognition (ASR) systems, where user-provided phrases must be recognized accurately under strict latency constraints. Although many context-biasing methods improve…
13 -
arXiv — NLP / Computation & Language research 6d ago
The Intrinsic Dimension of Prompts in Internal Representations of Large Language Models
arXiv:2501.10573v2 Announce Type: replace Abstract: We study the geometry of token representations at the prompt level in large language models through the lens of intrinsic dimension. Viewing transformers as mean-field particle systems, we estimate the intrinsic dimension of…
7 -
arXiv — NLP / Computation & Language research 6d ago
Explaining Intrinsic Moral Self-Correction with Mechanistic Interpretability
arXiv:2505.11924v4 Announce Type: replace Abstract: Intrinsic moral self-correction refers to the phenomenon where a language model refines its ethical judgments or aligns its outputs purely through prompting. While effective across diverse tasks, its mechanism remains unclear.…
36 -
arXiv — NLP / Computation & Language research 6d ago
Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
arXiv:2505.16222v2 Announce Type: replace Abstract: With the growing use of large language models(LLMs) as evaluators, their application has expanded to code evaluation tasks, where they assess the correctness of generated code without relying on reference implementations. While…
12 -
arXiv — NLP / Computation & Language research 6d ago
Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning
arXiv:2506.10903v2 Announce Type: replace Abstract: Statement autoformalization plays a crucial role in formal mathematical reasoning by enabling the automatic translation of natural language statements into formal languages. While recent advances using large language models…
18 -
arXiv — NLP / Computation & Language research 6d ago
GeoExplain: Multimodal Reasoning based on Hierarchy of Visual Information in Street View
arXiv:2506.16633v3 Announce Type: replace Abstract: Multimodal reasoning is a process of understanding, integrating and inferring information across different data modalities. It has recently attracted surging academic attention. Although there are various tasks for evaluating…
24 -
arXiv — NLP / Computation & Language research 6d ago
CPC-CMS: Cognitive Pairwise Comparison Classification Model Selection Framework for Document-level Sentiment Analysis
arXiv:2507.14022v3 Announce Type: replace Abstract: This study proposes the Cognitive Pairwise Comparison Classification Model Selection (CPC-CMS) framework for document-level sentiment analysis. The CPC, based on expert knowledge judgment, is used to calculate the weights of…
6 -
arXiv — NLP / Computation & Language research 6d ago
CulTrace: Tracing Internal Cultural Reasoning in Large Language Models
arXiv:2508.08879v4 Announce Type: replace Abstract: The growing deployment of large language models (LLMs) across diverse cultural contexts necessitates a deeper understanding of models' hidden representations of different cultures. Prior work has evaluated cultural awareness in…
33 -
arXiv — NLP / Computation & Language research 6d ago
SCOPE: A Generative Approach for LLM Prompt Compression
arXiv:2508.15813v2 Announce Type: replace Abstract: A big issue in modern LLM applications is they tend to feed long context to LLM, which results in high inference cost and latency, and may exceed the context limit. Prompt compression addresses this issue by reducing the length…
11 -
arXiv — NLP / Computation & Language research 6d ago
SKILL-RAG: Self-Knowledge Induced Learning and Filtering for Retrieval-Augmented Generation
arXiv:2509.20377v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) has significantly improved the performance of large language models (LLMs) on knowledge-intensive tasks in recent years. However, since retrieval systems may return irrelevant content,…
27 -
arXiv — NLP / Computation & Language research 6d ago
StruProKGR: A Structural and Probabilistic Framework for Sparse Knowledge Graph Reasoning
arXiv:2512.12613v2 Announce Type: replace Abstract: Sparse Knowledge Graphs (KGs) are commonly encountered in real-world applications, where knowledge is often incomplete or limited. Sparse KG reasoning, the task of inferring missing knowledge over sparse KGs, is inherently…
27 -
arXiv — NLP / Computation & Language research 6d ago
MedRAGChecker: Claim-Level Verification for Biomedical Retrieval-Augmented Generation
arXiv:2601.06519v2 Announce Type: replace Abstract: Biomedical retrieval-augmented generation (RAG) can ground LLM answers in medical literature, yet long-form outputs often contain isolated unsupported or contradictory claims with safety implications. We introduce…
15 -
arXiv — NLP / Computation & Language research 6d ago
SlidesGen-Bench: Evaluating Slides Generation via Computational and Quantitative Metrics
arXiv:2601.09487v2 Announce Type: replace Abstract: The rapid evolution of Large Language Models (LLMs) has fostered diverse paradigms for automated slide generation, ranging from code-driven layouts to image-centric synthesis. However, evaluating these heterogeneous systems…
15 -
arXiv — NLP / Computation & Language research 9d ago
A Virtual Member of a Community of Practice for the Society of Petroleum Engineers: From Prototype to Deployment
arXiv:2608.19199v1 Announce Type: new Abstract: We describe the evolution of a virtual assistant, called ATHENA, designed to support the capture, retrieval, and dissemination of knowledge for members of a Community of Practice (CoP) related to the Oil and Gas sector. An…
24