arXiv — NLP / Computation & Language
500 articles archived · Visit source ↗ · RSS
-
arXiv — NLP / Computation & Language research 2d ago
TreeGraft: Adaptive Multi-Drafter Grafting for Tree-Based Speculative Decoding
arXiv:2608.26112v1 Announce Type: new Abstract: Speculative decoding accelerates large language model inference through a draft-then-verify paradigm. Building on this, tree-structured methods improve inference by organizing proposals into multiple candidate paths, increasing the…
27 -
arXiv — NLP / Computation & Language research 2d ago
ElementCheck: Complexity-Aware Long-Form Text Factuality Evaluation via Sentence Elements
arXiv:2608.26118v1 Announce Type: new Abstract: Existing long-form factuality evaluation relies on the decompose-retrieve-verify pipeline. However, the pipeline suffers from noise from claim decomposition and fixed verification granularity, resulting in unreliable results. We…
26 -
arXiv — NLP / Computation & Language research 2d ago
DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs
arXiv:2608.26119v1 Announce Type: new Abstract: Whether large language models can be prompted to generate rhetorical fallacies on demand, and whether current safety post-training constrains this behavior, has received less attention than the related question of detecting…
30 -
arXiv — NLP / Computation & Language research 2d ago
Recipes for Steering and Scaling LLMs via Sampling
arXiv:2608.26120v1 Announce Type: new Abstract: Large Language Models (LLMs) are probabilistic models, typically defined by an autoregressive factorization. While recent work has begun to study richer target distributions beyond the base model, the sampling strategies remain…
4 -
arXiv — NLP / Computation & Language research 2d ago
Can a Model Catch Its Own Hallucinations for Free?: Label-Free Doubt Signals Hold Their Own Against a Labelled Dataset for Abstention
arXiv:2608.26121v1 Announce Type: new Abstract: Large language models state false facts as fluently as true ones, yet a model often "knows" internally when it is on shaky ground: the probability it assigns to its own answer tends to dip on the facts it gets wrong. The usual way…
19 -
arXiv — NLP / Computation & Language research 2d ago
Which India Survives Translation? Narrative Homogenisation Across Indian Oral Traditions in LLMs
arXiv:2608.26123v1 Announce Type: new Abstract: Large language models (LLMs) are trained predominantly on English-language internet text that over-represents certain cultural narratives, raising concerns that models flatten the diversity of non-Western storytelling traditions…
10 -
arXiv — NLP / Computation & Language research 2d ago
Natural-Language Policies to Executable Decisions: An Interpretable Large Language Model Framework
arXiv:2608.26124v1 Announce Type: new Abstract: Pricing automation in large-scale tourism is challenging because travel orders are highly unstructured, while pricing policies are complex, rapidly evolving, and inherently open-ended. Traditional rule engines are brittle and…
33 -
arXiv — NLP / Computation & Language research 2d ago
Training-Time Explainability for Multilingual Hate Speech Detection: Aligning Model Reasoning with Human Rationales
arXiv:2608.26125v1 Announce Type: new Abstract: Online hate against Muslim communities often appears in culturally coded, multilingual forms that evade conventional AI moderation. Such systems, though accurate, remain opaque and risk bias, over-censorship, or under-moderation,…
17 -
arXiv — NLP / Computation & Language research 2d ago
TelecomGPT-R1: A Unified Open-Source Reasoner for the Telecom Stack
arXiv:2608.26126v1 Announce Type: new Abstract: Telecommunications is a high-leverage domain for large language model (LLM)-based reasoning because routine engineering workflows require joint grounding in normative specifications, operational telemetry, vendor-specific fault…
5 -
arXiv — NLP / Computation & Language research 2d ago
FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial Outcomes
arXiv:2608.26129v1 Announce Type: new Abstract: Scientific peer review datasets have trained AI systems exclusively on Computer Science and Machine Learning venues, producing models that critique ablation studies yet have never seen a biology reviewer demand contamination…
8 -
arXiv — NLP / Computation & Language research 2d ago
Agents Don't Paginate: First-Chunk Selection for LLM Tool Responses
arXiv:2608.26130v1 Announce Type: new Abstract: Coding agents built on large language models (LLMs), such as Claude Code, Cursor, OpenAI Codex, GitHub Copilot, and Aider, receive tool responses that routinely exceed the agent's per-turn token budget. The standard remedy,…
32 -
arXiv — NLP / Computation & Language research 2d ago
Evaluating Language Models in Realistic Conversational Contexts
arXiv:2608.26131v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed to serve open-ended, multi-turn interactions, evaluating conversational quality at human scale has become a central challenge. Existing evaluation frameworks built for…
20 -
arXiv — NLP / Computation & Language research 2d ago
Agent Seer: Synthesizing Scenarios from Specification Understanding
arXiv:2608.26133v1 Announce Type: new Abstract: Evaluating AI agents that use external tools requires realistic test scenarios that capture how practitioners compose tools and iterate across conversation turns. Constructing such scenarios by hand demands deep domain expertise,…
18 -
arXiv — NLP / Computation & Language research 2d ago
Data Science Approaches to Evaluating Honours Candidates
arXiv:2608.26135v1 Announce Type: new Abstract: We present a modular data-science pipeline for estimating public sentiment towards individuals from fragmented, unstructured open-source intelligence (OSINT). The method chains web search, text extraction, relevance filtering,…
11 -
arXiv — NLP / Computation & Language research 2d ago
Reward-Informed Sparse Autoencoders and the Solution-Completeness Confound
arXiv:2608.26136v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) decompose language-model activations into sparse, interpretable features, and an appealing way to aim them at reasoning is to curate their data with a signal reinforcement learning already produces: the…
26 -
arXiv — NLP / Computation & Language research 2d ago
Interpretable, Fairly Evaluated Automated L2 Speaking Assessment that Beats the Single-Human Ceiling and Why Pause Encoding Does Not Change LLM Fluency Scores
arXiv:2608.26137v1 Announce Type: new Abstract: Second-language (L2) English learners can rarely rehearse speaking with a partner. Speaking is also the most anxiety-laden skill. These gaps drive a fast-growing market for automated speaking practice and scoring. But an automated…
24 -
arXiv — NLP / Computation & Language research 2d ago
Cross-Platform Generalisation Failure in Mental Health Natural Language Processing: A Five-Axis Fairness Audit of Transformer Models on Social Media
arXiv:2608.26138v1 Announce Type: new Abstract: We introduce the Cross-Platform Fairness Evaluation (CPFE) framework -- a five-axis audit protocol covering discriminative performance, calibration, statistical significance, prediction equity, and attribution stability -- and…
34 -
arXiv — NLP / Computation & Language research 2d ago
Syntax vs. Semantics: How Transformers Learn Deep Dependencies
arXiv:2608.26139v1 Announce Type: new Abstract: Large Language Models demonstrate remarkable syntactic fluency, yet the optimization dynamics governing their acquisition of deep semantic dependencies remain poorly understood. We propose a mechanistic framework that models this…
4 -
arXiv — NLP / Computation & Language research 2d ago
Affix Cache for Diffusion Large Language Models
arXiv:2608.26140v1 Announce Type: new Abstract: Diffusion Large Language Models (DLLMs) enable non-autoregressive decoding and bidirectional context modeling, but efficient inference remains challenging. Unlike autoregressive systems, whose key-value (KV) cache can be reused for…
38 -
arXiv — NLP / Computation & Language research 2d ago
AdaThinking-E: One-Token Entropy Regulation for Adaptive Thinking
arXiv:2608.26141v1 Announce Type: new Abstract: Multimodal large language models have demonstrated strong document reasoning capabilities by incorporating explicit thinking processes. While this capability significantly improves performance on challenging tasks, current models…
13 -
arXiv — NLP / Computation & Language research 2d ago
Position Is All You Need: A Free Lunch Token Compression Strategy for MLLM-based Referring Expression Segmentation
arXiv:2608.26142v1 Announce Type: new Abstract: Referring Expression Segmentation (RES) aims to generate pixel-wise segmentation masks from complex and implicit textual queries. While recent advances in Multimodal Large Language Models (MLLMs) have substantially boosted RES…
15 -
arXiv — NLP / Computation & Language research 2d ago
Beyond Accuracy: A Qualitative Analysis of Vision-Language Models for Hate Speech Detection in Memes
arXiv:2608.26143v1 Announce Type: new Abstract: Memes have turned out to be a powerful tool through which individuals share their ideas concerning contemporary social and political problems. Their anonymity, as well as their ability to go viral, make them a powerful medium for…
28 -
arXiv — NLP / Computation & Language research 2d ago
Why Current XAI Is Not Enough for Arabic NLP: A Critical Survey of the Explainability Gap
arXiv:2608.26144v1 Announce Type: new Abstract: Explainable AI (XAI) is now a major theme in NLP; however, Arabic NLP remains under-explained in three connected senses. First, there is a method gap: Arabic XAI relies heavily on a small set of post-hoc techniques such as LIME,…
21 -
arXiv — NLP / Computation & Language research 2d ago
Vagdhenu: A Vrutta (Meter) Aware Shloka-to-Chant (TTS) System for Sanskrit
arXiv:2608.26146v1 Announce Type: new Abstract: We present Vagdhenu, a vrutta (meter) aware shloka-to-chant system for Sanskrit: a text-to-speech system that maps a metrical verse to its chanted parayana recitation at high fidelity. This is an experience report, not a new…
33 -
arXiv — NLP / Computation & Language research 2d ago
CARE: Causally-Aligned Reasoning Exploration for Medical Large Language Models
arXiv:2608.26147v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown strong potential for medical reasoning, yet the scarcity and cost of expert-annotated data constrain their progress. While reinforcement learning offers a scalable alternative, standard…
21 -
arXiv — NLP / Computation & Language research 2d ago
Towards Interpretable Depression Detection: Linking Acoustic Features to DSM-5 Indicators
arXiv:2608.26148v1 Announce Type: new Abstract: Depression affects millions worldwide, yet diagnosis relies on subjective self-reports that may miss authentic behavior. This paper presents an approach linking speech acoustics to DSM-5 depressive-behavior indicators through a…
26 -
arXiv — NLP / Computation & Language research 2d ago
Artificial Intelligence Models Can Predict and Collaboratively Modulate Human Memory Search
arXiv:2608.26152v1 Announce Type: new Abstract: Large language models (LLMs) exhibit unprecedented natural language generation and many text-based problem-solving capabilities. Indeed, in many language-based tasks, for example routine coding, these artificial intelligence models…
28 -
arXiv — NLP / Computation & Language research 2d ago
Evaluating AI Generated Summaries for Cancer Patients
arXiv:2608.26154v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly being integrated into digital health platforms to generate summaries of complex medical data. Although these models can improve patient engagement and communication, these systems also…
7 -
arXiv — NLP / Computation & Language research 2d ago
VFA: Empowering Multilingual MLLMs via Vision-Free Adaptation
arXiv:2608.26155v1 Announce Type: new Abstract: Multimodal large language models have advanced rapidly, yet most remain English-centric, as scaling multilingual multimodal instruction tuning is limited by the scarcity and high cost of high-quality non-English image-text…
6 -
arXiv — NLP / Computation & Language research 2d ago
Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation
arXiv:2608.26159v1 Announce Type: new Abstract: Self-Generated Text Recognition (SGTR)--the ability of an LLM to identify its own outputs--poses risks to AI safeguards that rely on LLMs as evaluators or monitors. Specifically, an LLM may recognize outputs from other copies of…
34 -
arXiv — NLP / Computation & Language research 2d ago
Mutual Debiasing via Dual-Seed Comparison for Probabilistic Sampling in Large Language Models
arXiv:2608.26161v1 Announce Type: new Abstract: Although Large Language Models (LLMs) demonstrate remarkable capabilities in reasoning and decision-making, high-fidelity probabilistic sampling remains a persistent challenge. When generating random variables, LLMs consistently…
14 -
arXiv — NLP / Computation & Language research 2d ago
From Sound to Symptom: Real-Time Respiratory Signal Understanding for Conversational Healthcare Agents
arXiv:2608.26163v1 Announce Type: new Abstract: Cough events during live spoken conversations carry clinically valuable respiratory signals, yet existing dialogue systems treat them as acoustic noise to be discarded. We present HealthCUES (Clinical Understanding from Embodied…
31 -
arXiv — NLP / Computation & Language research 2d ago
Using Poly-Encoders for Computationally Efficient Automated Creativity Assessment
arXiv:2608.26165v1 Announce Type: new Abstract: Automated creativity assessment has been a long standing challenge, with traditional methods often being resource intensive or lacking practical accuracy. We introduce a novel approach by using Poly-Encoder for computationally…
8 -
arXiv — NLP / Computation & Language research 2d ago
Hallucinations in LLMs: A Lifecycle-Based Survey of Causes, Detection, Mitigation, and Prevention
arXiv:2608.26168v1 Announce Type: new Abstract: The lifecycle of hallucination in LLMs is a concept that enables building solid frameworks on the control and reliability of LLMs in high-stakes environments, including health, legal, and scientific research. Although previous…
16 -
arXiv — NLP / Computation & Language research 2d ago
Lost in Compression: A Controlled Cross-Lingual Audit of Extractive Prompt Compressors
arXiv:2608.26175v1 Announce Type: new Abstract: Extractive prompt compression promises to cut LLM inference costs by removing low-information tokens, and learned compressors such as LLMLingua-2 report strong results on English benchmarks. Most other languages already pay a token…
4 -
arXiv — NLP / Computation & Language research 2d ago
A Multi-Framework Comparison of Outline Stages in Long-Form Generation with LLMs
arXiv:2608.26177v1 Announce Type: new Abstract: Long-form generation exposes fundamental limitations of large language models. Even 70B-parameter models exhibit length collapse at 16k-token outputs, and multi-chapter stories frequently trigger the attribute drift characteristic…
11 -
arXiv — NLP / Computation & Language research 2d ago
PACEShop: Evaluating Personalized, Actionable, Compositional, and Evidence-grounded Shopping Assistants
arXiv:2608.26180v1 Announce Type: new Abstract: Shopping assistants are shifting from ranked product lists toward structured decision support, where systems must synthesize shopper context, product evidence, and next-step guidance into a coherent recommendation experience. This…
26 -
arXiv — NLP / Computation & Language research 2d ago
Investigating the Influence of Prompt and Response Languages on LLM Content Generation
arXiv:2608.26186v1 Announce Type: new Abstract: This study examines how prompt and response language influence the behavior of large language models. Using five models, we evaluated answers to 68 non translation questions across four language conditions: English to English,…
18 -
arXiv — NLP / Computation & Language research 2d ago
When the Canonical Completion Is Wrong: Formalizing and Measuring the Jump in Large Language Models
arXiv:2608.26187v1 Announce Type: new Abstract: Whether large language models (LLMs) can perform the abductive leap from evidence to a new system of axioms, commonly referred to as a jump, has recently attracted considerable debate. A prominent position holds that LLMs are…
17 -
arXiv — NLP / Computation & Language research 2d ago
Comparing Chunking and Embedding Strategies for Turkish RAG Systems
arXiv:2608.26192v1 Announce Type: new Abstract: How documents are segmented into retrievable chunks and how those chunks are embedded strongly affect Retrieval-Augmented Generation (RAG) quality, yet neither has been systematically studied for morphologically rich languages such…
8 -
arXiv — NLP / Computation & Language research 2d ago
A Reranker for Orchestrating Heterogeneous Speech and Text Retrievers
arXiv:2608.26194v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) systems have attracted significant interest for their ability to mitigate hallucinations in Large Language Models (LLMs). Although knowledge databases for RAG are increasingly diversifying to…
16 -
arXiv — NLP / Computation & Language research 2d ago
On Scope Classification and Current Knowledge-Editing Benchmarks: A Negative Result, with INLAY as a Gradient-Free Case Study
arXiv:2608.26292v1 Announce Type: new Abstract: Every memory-based knowledge editor in the SERAC lineage depends on a scope decision: given a query, does a stored edit apply? We report that current knowledge-editing benchmarks cannot measure this decision at all. Using INLAY, a…
27 -
arXiv — NLP / Computation & Language research 2d ago
MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models
arXiv:2608.26295v1 Announce Type: new Abstract: Tool-augmented LLMs must arbitrate between two fallible sources when a tool return conflicts with their parametric memory, yet existing evaluations measure source preference without establishing source correctness. We introduce…
25 -
arXiv — NLP / Computation & Language research 2d ago
When Is Noise Response Universal? Tokenization as the Hidden Variable in Language Models
arXiv:2608.26319v1 Announce Type: new Abstract: The performance of textual neural models often degrades when their inputs are corrupted by noise such as typos, OCR errors, or dropped words. We study the degradation rate across neural models, both sentence embeddings and…
18 -
arXiv — NLP / Computation & Language research 2d ago
How Unlikely Is "Unlikely"? Assessing Verbal Probability Perception Across Large Language Models
arXiv:2608.26327v1 Announce Type: new Abstract: Large language models increasingly produce and interpret verbal probability expressions, yet whether these expressions carry consistent meaning across models (or match human perceptions of uncertainty) remains unknown. We present a…
33 -
arXiv — NLP / Computation & Language research 2d ago
Neuro-symbolic PRM: Enhancing Scientific Reasoning via Structured Traces and Symbolic Verification
arXiv:2608.26329v1 Announce Type: new Abstract: While tool-augmented Large Language Models have significantly improved multi-step reasoning in quantitative STEM tasks, a critical residual failure mode remains: intermediate reasoning steps that are syntactically well-formed,…
17 -
arXiv — NLP / Computation & Language research 2d ago
MoganColBERT-TR: A Late-Interaction Multi-Vector Retrieval Model for Turkish
arXiv:2608.26344v1 Announce Type: new Abstract: We previously reported a ModernBERT encoder trained from scratch for Turkish (MoganBERT-TR) and a single-vector embedding model built on top of it (MoganBERT-embed). This work introduces the third model in that lineage:…
37 -
arXiv — NLP / Computation & Language research 2d ago
Cross-lingual Representation Learning via Centroid Intervention Fusion
arXiv:2608.26357v1 Announce Type: new Abstract: Large language models (LLMs) exhibit uneven multilingual performance, especially when dealing with low-resource languages. Inference-time intervention offers a lightweight way to improve cross-lingual transfer by modifying the…
32 -
arXiv — NLP / Computation & Language research 2d ago
Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives
arXiv:2608.26372v1 Announce Type: new Abstract: Large language models are increasingly deployed as autonomous agents serving users on behalf of companies, placing them in settings where user and deployer interests can conflict. When an agent knows that a user is owed something…
23 -
arXiv — NLP / Computation & Language research 2d ago
Survival-Guided Length Control for Efficient Diffusion Language Models
arXiv:2608.26374v1 Announce Type: new Abstract: Diffusion language models (DLMs) generate text by iteratively denoising masked sequences, but standard decoding either fixes the sequence length or relies on ad hoc stopping rules, often leading to unnecessary denoising steps. We…
5