News / #paper Tag Research papers 500 articles archived under #paper · RSS Sign in to follow arXiv — NLP / Computation & Language research 3d ago Loss-Based Active Learning for Neural Abstractive Summarization arXiv:2608.25881v1 Announce Type: new Abstract: Fine-tuning abstractive summarization models requires high-quality annotated data. However, obtaining such corpora is expensive and time-consuming, as it requires human annotators to read and comprehend long documents to create… 8 arXiv — NLP / Computation & Language research 3d ago From Passive Response to Proactive Correction: Enhancing LLM Robustness Against Input Fact Perturbations arXiv:2608.25894v1 Announce Type: new Abstract: Large language models (LLMs) frequently produce confident yet factually incorrect responses when user inputs contain misleading premises, a phenomenon we attribute to fact perturbations in the input. Existing approaches to… 4 arXiv — NLP / Computation & Language research 3d ago One Form to Transfer Them All: Pretraining Multilingual Language Models Beyond Native Orthography arXiv:2608.25904v1 Announce Type: new Abstract: Multilingual language models transfer knowledge across languages through shared subword vocabulary, a mechanism that breaks down when related languages use different writing systems. Prior work addresses this via script… 24 arXiv — NLP / Computation & Language research 3d ago SAMpLE: A SystemC-AMS Machine LEarning-based Framework for Virtual Prototyping arXiv:2608.25910v1 Announce Type: new Abstract: Machine Learning (ML) is increasingly used in virtual prototypes of embedded systems to model behaviors that are difficult to capture analytically. However, integrating ML models into virtual platform simulation is still typically… 21 arXiv — NLP / Computation & Language research 3d ago Query-Side Attacks on GNN-Based KGQA: Tracing Failures from Entity Linking to Answer Generation arXiv:2608.25922v1 Announce Type: new Abstract: GNN-based Knowledge Graph Question Answering (KGQA) pipelines process queries through four discrete stages: entity linking, subgraph retrieval, GNN reasoning, and answer generation. Standard robustness evaluations conflate… 15 arXiv — NLP / Computation & Language research 3d ago Unveiling Spectral Mechanisms in Training-Free LLM Text Detection arXiv:2608.25944v1 Announce Type: new Abstract: The rapid advancement of Large Language Models (LLMs) makes it increasingly difficult to distinguish human writing from machine-generated text. Training-free detection offers a scalable solution, yet common confidence-based metrics… 27 arXiv — NLP / Computation & Language research 3d ago Lost but not erased: Finding traces of a forgotten language in neural speech models arXiv:2608.25976v1 Announce Type: new Abstract: International adoptees retain phonological traces of a birth language they can no longer speak or comprehend, a persistence typically attributed to a biologically-timed critical period. We asked whether it could instead reflect the… 4 arXiv — NLP / Computation & Language research 3d ago When Personality Meets Quantization: A Layer-wise MBTI Analysis of Quantized LLMs arXiv:2608.25977v1 Announce Type: new Abstract: Personality is increasingly important in large language models (LLMs), as it shapes users' trust, engagement, and emotional experiences. While the Myers--Briggs Type Indicator (MBTI) has emerged as a common framework for assessing… 22 arXiv — NLP / Computation & Language research 3d ago Distinct dynamics of conceptual and referential disruptions in human reading and large language model processing arXiv:2608.25999v1 Announce Type: new Abstract: Linguistic meaning is grounded in conceptual content, from which reference to particular entities emerges as words enter discourse. To examine the processing dynamics associated with these two dimensions of meaning, we selectively… 7 arXiv — NLP / Computation & Language research 3d ago VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following arXiv:2608.26013v1 Announce Type: new Abstract: Multimodal instruction-following models require training data that is accurate, diverse, verifiable, and challenging. Existing synthesis pipelines typically follow a one-pass generate-and-filter paradigm, discarding feedback from… 7 arXiv — NLP / Computation & Language research 3d ago Beyond Local Surprise: Grounded Dialogue as Selective Belief Revision under Referential Uncertainty arXiv:2608.26035v1 Announce Type: new Abstract: When a speaker refers to a scene that the listener cannot directly see, the listener must decide whether to preserve its current understanding or revise it as new utterances arrive. Many language systems treat local mismatch as a… 16 arXiv — NLP / Computation & Language research 3d ago Fine-Tuning Whisper for Automatic Speech Recognition in Baniwa: A Preliminary Study arXiv:2608.26060v1 Announce Type: new Abstract: Automatic Speech Recognition (ASR) technologies have achieved remarkable performance in recent years through the use of large multilingual foundation models. However, most advances remain concentrated on high-resource languages,… 9 arXiv — NLP / Computation & Language research 3d ago Prefix Sliding for efficient test-time scaling arXiv:2608.26070v1 Announce Type: new Abstract: Test-time scaling uses extra test-time compute to improve performance, such as letting language models reason longer when solving a problem. As models keep the entire reasoning trace in memory via full attention, hard tasks that… 26 arXiv — NLP / Computation & Language research 3d ago Natural Language Input, Semantic Track Representation, and LLM Inference: Making the Maritime Information Exchange Model Tractable arXiv:2608.24892v1 Announce Type: cross Abstract: We describe a practical architecture for making the Maritime Information Exchange Model (MIEM) and the broader Rich Semantic Track model tractable using current large language model (LLM) technology. The barrier to adoption of… 19 arXiv — NLP / Computation & Language research 3d ago PA-CoT: Profile-Adaptive Chain-of-Thought for Personalized Nutritional Consulting arXiv:2608.24907v1 Announce Type: cross Abstract: In health and nutrition consulting, widely used prompting methods pass the user profile as an unstructured block without a dedicated analysis step, leaving personalization as a critical structural gap. We introduce PA-CoT… 6 arXiv — NLP / Computation & Language research 3d ago Can We Read the Mind of an Audio LLM? A Verbalizable, Multilingual Middle-Layer Workspace arXiv:2608.24958v1 Announce Type: cross Abstract: An audio language model is a black box in a specific way: we see what it says, never what it works out on the way there, and chain-of-thought monitoring helps only if the model writes its reasoning down. Reading a base Qwen3-Omni… 8 arXiv — NLP / Computation & Language research 3d ago Retrieved But Not Reliable: A Survey on Attacks, and Defenses in Retrieval-Augmented Generation arXiv:2608.24977v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) enhances large language models by grounding outputs in external knowledge, improving factuality and reducing hallucinations. At the same time, the retrieval-augmented pipeline introduces new… 32 arXiv — NLP / Computation & Language research 3d ago FrontierChallenge: Evaluating Scientific Workflow Completion arXiv:2608.24979v1 Announce Type: cross Abstract: Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated programs, or a single domain. We introduce FrontierChallenge, a cross-domain… 11 arXiv — NLP / Computation & Language research 3d ago Retrieve, Match, Escalate: Accurate and Scalable Product Linking with VLM-Distilled Cross-Encoders and Agentic VLMs arXiv:2608.25037v1 Announce Type: cross Abstract: Product linking, the entity-resolution task of mapping merchant product records to canonical catalog products, consolidates fragmented listings so downstream search, recommendation, and advertising see one clean entry per… 19 arXiv — NLP / Computation & Language research 3d ago RefLAM: A Reference-Grounded Line Annotation Pipeline for Historical Arabic Manuscripts arXiv:2608.25140v1 Announce Type: cross Abstract: Existing approaches to building line-level Arabic handwritten-text-recognition (HTR) training data either rely on fully manual annotation, which does not scale, or on automatic OCR-to-reference alignment methods not yet extended… 33 arXiv — NLP / Computation & Language research 3d ago TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue arXiv:2608.25218v1 Announce Type: cross Abstract: Speakers in natural conversation take turns speaking and listening, deciding in real time when to take, hold, or yield the floor. However, turn-taking evaluation remains limited due to the lack of a consistent, linguistically… 14 arXiv — NLP / Computation & Language research 3d ago Rare Diseases, Common Dilemmas: LLMs Prioritize Equal Resource Distribution over Patient Benefit in Decision-Making arXiv:2608.25236v1 Announce Type: cross Abstract: Clinical decision-making often involves prioritizing ethical values, such as beneficence, non-maleficence, respecting a patient's autonomy, and justice. Recent work has begun to assess how large language models (LLMs) make such… 12 arXiv — NLP / Computation & Language research 3d ago The "Curse of Knowledge" in LLM Query Simulation: Concept Provenance for Tracing Answer-Side Intrusion arXiv:2608.25245v1 Announce Type: cross Abstract: LLM-generated search queries are widely used to augment IR evaluation, yet they may contain concepts that presuppose answer-side document knowledge, violating the information-access boundary of pre-search users. Existing… 25 arXiv — NLP / Computation & Language research 3d ago FinRiskAtlas: Decision-Aligned Evaluation of Large Language Models for Financial Risk Review arXiv:2608.25325v1 Announce Type: cross Abstract: Deploying large language models for professional financial review requires more than measuring general financial competence: models must perform the specific review operation required by a workflow and determine whether available… 38 arXiv — NLP / Computation & Language research 3d ago Learning What to Share and What to Personalize: Hierarchical Strategy Co-Evolution for Agent Memory arXiv:2608.25329v1 Announce Type: cross Abstract: Memory-augmented agents maintain compact user profiles throughout extended conversations, enabling personalized and consistent responses without the need to process the entire dialogue history. The quality of these user profiles… 5 arXiv — NLP / Computation & Language research 3d ago GGSS: Geodesic-Gated Spherical Steering for Inference-Time Debiasing of Generative Vision-Language Models arXiv:2608.25375v1 Announce Type: cross Abstract: Generative vision-language models (VLMs) are increasingly used in human-centered settings, yet they can produce demographically biased outputs even when images differ only in controlled attributes such as perceived race or… 7 arXiv — NLP / Computation & Language research 3d ago PonsRAG: A Pons-Inspired RAG Bridging Cognitive Islands for Coordinated Long Narrative Reasoning arXiv:2608.25486v1 Announce Type: cross Abstract: Long Narrative Reasoning is an essential capability for processing and reasoning over complex narratives. While retrieval-augmented generation provides a promising framework, existing methods still face two critical challenges:… 27 arXiv — NLP / Computation & Language research 3d ago CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval arXiv:2608.25500v1 Announce Type: cross Abstract: Reusable skill libraries allow large language model (LLM) agents to reuse procedural knowledge across tasks, but they also turn memory access into a challenging retrieval problem. Full-library prompting preserves coverage at high… 25 arXiv — NLP / Computation & Language research 3d ago Conditional Total Correlation and the Serial Depth of Adaptive Parallel Sampling arXiv:2608.25505v1 Announce Type: cross Abstract: Motivated by parallel decoding in masked diffusion models, we study adaptive parallel sampling of discrete vectors: in each round, a deterministic policy selects unrevealed coordinates on the basis of the values observed so far,… 14 arXiv — NLP / Computation & Language research 3d ago When Stale Constraints Go Unchecked: Budgeted Verification Failures in Inherited Agent Memory arXiv:2608.25553v1 Announce Type: cross Abstract: An agent that inherits a consolidated memory may inherit a constraint that was true when written and has since been withdrawn by a newer authoritative record. Under a scarce verification budget, does the agent recover the… 10 arXiv — NLP / Computation & Language research 3d ago Plans You Can Check: Verifier-Grounded Learning of an Open-Weight Planner for Executable Video-Editing arXiv:2608.25622v1 Announce Type: cross Abstract: Practical video editing is not only pixel generation: an editor must turn a brief, a clip pool, music metadata, and hard constraints into an executable timeline. We study this decision layer as \emph{executable video-editing… 32 arXiv — NLP / Computation & Language research 3d ago Formal, Executable and Explainable Runtime Monitoring of Spoken Air Traffic Control Operational Procedures arXiv:2608.25926v1 Announce Type: cross Abstract: Air traffic control procedures are executed through spoken exchanges between controllers and pilots. These interactions are essential to the safety of air transportation: failures in their execution can create severe operational… 21 arXiv — NLP / Computation & Language research 3d ago Code World Model: Coding Agent as World Brain arXiv:2608.25927v1 Announce Type: cross Abstract: World models aim to simulate how complex environments evolve under actions and events, yet existing video-based world models primarily learn dynamics from visual observations, which reveal outcomes rather than the underlying… 16 Hugging Face Daily Papers research 4d ago LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training Abstract LAION-BVD is a large-scale open video dataset enabling multimodal pre-training across video, audio, and image modalities with synthetic captions and strong benchmark performance. Generated by thinkingmachines/Inkling-Small We present LAION-BVD, a large-scale open video… 22 arXiv — Machine Learning research 4d ago Equivariant Cellular Sheaves for Molecular Electronic Structure: Bridging Sheaf Cohomology and E(3)-Equivariant Hamiltonian Learning arXiv:2608.23571v1 Announce Type: new Abstract: Equivariant message-passing networks are the standard model for molecular property and interatomic-potential prediction, and recent work predicts the electronic Hamiltonian itself in an E(3)-equivariant way. Separately, topological… 32 arXiv — Machine Learning research 4d ago Data Predictability Shapes Weibull Weight-Scale Growth in Transformer Training arXiv:2608.23573v1 Announce Type: new Abstract: A trained transformer's weight magnitudes can be summarized by a two-parameter Weibull distribution whose shape $k \approx 1.2$ is stable across layers and models, so the scale $\lambda$ carries most training-induced movement. What… 10 arXiv — Machine Learning research 4d ago From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers arXiv:2608.23660v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to provide prior causal knowledge for structural causal discovery, yet whether their direct-edge judgments and confidence can be trusted remains unclear. We systematically evaluate… 25 arXiv — Machine Learning research 4d ago Renormalization Group Flow Matching for Scalable Local Generative Modeling arXiv:2608.23696v1 Announce Type: new Abstract: Despite their remarkable success in modeling complex data, generative models face a fundamental tradeoff. Global approaches can capture full structural coherence but suffer from high computational costs, while local models are… 26 arXiv — Machine Learning research 4d ago Response Renormalization for Critical Deep Equilibrium Models arXiv:2608.23725v1 Announce Type: new Abstract: Deep Equilibrium Models (DEQs) compute predictions from a hidden representation unchanged by the model update. Training through this equilibrium uses implicit differentiation and requires solving an adjoint system built from the… 8 arXiv — Machine Learning research 4d ago Calibration-Preserving Pruning: Compression as a Reliability Contract arXiv:2608.23744v1 Announce Type: new Abstract: Split conformal prediction, not the pruning rule, supplies finite-sample marginal coverage once a pruned model is fixed independently of the conformal calibration split. We study the separate efficiency problem: can pruning… 31 arXiv — Machine Learning research 4d ago Tight Majorizations and Convergence Rates of Nuclear Norm Minimization IRLS arXiv:2608.23765v1 Announce Type: new Abstract: Iteratively reweighted least squares (IRLS) methods constitute a natural approach to nuclear norm minimization, but their convergence rates and the role of the weight operator have remained poorly understood. This paper establishes… 7 arXiv — Machine Learning research 4d ago Disentangled Skill Representations for Predictive Human Modeling arXiv:2608.23776v1 Announce Type: new Abstract: Understanding human skill is important for AI systems that collaborate with, coach, or assist people. Unlike typical latent variable estimation problems which rely on single observations, skill is a persistent, compositional, and… 21 arXiv — Machine Learning research 4d ago GAP-Prompt: Gated Adaptive Prompting for Efficient Continual Learning arXiv:2608.23782v1 Announce Type: new Abstract: Continual learning faces the persistent challenge of catastrophic forgetting, where sequential task updates degrade previously acquired knowledge. While prompt-based methods integrated with pre-trained models offer a compelling… 15 arXiv — Machine Learning research 4d ago Mixture of Channel Experts: Static Sparse Supports with Input-Adaptive Mixing for Pointwise Projections arXiv:2608.23794v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) scales language models by routing each input through a small set of independently parameterized experts. We show that copying this design into convolutional networks fails for a structural reason: parallel… 15 arXiv — Machine Learning research 4d ago A Theory of Speciation in Generative Diffusion Models on Compact Riemannian Manifolds arXiv:2608.23798v1 Announce Type: new Abstract: Speciation in generative diffusion models denotes the emergence of distinct stable branches during denoising, through which initially undifferentiated trajectories progressively commit to different data classes. In this work we… 6 arXiv — Machine Learning research 4d ago Discovering Cross-Language Reasoning Invariance in LLMs with Geometry-Invariant Sparse Autoencoders arXiv:2608.23809v1 Announce Type: new Abstract: Multilingual language models can solve the same mathematical problem in different languages, but it remains unclear whether they rely on shared features or on language-specific computations that only produce similar outputs. We… 31 arXiv — Machine Learning research 4d ago Learning to Grade Efficiently: A Bandit-Driven Prompt-Selection Framework for Low-Cost LLM Essay Scoring arXiv:2608.23814v1 Announce Type: new Abstract: Large Language Models (LLMs) demonstrate strong capabilities in automated essay scoring (AES), but contemporary approaches typically employ fixed prompt selection, failing to address operational cost concerns and evolving optimal… 26 arXiv — Machine Learning research 4d ago AQLoRA: A Zero-Search Recipe for Fast Quantized LoRA Fine-Tuning arXiv:2608.23816v1 Announce Type: new Abstract: Quantized fine-tuning (QLoRA) saves memory but not time. It dequantizes every 4-bit weight on the fly, so it trains more slowly than fp16 LoRA. We present AQLoRA (Adaptive-Quantization LoRA), a recipe that buys part of that time… 10 arXiv — Machine Learning research 4d ago Generating Intervention Hypotheses using Explainable Explanations on Graphs: G2I, a Two-Stage Greedy Framework arXiv:2608.23835v1 Announce Type: new Abstract: Real-world decision-making in public health and social science can greatly benefit from predictive models, yet translating predictions into effective interventions requires explaining the model behavior. While Graph Neural Networks… 36 arXiv — Machine Learning research 4d ago PuzzleKV: Page-Wise Low-Rank Decomposition for KV Cache Compression arXiv:2608.23843v1 Announce Type: new Abstract: Long-context inference in large language models (LLMs) is increasingly limited by the memory required for the key-value (KV) cache. KV cache compression addresses this problem by reducing the storage cost of previous tokens. Among… 9 Page 8 of 10 · 500 articles ← Newer Older →