News / #rag Tag Rag 500 articles archived under #rag · RSS Sign in to follow arXiv — NLP / Computation & Language research 6d ago EnSI-RAG: Entity-Structure-Indexed Retrieval-Augmented Generation for Long-Document Question Answering arXiv:2608.21252v1 Announce Type: new Abstract: Question answering (QA) over long, connected documents remains challenging because relevant evidence may span multiple entities and their relationships. Existing retrieval-augmented generation (RAG) methods typically index… 23 arXiv — NLP / Computation & Language research 6d ago Profiling What Matters: Context-Aware Item Profiles from Large-Scale Metadata for LLM Recommenders arXiv:2608.20801v1 Announce Type: cross Abstract: While Large Language Models (LLMs) have significantly advanced reranking in recommendation, effectively leveraging item-side information remains challenging. Real-world items are described by vast, heterogeneous, and unstructured… 18 arXiv — NLP / Computation & Language research 6d ago Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems arXiv:2608.21095v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) grounds Large Language Model (LLM) outputs in external knowledge, but RAG systems usually trust whatever they retrieve, creating a Security-Reliability Gap: high semantic relevance does not… 12 arXiv — NLP / Computation & Language research 6d ago SKILL-RAG: Self-Knowledge Induced Learning and Filtering for Retrieval-Augmented Generation arXiv:2509.20377v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) has significantly improved the performance of large language models (LLMs) on knowledge-intensive tasks in recent years. However, since retrieval systems may return irrelevant content,… 27 arXiv — NLP / Computation & Language research 6d ago MedRAGChecker: Claim-Level Verification for Biomedical Retrieval-Augmented Generation arXiv:2601.06519v2 Announce Type: replace Abstract: Biomedical retrieval-augmented generation (RAG) can ground LLM answers in medical literature, yet long-form outputs often contain isolated unsupported or contradictory claims with safety implications. We introduce… 15 Vercel — AI dev-tools 6d ago Vercel Sandbox is now globally available Vercel Sandbox now runs globally, starting with four regions: iad1 (Washington, D.C.), sfo1 (San Francisco), cle1 (Cleveland), and cdg1 (Paris). iad1 remains the default. Support for all Vercel regions is coming soon. Choose a region close to the databases, object storage, and… 8 r/LocalLLaMA community 6d ago Qwen3.8-27B NVFP4 with vision + 451K token KV-cache on one RTX 5090 (power limited to 400W) at 120 tokens/s average Hello, So I've been trying lots of combinations in that never-ending landscape of options and settings. I wanted a proper quant of 3.8 27B running as fast as possible on my 5090 at 400W, with vision and with as much KV-cache as possible and with concurrency enabled (aiming at 3… 25 r/LocalLLaMA community 7d ago I hosted Kimi K3 (2.8T parameters) using 8 B300s. 92 tok/s, $190 per million tokens What I ran: 8x B300 on Modal, $56.79 per hour, vLLM, tensor parallel 8, native MXFP4 Cold boot ~27 min (1.56 TB load, JIT, 51 CUDA graph captures) TTFT 0.92 to 1.02 s, decode 92 tok/s steady, 83 tok/s average over 4 prompts $190 per million output tokens. One clean run is about… 10 r/MachineLearning community 7d ago Ablating 1 of a chess transformer's 128 attention heads causes the model to stop finding the queen sacrifice in a famous chess game. [P] Hooks and reads out Maia-3 23m model with chessformer_lens library: github.com/chessformer-lens/chessformer_lens DOI: 10.5281/zenodo.21986988   submitted by   /u/Weird-Asparagus4136 [link]   [comments] 18 llama.cpp releases dev-tools 8d ago b10577 common : fix draft-mtp with embeddings ( #26352 , #27299 ) ( #27400 ) common: fix draft-mtp with embeddings ( #26352 ) --whitespace Co-authored-by: Sigbjørn Skjæret sigbjorn.skjaeret@huggingface.co Website: https://llama.app Attestations:… 25 r/LocalLLaMA community 8d ago Sharp template to NInfer: -42% output tokens, same speed Sharp is u/peculiar-ragdoll 's system prompt that makes Qwen answer way more tersely without losing correctness. It's built on top of froggeric's fixed chat templates for Qwen; several fixes now in the v22.x templates (error-escalation tiers, false retry-loop kills, multi-system… 20 Vercel — AI dev-tools 8d ago Deployment Storage keeps your deployments rollback-ready Every deployment produces a set of files, including the pages, functions, and assets Vercel serves. Deployment Storage keeps those files available so you can inspect previous deployments and roll back when needed. Instantly roll back to previous deployments in seconds If a… 19 The Information — AI news-outlet 8d ago Dragoneer Founder Stad to Buy Timberwolves Controlling Stake Marc Stad, the founder of Dragoneer Investment Group, is buying a controlling stake in the Minnesota Timberwolves and Minnesota Lynx professional basketball teams, according to The New York Times’ Athletic publication. Stad is buying the interest at a $4.5 billion valuation from… 15 Hugging Face Daily Papers research 8d ago The Embedder's Dilemma: LLMs Are Better, but at What Cost? Abstract Large language models and dedicated embedding models achieve nearly identical aggregate performance across diverse tasks, but embedding models are far cheaper and faster, supporting a division of labor by task type. Generated by thinkingmachines/Inkling-Small Should you… 33 Hugging Face Daily Papers research 9d ago Listening Forward: Next Patch Embedding Prediction Enables Scalable Audio Learners Abstract NAPE uses causal Transformers to predict successive spectrogram patch embeddings for self-supervised audio learning without auxiliary components. Generated by thinkingmachines/Inkling-Small Self-supervised learning (SSL) has driven substantial progress in audio… 6 arXiv — NLP / Computation & Language research 9d ago DeltaMomentum: A Key-Value based Anisotropic Momentum Update via Delta Rule arXiv:2608.19491v1 Announce Type: cross Abstract: Most modern optimizers form their momentum as an exponential moving average (EMA) of past gradients, forgetting every direction at one fixed rate. However, the inputs a deep network sees during training can be highly anisotropic,… 37 arXiv — Machine Learning research 9d ago SAGE-XGBoost: Spatially Augmented Graph Embeddings--Machine Learning Framework for Natural Hazards Susceptibility Mapping under Data Scarcity arXiv:2608.19672v1 Announce Type: new Abstract: Natural hazard susceptibility mapping is often constrained by limited labeled data, reducing the generalizability of conventional machine learning and limiting the applicability of complex deep learning models. This study proposes… 22 arXiv — Machine Learning research 9d ago Orthogonal JEPA: Factorized Predictive States for Latent World Models arXiv:2608.20065v1 Announce Type: new Abstract: World models construct latent states that support prediction, planning, and reasoning about an underlying system. Joint-embedding predictive architectures (JEPAs) offer a direct way to learn such states by predicting targets in… 15 arXiv — Machine Learning research 9d ago Learning Hierarchical Skill Policies with Offline Quality-Diversity Reinforcement Learning arXiv:2608.19684v1 Announce Type: cross Abstract: Recent studies investigate how to leverage pre-collected datasets to improve the policy performance and sample efficiency of RL. One promising approach to achieve this goal is to employ a two-stage strategy: In the first stage,… 36 arXiv — NLP / Computation & Language research 9d ago Automated Summarization of Financial News Using Large Language Models and Retrieval-Augmented Generation: An Early Empirical Study (Fall 2023) arXiv:2608.19526v1 Announce Type: new Abstract: Stock market analysts and investors face a daily challenge: too much financial news, too little time. Manually reading and synthesizing hundreds of company-specific articles is impractical, yet missing key information can directly… 38 arXiv — NLP / Computation & Language research 9d ago SABET-QA: Temporal Knowledge Graph Question Answering arXiv:2608.20083v1 Announce Type: new Abstract: Question Answering over Temporal Knowledge Graphs (TKGQA) requires reasoning over time-sensitive facts, yet existing embedding-based methods struggle with multi-step queries due to single-pass reasoning pipelines. We propose… 21 arXiv — NLP / Computation & Language research 9d ago From Retrieved Context to Runtime Control: Adaptive Compression for Edge-based RAG arXiv:2608.19535v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) improves language-model responses by grounding generation in external passages, which comes with overhead: retrieved context lengthens the prompt, increasing prefill work, KV-cache footprint,… 32 Hugging Face Daily Papers research 9d ago Repo0: Design-Driven Zero-to-All Code Generation Abstract Repo0 uses a dual-graph architectural state and modularity-guided structural evolution to generate complete software repositories from natural-language requirements with high functionality coverage. Generated by thinkingmachines/Inkling-Small Large language model agents… 15 Hugging Face Daily Papers research 9d ago VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation Abstract A human-aligned chain-of-thought reward model and preference dataset improve joint video-audio generation by replacing fragmented metrics with coherent, dimension-wise reinforcement learning. Generated by thinkingmachines/Inkling-Small Using reinforcement learning to… 29 r/MachineLearning community 9d ago Is KV Cache in a high dimensional vector space? [D] I've been doing some research on this question: At inference time a large part of a model's working memory lives in the KV cache, plus whatever external memory the harness bolts on. I've been poking at the storage-and-retrieval side of this, treating that cache as an index, and… 34 llama.cpp releases dev-tools 9d ago b10517 vulkan : dequant q8_0 KV once in coopmat1 ( #25494 ) vulkan : dequant q8_0 KV once in coopmat1 Assisted-by: Claude (Opus 4.8) vulkan : fall back instead of aborting when FA scratch exceeds maxStorageBufferRange vulkan : require KV-cache layout in FA dequant path Assisted-by:… 19 arXiv — Machine Learning research 10d ago Physics-Unrolled Neural Operator for Wireless Field Modeling arXiv:2608.18495v1 Announce Type: new Abstract: Radio maps are essential for wireless decision-making tasks such as access-point placement, coverage planning, and localization, but their fine spatial details are governed by complex propagation effects and are costly to simulate… 36 arXiv — Machine Learning research 10d ago Denoising-Aware Inversion: Revealing Privacy Risks in Noise-Protected Text Embeddings arXiv:2608.18610v1 Announce Type: new Abstract: Dense text embeddings are widely used in data mining, retrieval, and downstream machine learning systems due to their compact and semantically rich representations, but recent embedding inversion attacks have shown that they can… 24 arXiv — Machine Learning research 10d ago FedLNS: Leverage LayerNorm Signature Modeling to Mitigate Adversarial Manipulation in Federated LLMs arXiv:2608.18736v1 Announce Type: new Abstract: Federated training enables language models to learn from distributed private text, but the server cannot directly verify the local supervision or optimization process that produces each client update. A malicious client can… 25 arXiv — Machine Learning research 10d ago Enhancing Distance-Based Graph Autoencoders with Structural Penalties for Dynamic Graph Embedding arXiv:2608.18762v1 Announce Type: new Abstract: Graph autoencoders (GAEs) are widely used for learning representations of dynamic graphs. However, their optimisation objectives typically do not take structural heterogeneity across nodes into account. We propose three… 25 arXiv — Machine Learning research 10d ago On the Slow Convergence to Trivial Solutions of Algorithms for Hard Optimization Problems arXiv:2608.18910v1 Announce Type: new Abstract: Hard combinatorial optimization problems, many of which are NP-hard, present fundamental algorithmic challenges. Average-case analysis on random instances has emerged as a powerful framework for understanding typical algorithmic… 36 arXiv — Machine Learning research 10d ago Lost in Aggregation: How Benchmarks Overlook Irreplaceable Model Strengths arXiv:2608.18919v1 Announce Type: new Abstract: Tabular machine learning benchmarks typically summarize performance by averaging scores, ranks, or pairwise wins across datasets. Such aggregates are useful for selecting robust default models, but they can obscure a different… 6 arXiv — Machine Learning research 10d ago Multi-Agent Off-Policy Deep Reinforcement Learning for Smart Campus Coverage arXiv:2608.19049v1 Announce Type: new Abstract: Deep reinforcement learning (DRL) has recently gained a great attention due to its real-time adaptation and effectiveness in complex optimization problems. This paper investigates the optimal deployment of millimeter-wave (mmWave)… 6 arXiv — Machine Learning research 10d ago Beyond Trial Averaging: Anchoring Neural and Visual Representations for Few-Repetition Brain-to-Image Retrieval arXiv:2608.19128v1 Announce Type: new Abstract: Decoding visual information from brain signals probes neural representations and enables neuro-rehabilitation and dream decoding. Recent brain-to-image retrieval approaches have achieved promising performance, typically by… 22 arXiv — Machine Learning research 10d ago Optimizing Energy Efficiency and Grid Stability via Public EV Charging Flexibility arXiv:2608.18126v1 Announce Type: cross Abstract: This study evaluates the potential of electric vehicle (EV) charging flexibility to enhance both energy efficiency and power grid stability. Using real-world data from public charging stations in Prague, we analyze individual and… 20 arXiv — NLP / Computation & Language research 10d ago MissDiag: Diagnostic Evaluation of Incomplete-Knowledge Robustness in KGQA and KG-RAG arXiv:2608.18489v1 Announce Type: new Abstract: Knowledge graph question answering (KGQA) and knowledge-graph-based retrieval-augmented generation (KG-RAG) aim to ground answers in explicit graph evidence, but real-world knowledge graphs are often sparse, outdated, and… 28 arXiv — NLP / Computation & Language research 10d ago From Storage to Access: Verifiable Activation of Parametric Knowledge in LLMs via Explicit Priming and Implicit Reasoning arXiv:2608.18581v1 Announce Type: new Abstract: Although Large Language Models (LLMs) encode rich factual knowledge in their parameters, reliably recalling and verifying such knowledge remains a key bottleneck in factual question answering. Existing end-to-end methods entangle… 37 arXiv — NLP / Computation & Language research 10d ago Learning What to Fail On: Failure-Mode Contextual Bandits for Adversarial Data Curation arXiv:2608.18681v1 Announce Type: new Abstract: We introduce a failure-aware adversarial retrieval-augmented framework for improving robustness in natural language understanding. Rather than selecting synthetic examples with a fixed reward threshold, our method formulates… 26 arXiv — NLP / Computation & Language research 10d ago MemFuse: Multi-Source Memory Fusion from Fragmented Observations arXiv:2608.18704v1 Announce Type: new Abstract: Long-term memory is essential for agents that operate across extended interactions, yet existing memory systems and benchmarks predominantly focus on single-source textual histories. In realistic settings, however, relevant… 10 arXiv — NLP / Computation & Language research 10d ago Gradient Mirage: Trainable yet Label-Unidentifiable Gradients in Large Language Model Split Learning arXiv:2608.18767v1 Announce Type: new Abstract: Gradient matching attacks (GMAs) in LLM split learning (SL) rely on a critical yet underexplored assumption: the gradient exposed at the split interface is a faithful derivative of the client's full-label training objective. This… 19 arXiv — NLP / Computation & Language research 10d ago DeepWeaver: Bridging the Evidence Synthesis Gap in Open-Ended Question Answering arXiv:2608.18988v1 Announce Type: new Abstract: Retrieve-then-generate pipelines are commonly used to produce deep-research answers for open-ended questions, but retrieval alone is insufficient: LLMs must organize noisy and fragmented evidence into comprehensive, well-cited… 27 arXiv — NLP / Computation & Language research 10d ago Comment-level Topic Drift Analysis in the Reddit Corpus arXiv:2608.19133v1 Announce Type: new Abstract: We present a novel application of embedding-based dynamic topic modeling techniques to detect and quantify topic drift at the comment level in a massive corpus. By leveraging pretrained language models to generate contextualized… 23 arXiv — NLP / Computation & Language research 10d ago Building real-time digital twin instances with Function+Data Flow: user evaluation and extension for iterative pipelines arXiv:2608.18480v1 Announce Type: cross Abstract: Digital twins (DTs) increasingly leverage artificial intelligence (AI) and machine learning (ML) pipelines, both to build real-time DTs from high-fidelity simulations and to instantiate them with historical data. However,… 10 r/MachineLearning community 10d ago Is it possible to fine-tune gemma4 A4B to generate complex legal principles of court decision? [D] I have a big database of local court decisions with a legal sentence which is like a paragraph summary of the doc. I've been tinkering with FTing for many days now, all results inconclusive never beating base except for a highly specific task where the eval was built around a… 34 r/LocalLLaMA community 10d ago DFlash2 speeds Qwen 3.8 27B up to 4 times llama.cpp pr #27342 adds dflash2, so i rented an rtx 6000 and ran the same four prompts through four decoding setups on qwen3.8 27B median results over the four tasks: baseline 47.4 tok/s mtp 114.7 tok/s dflash 99.3 tok/s dflash2 140.6. tok/s so on average 3x for dflash2 though… 33 The Information — AI news-outlet 10d ago For AI Companies, Unanswered Questions About White House Model Testing Plan In early August, the White House hosted tech companies including OpenAI, Anthropic and Google to share details on its new AI framework, which would encourage labs to voluntarily share their most advanced new models with the government before releasing them to the public. But… 14 arXiv — Machine Learning research 11d ago Hierarchical Data Selection via Manifold Coverage and Sparse Feature Coverage in LLM Post-training arXiv:2608.16927v1 Announce Type: new Abstract: As supervised fine-tuning data continues to scale, selecting high-value subsets from large candidate pools is crucial for reducing training cost and improving model performance. Existing methods often measure diversity directly in… 17 arXiv — Machine Learning research 11d ago Task Specialization Fine-Tuning for Contextual Reinforcement Learning arXiv:2608.17180v1 Announce Type: new Abstract: Contextual Reinforcement Learning (CRL) seeks to generalize classical RL by maximizing task coverage across a context space of related tasks. While prior works often train from scratch and rely on either multi-task learning for a… 38 arXiv — Machine Learning research 11d ago No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models arXiv:2608.17542v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs) learn world models by predicting future embeddings, but the objective admits a trivial solution of a constant encoder, so every practical system adds an anti-collapse mechanism… 23 arXiv — Machine Learning research 11d ago Leveraging Association Context Retrieval in Knowledge Edit- ing to Build White-Box Attacks on LLMs arXiv:2608.17836v1 Announce Type: new Abstract: As large language models (LLMs) are granted increasing autonomy, it is essential to investigate methods that can induce unsafe behavior. We propose a novel white-box attack inspired by locate-then-edit approaches from the field of… 25 Page 3 of 10 · 500 articles ← Newer Older →