News / #rag Tag Rag 500 articles archived under #rag · RSS Sign in to follow arXiv — Machine Learning research 13d ago BCIJelly: An integrated ecosystem for brain-computer interface research arXiv:2608.13576v1 Announce Type: cross Abstract: Brain-computer interface (BCI) research relies on multistage computational pipelines, yet progress remains constrained by fragmented data formats, heterogeneous decoder implementations and hardware-specific deployment toolchains,… 38 arXiv — Machine Learning research 13d ago Language-Specific Gaps in AI Safety Training Datasets arXiv:2608.13695v1 Announce Type: cross Abstract: Large language model providers routinely cite multilingual safety benchmarks spanning a dozen or more languages as evidence that their models are safe for non-English-speaking users. We show that these collection-level coverage… 32 arXiv — NLP / Computation & Language research 13d ago GALA: Generation-Aware Cross-Modal Alignment for Text-to-Time-Series Synthesis arXiv:2608.13741v1 Announce Type: new Abstract: Synthesizing time series from natural language is emerging as the most expressive form of controllable time series generation. However, existing text-conditioned generators either take caption embeddings frozen from off-the-shelf… 8 arXiv — NLP / Computation & Language research 13d ago IterCOMP: Reasoning-aware Adaptive Prompt Compression for Multi-hop Question Answering arXiv:2608.13588v1 Announce Type: new Abstract: Multi-hop question answering requires complex reasoning across multiple evidence segments, which often overwhelms retrieval-augmented generation systems with lengthy and noisy contexts, thereby undermining both efficiency and… 18 arXiv — NLP / Computation & Language research 13d ago CLAIR-Fin: An Adversarial Multi-Agent Framework for Claim-Level Verification and Adaptive Debate in Cross-Modal Financial QA arXiv:2608.13706v1 Announce Type: new Abstract: Existing defenses against hallucination in retrieval-augmented and multi-agent pipelines remain partial: evidence is trusted despite modality disagreement, debate verifies an aggregate report rather than individual claims, and such… 8 arXiv — NLP / Computation & Language research 13d ago TeachMateGPT: A Multi-Agent Knowledge-Grounded Framework for Pedagogical Assessment Generation from Science Curriculum Materials arXiv:2608.13708v1 Announce Type: new Abstract: Automatically generating textbook-grounded assessment items can reduce science teachers' workload, but existing retrieval-augmented generation (RAG) systems rely on flat retrieval, support only single-question generation, lack… 21 arXiv — NLP / Computation & Language research 13d ago BM25-Augmented Many-Shot Translation for Low-Resource North-Eastern Indian Languages arXiv:2608.13722v1 Announce Type: new Abstract: This paper describes the University of Florida Gators submission to the WMT26 Low-Resource Indic Language Translation shared task. We adapt the retrieval-augmented many-shot translation pipeline from our AmericasNLP 2026 system to… 9 arXiv — NLP / Computation & Language research 13d ago How Much Do Legal RAG Systems Still Hallucinate? arXiv:2608.14210v1 Announce Type: new Abstract: Hallucination is a major challenge for retrieval-augmented generation (RAG) systems in the legal domain, where ungrounded answers can lead to serious consequences. To better understand this problem, we conduct a fine-grained… 26 arXiv — NLP / Computation & Language research 13d ago Stable Miscalibration in Large Language Models: A Practical View of High-Confidence Errors arXiv:2608.13591v1 Announce Type: cross Abstract: High-confidence errors in large language models are often treated as evidence of fragile internal inference. We study a different possibility: stable miscalibration, where a confident wrong answer remains locally stable under… 33 arXiv — NLP / Computation & Language research 13d ago Cross-Disciplinary Taxonomy and Modeling of Misunderstanding Generation, Amplification, and Detection, from Pragmatics to AI Agents arXiv:2608.13604v1 Announce Type: cross Abstract: Detection of misunderstanding is an urgent problem to solve because communication has moved away from real-time, in-person interaction and is increasingly handled by AI-mediated channels. This shift cuts communicators off from… 25 arXiv — NLP / Computation & Language research 13d ago Leveraging Few-Shot Learning and Large Language Models for Analyzing Blood Pressure Variations Across Biological Sex from Scientific Literature arXiv:2402.01826v2 Announce Type: replace Abstract: Current blood pressure (BP) technologies and standards were established decades ago, and these standards are still used worldwide today, often without adjusting BP readings for individual demographic factors such as sex and… 34 arXiv — NLP / Computation & Language research 13d ago Research-Oriented Human-Centric Evaluation for Foundation Models arXiv:2506.01793v2 Announce Type: replace Abstract: Most current evaluations of foundation models focus on objective benchmarks, such as knowledge coverage and reasoning accuracy, often overlooking users' subjective experiences in human-AI collaboration. To address this gap, we… 14 arXiv — NLP / Computation & Language research 13d ago Decoding Student Minds: Leveraging Conversational Agents for Psychological and Learning Analysis arXiv:2512.10441v2 Announce Type: replace Abstract: This paper presents a psychologically-aware conversational agent designed to enhance both learning performance and emotional well-being in educational settings. The system combines Large Language Models (LLMs), a knowledge… 38 arXiv — NLP / Computation & Language research 13d ago Adaptive Stopping for Multi-Turn LLM Reasoning arXiv:2604.01413v3 Announce Type: replace Abstract: Large Language Models (LLMs) increasingly rely on multi-turn reasoning and interaction, such as adaptive retrieval-augmented generation (RAG) and ReAct-style agents, to answer difficult questions. These methods improve accuracy… 35 r/LocalLLaMA community 15d ago GPU prices haven't stopped climbing for 3 weeks straight across the EU, here's the data hey again! I run a EU PC hardware price tracker PriceSquirrel , 25+ stores across 9 countries, and wanted to check: are GPU prices actually rising, or does it just feel that way? To make this defensible, I didn't just average "whatever's in stock" each day, that inflates the… 13 llama.cpp releases dev-tools 15d ago b10437 model : add support for MiniMaxText01ForCausalLM and MiniMaxM1ForCausalLM ( #27018 ) llama : support for MiniMax-Text-01 model chore : renames to match the other MiniMax models model : add logits mask as MiniMax-Text-01 embeddings tensor has zero-valued embeddings for tokens >=… 22 arXiv — Machine Learning research 16d ago MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning arXiv:2608.12724v1 Announce Type: new Abstract: Few-shot in-context learning (ICL) with multi-modal large language models (MLLMs) enables task adaptation without parameter updates, but its performance is highly sensitive to the quality and coverage of the selected… 27 arXiv — Machine Learning research 16d ago Diagnosing JEPA World Models with Action-Conditioned Predictive Consistency arXiv:2608.12939v1 Announce Type: new Abstract: Joint-embedding predictive architectures (JEPAs) learn world models that predict in a compact latent space rather than in pixels, reducing the pressure to model nuisance appearance. Yet this provides no guarantee against visual… 15 arXiv — NLP / Computation & Language research 16d ago LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning arXiv:2608.12321v1 Announce Type: new Abstract: When a salient surface cue competes with an implicit feasibility constraint, LLMs often fail -- but aggregate accuracy conflates genuine constraint inference with conservative defaulting. We formalize the distinction as conditional… 10 arXiv — NLP / Computation & Language research 16d ago Vision-Language Models are Fragile Multilingual Associators arXiv:2608.12333v1 Announce Type: new Abstract: Vision-language models must associate visual entities with textual attributes. Whether these associations or concept bindings remain stable when the language of the input changes is unexplored. We introduce M$^2$BIND, a benchmark… 37 arXiv — NLP / Computation & Language research 16d ago HC-RAG: Evidence-Centric Retrieval-Augmented Generation over Heterogeneous Financial Filings arXiv:2608.12335v1 Announce Type: new Abstract: Financial question answering over annual reports requires more than retrieving semantically similar passages. It often involves identifying relevant companies and fiscal years, locating standardized filing sections, collecting… 10 arXiv — NLP / Computation & Language research 16d ago The Embedder's Dilemma: LLMs Are Better, but at What Cost? arXiv:2608.12875v1 Announce Type: new Abstract: Should you replace your text-embedding pipeline with a large language model? We answer this with a controlled, cost-aware comparison of ten LLMs across six families and 26 embedding models (118M to 14B parameters) on 37 tasks… 24 arXiv — NLP / Computation & Language research 16d ago When Your Agent Opens the Chat App: Agent-Controlled Search over Raw Chat Logs Rivals Structured Memory arXiv:2608.12888v1 Announce Type: new Abstract: Agent-memory systems increasingly buy retrieval quality with structure, transforming raw conversation histories into summaries, embeddings, trees, or knowledge graphs before any question is asked. We ask how much of that benefit… 32 arXiv — NLP / Computation & Language research 16d ago HybridRAG-BN: A Retrieval-Augmented Framework with Fine-Tuned Verification for Bangla KBQA arXiv:2608.13004v1 Announce Type: new Abstract: Knowledge-base question answering (KBQA) systems rely on effective retrieval and reasoning mechanisms to generate accurate answers from external knowledge sources. However, developing reliable KBQA systems for low-resource… 19 arXiv — NLP / Computation & Language research 16d ago RAGSieve: Self-Referenced Local Contrast for Knowledge-Poison Detection in Retrieval-Augmented Generation arXiv:2608.13010v1 Announce Type: new Abstract: Retrieval-augmented generation treats an external corpus as inference evidence, allowing injected documents to promote attacker-chosen claims. Existing detectors depend on trusted references, specific attack artifacts, or global… 19 arXiv — NLP / Computation & Language research 16d ago Better Decomposition, Free Aggregation: A Synthesizer-Folding Framework for Multilingual Multi-Hop Question Answering arXiv:2608.13160v1 Announce Type: new Abstract: Multilingual retrieval-augmented generation (mRAG) equips large language models with access to globally distributed external knowledge for complex multilingual question answering. Recent approaches either translate retrieved… 15 arXiv — NLP / Computation & Language research 16d ago GEM: A Generative Embedding Model Bridging Reasoning and Retrieval arXiv:2608.13200v1 Announce Type: new Abstract: Modern LLMs excel at reasoning and instruction following, enabling users to express complex and diverse information needs. However, conventional retrievers largely rely on surface-level matching between queries and documents,… 32 arXiv — NLP / Computation & Language research 16d ago When Should Multi-Round RAG Stop? Structured Stopping Judgments and Retrieval Reduction in Search-R1 arXiv:2608.13237v1 Announce Type: cross Abstract: Multi-round retrieval-augmented generation (RAG) must decide when to stop searching as evidence accumulates. Because the deployed policy is determined by the first STOP on each trajectory, this is a sequential selection problem… 19 arXiv — NLP / Computation & Language research 16d ago OmniScientist: An Omni-Modal Omni-Discipline AI Scientist arXiv:2608.13558v1 Announce Type: cross Abstract: Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to manuscript preparation. Yet workflow coverage alone does not… 6 Simon Willison community 16d ago llm-gemini 0.33 Release: llm-gemini 0.33 It's been a while since the last llm-gemini release. This version of the plugin adds support for today's Gemini 3.7 Flash release, plus gemini-3.6-flash , gemini-3.5-flash-lite and two embedding models gemini-embedding-2 and gemini-embedding-001 . The… 19 Hugging Face official-blog 16d ago Record, train, and deploy from one place with Strands Agents, LeRobot, and Hugging Face Storage Buckets Back to Articles a]:hidden"> Record, train, and deploy from one place with Strands Agents, LeRobot, and Hugging Face Storage Buckets Enterprise Article Published August 13, 2026 Upvote 4 Sundar Raghavan rsundaraws amazon Steven Palma imstevenpmwork amazon Cagatay Cali cagataydev… 34 arXiv — NLP / Computation & Language research 17d ago Weightless Fine-Tuning: Personalizing LLMs via Logit-Space Transport arXiv:2608.11342v1 Announce Type: cross Abstract: Supervised fine-tuning (SFT) is a standard approach for adapting LLMs to a target distribution, but in settings such as personalization, where each author requires separate weight access, optimization, storage, and retraining,… 35 arXiv — Machine Learning research 17d ago Disentangling the Expressivity of RoPE arXiv:2608.11909v1 Announce Type: new Abstract: Two accounts recur in explanations of the success of rotary position embeddings (RoPE). Expressivity studies associate periodic position information with modular predicates, whereas mechanistic and long-context studies emphasize… 16 arXiv — Machine Learning research 17d ago Earth observation embeddings are effective sub-grid descriptors for probabilistic weather downscaling arXiv:2608.12271v1 Announce Type: new Abstract: Global weather reanalyses and forecasts resolve the evolving atmospheric state on coarse grids, but site-specific applications require predictions at arbitrary locations where near-surface conditions also depend on unresolved… 21 arXiv — NLP / Computation & Language research 17d ago Diffuse to Compress: Leveraging Diffusion LMs for Lossless Compression arXiv:2608.11249v1 Announce Type: new Abstract: We study the problem of lossless text compression, motivated by the rapid growth in the collection and storage of digital textual data - including plain text, source code, and structured formats such as XML - and by recent advances… 37 arXiv — NLP / Computation & Language research 17d ago LazyTrain: Limited-resource Allocation toward Zero-waste Yield Optimization in Large Language Model Training arXiv:2608.11919v1 Announce Type: new Abstract: Training large language models on limited hardware is increasingly a scheduling problem across GPU compute, host memory, PCIe transfer, and storage bandwidth. Existing offloading systems reduce GPU residency, and MegaTrain shows… 19 arXiv — NLP / Computation & Language research 17d ago LODESTAR: Trustworthy Entropy Is Navigated, Not Merely Measured -- Reinforced Polarizer Keeps a Frozen LLM from Being Confidently Misled by the Wrong Evidence arXiv:2608.11922v1 Announce Type: new Abstract: Predictive-distribution entropy makes a strong selection rule in retrieval-augmented question answering: across five QA benchmarks, keeping the candidate answer that a frozen respondent LLM produces with the lowest answer-token… 8 arXiv — NLP / Computation & Language research 17d ago QV-PIC: Query-Aware Visual Position-Independent Caching for Efficient RAG Serving arXiv:2608.12121v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) repeatedly prefills identical text chunks across queries, incurring redundant computations. Position-Independent Caching (PIC) mitigates it by reusing precomputed Key-Value (KV) across… 27 arXiv — NLP / Computation & Language research 17d ago SAG: SQL-Retrieval Augmented Generation with Query-Time Dynamic Hyperedges arXiv:2608.12129v1 Announce Type: new Abstract: While retrieval-augmented generation (RAG) has proven effective at giving LLMs access to external knowledge, mainstream dense-retrieval implementations remain inherently limited in handling structured constraints and multi-hop… 29 arXiv — NLP / Computation & Language research 17d ago A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench arXiv:2608.12138v1 Announce Type: new Abstract: General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed… 10 arXiv — NLP / Computation & Language research 17d ago Investigating Learner-Aware Design of LLM-Generated Educational Feedback arXiv:2602.11650v2 Announce Type: replace Abstract: Although large language models (LLMs) show promise for generating educational feedback, it remains unclear how feedback should be designed (e.g., tone and information coverage) to support answer revision and learner acceptance… 12 Hugging Face Daily Papers research 17d ago CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG Abstract CoinRAG improves retrieval-augmented generation efficiency and accuracy by reusing fine-grained semantic nugget caches instead of full chunks. Generated by thinkingmachines/Inkling-Small Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited… 4 r/MachineLearning community 17d ago chessformer_lens demo: ablating 1 of a chess transformer's 128 attention heads makes the model stop finding Morphy's queen sacrifice [P] https://i.redd.it/ipz7i6ife1jh1.gif Notebooks to replicate on github!   submitted by   /u/Weird-Asparagus4136 [link]   [comments] 22 Hugging Face official-blog 17d ago Introducing OlmoEarth embeddings: Custom embedding exports from OlmoEarth Studio for downstream analysis Back to Articles a]:hidden"> Introducing OlmoEarth embeddings: Custom embedding exports from OlmoEarth Studio for downstream analysis Enterprise Article Published August 12, 2026 Upvote 1 Kyle Wiggers Ai2Comms allenai 📄 Tech Report: https://allenai.org/papers/olmoearth | 📊… 9 NVIDIA Developer Blog official-blog 17d ago How to Choose Full-Stack Observability for NVIDIA AI Factories AI infrastructure spans multiple layers, from compute and networking to storage, orchestration, and applications. When performance degrades, identifying the... 34 Hacker News — AI on Front Page community 18d ago Facebook is paying controversial creators to produce rage-bait content Article URL: https://www.abc.net.au/news/2026-08-06/ragebait-how-facebook-is-paying-controversial-creators/106940696 Comments URL: https://news.ycombinator.com/item?id=49269818 Points: 234 # Comments: 130 10 r/LocalLLaMA community 18d ago RAG for regular users? One of the reasons I got into local LLMs was the possibility of getting answers using my own documents and books (a few hundreds) instead of having to search through them manually. However since I'm not a data specialist or an engineer, RAG projects were too difficult for me,… 17 arXiv — NLP / Computation & Language research 18d ago Procedural Fairness Failures in RLHF from Preference Averaging arXiv:2608.10126v1 Announce Type: cross Abstract: Reinforcement Learning from Human Feedback (RLHF) aggregates heterogeneous preferences into a single reward model, assuming preference homogeneity. When preferences are heterogeneous, this aggregation induces a procedural… 13 arXiv — Machine Learning research 18d ago Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-Tuning arXiv:2608.10473v1 Announce Type: new Abstract: Offline-to-online (O2O) reinforcement learning aims to leverage policies pretrained on static datasets while improving them through online interaction. However, directly reusing an offline-trained critic can hinder online… 19 arXiv — NLP / Computation & Language research 18d ago The Multilingual Quantization Tax: Structural Collapse and Typological Fragility in Edge SLMs arXiv:2608.09941v1 Announce Type: new Abstract: While 4-bit weight quantization is critical for deploying Small Language Models (SLMs) on edge devices, evaluations of the resulting performance degradation-the quantization tax-remain overwhelmingly English-centric. We present a… 30 Page 5 of 10 · 500 articles ← Newer Older →