News / #rag Tag Rag 500 articles archived under #rag · RSS Sign in to follow arXiv — NLP / Computation & Language research 25d ago ATFlash: Per-RoPE-Wavelength Attention Windows for Compute/Memory-Efficient LLM Inference arXiv:2608.02947v1 Announce Type: cross Abstract: The attention score with rotary position embeddings (RoPE) decomposes exactly into a sum over its 2D-rotation frequency pairs, and each pair's wavelength limits how far it can discriminate position. Aligned with this structure,… 19 arXiv — Machine Learning research 25d ago Lightweight Chunk Selection for Mobile Retrieval-Augmented Generation arXiv:2608.03148v1 Announce Type: new Abstract: RAG improves the factual grounding of LLM by incorporating external knowledge, but deploying RAG on mobile and edge devices remains challenging because retrieved context increases computation and memory. A direct way to reduce this… 26 arXiv — Machine Learning research 25d ago The Tell-Tale Trace: Detecting Reasoning Failures in LLMs Using Chain-of-Thought Dynamics arXiv:2608.03291v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning improves large language model (LLM) performance while also providing an observable interface to the model's reasoning process. Existing approaches that leverage verbalized CoTs to monitor reasoning… 6 arXiv — Machine Learning research 25d ago Tight Worst-Case Bounds for the Smallest Eigenvalue of ReLU NTK Gram Matrices arXiv:2608.03368v1 Announce Type: new Abstract: For $n$ unit vectors $x_1,\ldots,x_n \in \mathbb{R}^d$, we study the continuous ReLU derivative Gram matrix $H$, whose entries are obtained by averaging pairwise gated inner products over a standard Gaussian direction. Writing $… 10 arXiv — NLP / Computation & Language research 25d ago JudgeArena: A Unified Framework for Reproducible LLM-Judge Evaluation arXiv:2608.02620v1 Announce Type: new Abstract: LLM-as-a-judge evaluation has become a dominant paradigm for ranking language models, yet the ecosystem remains fragmented: most benchmarks ship their own code base, hardcode a specific closed-model judge, and support a single… 36 arXiv — NLP / Computation & Language research 25d ago ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads arXiv:2608.02703v1 Announce Type: new Abstract: Weight-only quantization substantially reduces the storage of large language model (LLM) transformer blocks, but practical backends often retain the final language-modeling head (LM-head) in BF16 or FP16. Quantizing this projection… 21 arXiv — NLP / Computation & Language research 25d ago MoEGen: Mixture-of-Experts for Instance-Adaptive LoRA Generation arXiv:2608.03275v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) enables efficient adaptation of large language models, but existing MoE-based PEFT methods typically improve capacity by storing multiple full LoRA experts, causing adapter storage to grow… 15 arXiv — NLP / Computation & Language research 25d ago Efficient Multilingual Neural Machine Translation via Corpus-Driven Vocabulary Pruning: An English-Arabic Case Study arXiv:2608.03480v1 Announce Type: new Abstract: The adoption of large pre-trained multilingual models for neural machine translation (MNMT) faces a major challenge: excessive memory and computational consumption due to overly large vocabularies and embedding layers. Although… 20 arXiv — NLP / Computation & Language research 25d ago Beyond Initialization Loss: A Systematic Study of Token Embedding Initialization Strategies for LLM Vocabulary Extension arXiv:2608.03494v1 Announce Type: new Abstract: Vocabulary extension is an efficient way to adapt pretrained large language models (LLMs) to new languages, but the initialization of newly added token embeddings can strongly affect continued pre-training (CPT) efficiency. We… 11 arXiv — NLP / Computation & Language research 25d ago ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels arXiv:2608.03507v1 Announce Type: new Abstract: Historical language change affects morphology, syntax, semantics, and pragmatics, yet computational studies typically examine these levels with incompatible representations and therefore cannot determine whether they evolve… 23 arXiv — NLP / Computation & Language research 25d ago Language-Specialized Multi-Teacher On-Policy Distillation for Multilingual LLM-Based ASR arXiv:2608.03610v1 Announce Type: new Abstract: Modern LLM-based ASR systems have established multilingual capability as a standard feature, leveraging large-scale multilingual corpora and LLMs' cross-lingual knowledge to achieve competitive performance across multilingual… 9 arXiv — NLP / Computation & Language research 25d ago SciRet: A Compute-Aware Empirical Study of Retrieval and Reranking for Scientific RAG arXiv:2608.03860v1 Announce Type: new Abstract: We introduce SciRet, a compute-aware empirical study of retrieval-augmented generation for scientific question answering over CORD-19. Rather than proposing a new model, we evaluate a fixed scientific RAG pipeline across three… 38 Hugging Face Daily Papers research 25d ago AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling Abstract Language remains an outlier in generative modeling: while images, video, and audio are increasingly modeled in continuous latent spaces, text generation still relies predominantly on discrete tokens. Existing continuous language models either inherit embedding spaces… 32 r/LocalLLaMA community 25d ago Kimi K3 full model running on 16x GB10 cluster at 20+tps Kimi K3 full model running on 16x GB10 cluster at 20+tps average (llama-benchy coherent corpus) 38tps peak, 750tps prefill. This is the first run of full k3 with dspark on my cluster. I will be doing some tests and try tp speed this up. As soon as it looks ready I'll publish the… 13 MIT News — AI research 25d ago Solving the solvent problem By focusing on electrolytes, MIT scientists are making sodium-metal batteries a more practical energy storage option. 13 r/MachineLearning community 25d ago NeurIPS 2026 post-rebuttal score distribution poll [D] As the title suggests, because there's no data on Papercopilot yet, and people have been talking about the scores being lower in general than last year, I thought it could be interesting to survey the average score distribution after the rebuttal phase (not considering… 26 Hacker News — AI on Front Page community 26d ago AI-Generated Images Discourage Me from Reading Your Blog Article URL: https://nelson.cloud/ai-generated-images-discourage-me-from-reading-your-blog/ Comments URL: https://news.ycombinator.com/item?id=49167113 Points: 387 # Comments: 229 21 arXiv — Machine Learning research 26d ago A Physics-Chemistry-Informed Neural Network (PCINN) for Real-Time Spatial-ALD Coverage Prediction and Reliable Kinetics Inversion arXiv:2608.00212v1 Announce Type: new Abstract: Spatial atomic layer deposition (SALD) is a leading atmospheric-pressure, high-throughput route to industrial ALD, but design and control are limited by the cost of predicting surface coverage: high-fidelity CFD is far too slow for… 32 arXiv — Machine Learning research 26d ago HP-JEPA: Hierarchical Partitioning for Multi-Resolution Graph Joint-Embedding Predictive Learning arXiv:2608.00491v1 Announce Type: new Abstract: Graph self-supervised learning aims to learn transferable representations from large-scale unlabeled graph data. Joint-embedding predictive architectures (JEPAs) avoid explicit negative-pair construction and raw-input… 19 arXiv — Machine Learning research 26d ago Deep Learning CNN and Recurrence Analysis for Alpha Gamma EEG Biomarkers in Fragile X Syndrome arXiv:2608.00835v1 Announce Type: new Abstract: Fragile X Syndrome (FXS) is a neurodevelopmental disorder caused by reduced expression of fragile X mental retardation protein (FMRP), leading to disrupted synaptic plasticity, cortical hyperexcitability, and impaired network… 37 arXiv — Machine Learning research 26d ago UpliftBench: Revealing Outcome-Regime and Objective Mismatch in Uplift Evaluation arXiv:2608.00915v1 Announce Type: new Abstract: Uplift modeling (conditional-average-treatment-effect estimation) drives personalized targeting, yet published uplift benchmarks frequently disagree on which estimator performs best; we show the disagreement is substantially about… 33 arXiv — Machine Learning research 26d ago Beyond Gene Reconstruction: Learning Cell Representations through Complementary Transcriptomic Views arXiv:2608.00985v1 Announce Type: new Abstract: The rapid growth of single-cell transcriptomic data has enabled the development of foundation models pretrained primarily by reconstructing masked expression values. This objective encourages these models to learn gene dependencies… 35 arXiv — Machine Learning research 26d ago Differentiable Lifting for Topological Neural Networks arXiv:2608.01160v1 Announce Type: new Abstract: Topological neural networks (TNNs) enable leveraging high-order structures on graphs (e.g., cycles and cliques) to boost the expressive power of message-passing neural networks. In turn, however, these structures are typically… 36 arXiv — Machine Learning research 26d ago FedChronos: Federated Fine-Tuning of Time-Series Foundation Models for Privacy-Preserving Commodity Price Forecasting arXiv:2608.01290v1 Announce Type: new Abstract: Time-series foundation models (TSFMs) such as Chronos have demonstrated strong forecasting capabilities across domains, yet adapting them to institutionally fragmented settings, where data cannot be centralized due to regulatory,… 18 arXiv — Machine Learning research 26d ago Conformalized Large Language Models under Configuration Shift arXiv:2608.01460v1 Announce Type: new Abstract: Conformal prediction (CP) is a distribution-free framework for uncertainty quantification that has recently been adapted to large language models (LLMs), providing prediction sets with finite-sample coverage guarantees under… 22 arXiv — Machine Learning research 26d ago Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval arXiv:2608.01481v1 Announce Type: new Abstract: Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2.0 audio embeddings. Yet their weights do not map… 38 arXiv — NLP / Computation & Language research 26d ago Averaging Bias: Human Faithfulness Annotations are not Locally Faithful arXiv:2608.00205v1 Announce Type: new Abstract: Evaluation of faithfulness of text summarization treats a model generated summary as faithful only if every of its sentences is supported by the source document: a strict conjunctive rule under which a single unsupported sentence… 29 arXiv — NLP / Computation & Language research 26d ago Cross-Task Dissociation in Frontier Vision-Language Model Theory of Mind arXiv:2608.00261v1 Announce Type: new Abstract: Do frontier vision-language models present a coherent Theory-of-Mind (ToM) profile across tasks, matching the same human reference group, or does that profile fragment from one paradigm to the next? We evaluate a shared panel of… 22 arXiv — NLP / Computation & Language research 26d ago Sixteen models, fewer than two voices: measuring ensemble dispersion where no answer is uniquely correct arXiv:2608.00285v1 Announce Type: new Abstract: Sixteen language models drawn from ten families produced, on average, the semantic diversity of 1.69 distinct formulations of a psychotherapeutic case, against a single-model baseline of 1.43 from one model's own runs. Ensembles… 11 arXiv — NLP / Computation & Language research 26d ago Writing-System-Level Tokenizer Adaptation for Byte-Level BPE arXiv:2608.00582v1 Announce Type: new Abstract: Pretrained byte-level BPE tokenizers can segment underrepresented languages inefficiently. Replacing a tokenizer changes the meaning of nearly every token ID, while vocabulary expansion enlarges the model's embedding and output… 6 arXiv — NLP / Computation & Language research 26d ago Verification Without Sufficiency: Per-Chunk Filtering Fails on Multi-Hop RAG, and Decomposition Repairs It arXiv:2608.00585v1 Announce Type: new Abstract: Verification for retrieval-augmented generation usually scores each retrieved chunk and drops the ones that fail. We show this cannot work for multi-hop questions, and show what does. Per-chunk scoring assumes one chunk is a… 9 arXiv — NLP / Computation & Language research 26d ago TreeProbe : A Tibetan Medicine Benchmark for Cultural Bias in LLMs arXiv:2608.00640v1 Announce Type: new Abstract: Large language models are increasingly viewed as a potential means of mitigating global health inequities, yet their outputs often reflect dominant high-resource medical traditions and provide limited coverage of traditional… 38 arXiv — NLP / Computation & Language research 26d ago Select-And-Extract: A Lightweight Plugin for Retrieval-Augmented Generation arXiv:2608.00658v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) for language model (LM) systems fundamentally has two failure modes: retrieval failure and reading failure. The former fails to recall the right pieces of information from the external corpus,… 35 arXiv — NLP / Computation & Language research 26d ago RAGOCR: Optical Compression of Retrieval-Augmented Text via Visual Representation arXiv:2608.00765v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has become essential for knowledge-intensive question answering, yet scaling RAG pipelines remains challenging due to the prohibitive computational cost of processing lengthy retrieved contexts.… 31 arXiv — NLP / Computation & Language research 26d ago Does Machine "know" interpersonal pragmatics? Evidence from MARBERT's learning of emoji pragmatics in Arabic digital discourse arXiv:2608.01174v1 Announce Type: new Abstract: This study examines Transformer-based models' ability to learn emoji pragmatics in Arabic digital discourse (ADD), providing evidence from MARBERT's behavior with interpersonal pragmatic functions (IPFs). A corpus of 8,504 unique… 19 arXiv — NLP / Computation & Language research 26d ago ACE-GraphRAG: Agentic Context Engineering for Hierarchical GraphRAG arXiv:2608.01269v1 Announce Type: new Abstract: Hierarchical Graph Retrieval-Augmented Generation (GraphRAG) organizes corpus knowledge at multiple levels of granularity, yet fixed context construction may fail to translate these multi-resolution representations into a context… 28 arXiv — NLP / Computation & Language research 26d ago RH-RAG: Trustworthy Long-Form Generation for Privacy-Constrained Settings arXiv:2608.01311v1 Announce Type: new Abstract: Generating long-form content from extensive internal reports remains challenging for organizations operating under strict privacy and security constraints, where proprietary cloud-based LLM APIs are often not viable. While locally… 36 arXiv — NLP / Computation & Language research 26d ago Discriminative Axis, Not Data Volume: What a Contrastive Corpus Teaches an Audio Embedding arXiv:2608.01560v1 Announce Type: new Abstract: Scaling the corpus is the default remedy when a contrastive representation lacks an attribute. We report a case where it does nothing, and identify what does: adding a lexical-speech round to a frozen-base multimodal embedding… 29 arXiv — NLP / Computation & Language research 26d ago DocNavRAG: Document-Structured Graph RAG with Stateful Evidence Construction for Complex Document Question Answering arXiv:2608.01565v1 Announce Type: new Abstract: Answering complex questions over large document collections requires assembling complementary evidence across sections and documents. GraphRAG offers structured retrieval but typically uses fixed traversal, while agentic RAG… 21 arXiv — NLP / Computation & Language research 26d ago RING: Retrieval-Internalized Generation for Continual Large-Scale Knowledge Injection arXiv:2608.01630v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) improves factuality but adds latency and engineering overhead at serving time. We propose RING (Retrieval-Internalized Generation), a holistic paradigm spanning both architecture and training… 8 Hugging Face Daily Papers research 26d ago UEmbed: Unified Sparse and Dense Multimodal Embeddings Abstract Sparse retrieval underpins modern search systems, from web search to retrieval-augmented generation. Existing work has introduced Learned Sparse Retrieval (LSR) to push beyond exact lexical matching toward richer semantics. Yet LSR has so far remained tied to… 5 OpenAI Python SDK releases dev-tools 26d ago v2.53.0 2.53.0 (2026-08-03) Features api: Add gpt-5.5 and tool name/namespace to Responses types ( #3569 ) ( dd1202d ) Bug Fixes ci: avoid NumPy source builds and duplicate HTTPX coverage ( #3573 ) ( b58332f ) 8 r/LocalLLaMA community 26d ago 'I ran my own benchmarks on it' seems to be pretty common comment around here. How about dedicating a thread for this and sharing? Of course, the concern is that in the end, this thread will be fed into the models' training data, but I feel benchmarking isn't so open and very fragmented.   submitted by   /u/jinnyjuice [link]   [comments] 38 r/MachineLearning community 26d ago ARPL — runtime ISA/topology detection for llama.cpp on ARM (built for Snapdragon 8 Elite) [r] I've been working on this for a while and finally pushed a public version. The problem: llama.cpp runs fine on ARM phones, but it doesn't know anything about the specific chip it's on. Same thread count, same context params, whether you're on a Snapdragon 8 Elite or a… 24 r/LocalLLaMA community 27d ago I CANNOT believe I've got DeepSeek-V4-Flash-0731, a frontier model, running on my home PC. Insane! So this is the stuff of absolute insanity. In less than 20 months we've gone from super expensive cloud models only, to being able to run a Q3 quant of DeepSeek on an Intel Windows PC with a very average 24GB of VRAM. No wonder the big boys are panicking (and yes it's slow as… 10 NVIDIA Developer Blog official-blog 27d ago NVIDIA Vera Storage Benchmarks: Faster Encryption, Compression, Integrity Checking, and Recovery for AI-Native Storage Storage is an active part of every agentic AI workflow. As agents retrieve enterprise knowledge, access persistent memory, reuse key-value (KV) cache data,... 17 Hugging Face Daily Papers research 27d ago SAF-OPD: Stable Advantage Fusion for On-Policy Distillation Abstract Reinforcement learning with verifiable rewards (RLVR) broadcasts a single response-level reward to every token, while on-policy distillation (OPD) scores each token against a stronger teacher for a dense advantage but caps performance at teacher quality and discourages… 8 Hugging Face Daily Papers research 27d ago One Future, Every Robot: Label-Efficient Collective-State Prediction with Decentralized JEPA Abstract Can every robot in a swarm predict the same future collective state from only local observations and bandwidth-limited messages? We formulate this as decentralized shared-state prediction and introduce Collective-State JEPA (CS-JEPA), a recurrent joint-embedding… 29 r/LocalLLaMA community 27d ago I benchmarked classic vector RAG vs Google's new OKF format vs both combined — same corpus, same 7 questions, all local (Ollama + ChromaDB) Google Cloud published OKF (Open Knowledge Format) on June 12th — a spec for storing curated knowledge as a directory of markdown files with YAML frontmatter. One concept per file, linked to each other, with an index.md for progressive disclosure. The only required field is… 7 arXiv — Machine Learning research 27d ago Guarantees on Dynamical System Distinguishability for LLM Token Generation arXiv:2607.28667v1 Announce Type: new Abstract: Recent work has shown that classifying large language models (LLMs)' responses can be distinguished by modeling token embeddings as trajectories of a black-box dynamical system (DS) and comparing prediction residuals of two DSs.… 28 Page 8 of 10 · 500 articles ← Newer Older →