News / #benchmark Tag Benchmark 500 articles archived under #benchmark · RSS Sign in to follow arXiv — NLP / Computation & Language research 19d ago Unified Hallucination Fuzzing for Multimodal Large Language Models arXiv:2608.07525v1 Announce Type: new Abstract: Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high-stakes applications. Existing evaluations, predominantly based on static benchmarks, suffer from… 16 arXiv — NLP / Computation & Language research 19d ago WuYuEval: A Multi-Level Benchmark for Large Language Models in Solid Waste Management arXiv:2608.07529v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as technical assistants, but their competence in solid waste management (SWM) remains difficult to assess because existing benchmarks emphasize general knowledge rather than… 6 arXiv — NLP / Computation & Language research 19d ago SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators arXiv:2608.07641v1 Announce Type: new Abstract: The rapid advancement of large language models has transformed survey writing from a months-long manual effort into an automated process. As generation scales, reliable evaluation becomes the bottleneck, and LLMs are increasingly… 36 arXiv — NLP / Computation & Language research 19d ago SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs arXiv:2608.07862v1 Announce Type: new Abstract: Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguistic diversity and culturally grounded safety risks present in other languages.… 21 arXiv — NLP / Computation & Language research 19d ago Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives arXiv:2608.08160v1 Announce Type: new Abstract: The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling. However, existing research has largely overlooked the critical challenge of… 14 arXiv — NLP / Computation & Language research 19d ago OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents arXiv:2608.08775v1 Announce Type: new Abstract: Agentic benchmarks aim to measure how well AI agents plan, search, execute, and recover within realistic multi-tool environments, but they are almost exclusively in English. As AI agents are globally deployed to a linguistically… 18 arXiv — NLP / Computation & Language research 19d ago LexKairos: Benchmarking Legal Temporal Capabilities in LLMs arXiv:2608.09106v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong performance across a wide range of legal tasks. In legal practice, time is a critical concept that governs the validity of statutes, the progression of legal cases, and the… 37 arXiv — NLP / Computation & Language research 19d ago Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments arXiv:2608.09128v1 Announce Type: new Abstract: LLM agents are increasingly deployed in multi-agent social settings where they must cooperate, negotiate, and adapt to other agents. Measuring and improving these social skills is hard because, unlike math or logic, social… 29 arXiv — NLP / Computation & Language research 19d ago UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers arXiv:2608.09209v1 Announce Type: new Abstract: Neural language models trained on large crowdsourced corpora frequently exploit spurious surface patterns tied to target labels without true linguistic or causal relevance, boosting benchmark performance while failing on… 33 arXiv — NLP / Computation & Language research 19d ago Temporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax Law arXiv:2608.09393v1 Announce Type: new Abstract: We identify and quantify temporal misgrounding: the systematic retrieval and citation of the currently in-force version of a legal article when the applicable version is an earlier or future one. Standard legal RAG treats the… 17 arXiv — NLP / Computation & Language research 19d ago Build it, Break it, Repeat: Benchmarking and improving LLM-manipulated disinformation detection in social media posts arXiv:2608.09510v1 Announce Type: new Abstract: Detecting machine-generated disinformation on social media is increasingly difficult as large language models (LLMs) make it easier to generate and rewrite misleading content at scale. Static benchmark evaluations, measuring… 38 arXiv — NLP / Computation & Language research 19d ago TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability arXiv:2608.09538v1 Announce Type: new Abstract: We introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof generation. TCS-Bench consists of theorem-proving tasks from papers published at top… 16 arXiv — NLP / Computation & Language research 19d ago Mawqif-v2: An Arabic Benchmark Dataset for Cross-Target Stance Detection arXiv:2608.09539v1 Announce Type: new Abstract: Publicly available Arabic datasets for target-specific stance detection remain limited, particularly for evaluating cross-target generalization. This paper presents the Mawqif-v2 Extension, consisting of 996 manually annotated… 9 arXiv — NLP / Computation & Language research 19d ago ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models arXiv:2608.09548v1 Announce Type: new Abstract: Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that ordinary question answering does not: a usable education-facing model is supposed… 32 arXiv — NLP / Computation & Language research 19d ago MDB-Link: Hierarchical Schema Linking for Multi-Database Text-to-SQL arXiv:2608.09588v1 Announce Type: new Abstract: Traditional Text-to-SQL research and benchmarks assume a known target database, overlooking settings in which a query must be routed within a large, heterogeneous database collection. We therefore study schema linking in a… 31 arXiv — NLP / Computation & Language research 19d ago Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness arXiv:2608.09766v1 Announce Type: new Abstract: Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlooks locale… 30 arXiv — NLP / Computation & Language research 19d ago PragMatch: Separating Pragmatic Incongruity from Cross-Modal Mismatch in Large Vision-Language Models arXiv:2608.09772v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) have demonstrated strong performance on multimodal benchmarks, yet it remains unclear whether they genuinely reason about relationships between images and text or rely on superficial… 8 Hugging Face Daily Papers research 19d ago Business Arena: Benchmarking LLM Agents in a Realistic Marketplace Abstract Business Arena evaluates LLM agents running a realistic cross-border shop, revealing large performance gaps versus human strategies and enabling detailed attribution of business decisions. Generated by thinkingmachines/Inkling-Small Running a business is a challenging… 21 Hugging Face Daily Papers research 19d ago SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Abstract SWE-Bench ProMax is a rigorously curated multilingual benchmark of large-scale code refactoring tasks that reveals substantial unsolved challenges for current AI coding agents. Generated by thinkingmachines/Inkling-Small As AI coding agents take on increasingly complex,… 21 r/LocalLLaMA community 19d ago Muse glimmer benchmark Little less smart than Qwen, but way fewer tokens per task.   submitted by   /u/NoFaithlessness951 [link]   [comments] 18 r/LocalLLaMA community 19d ago I made a web-design benchmark for local models (Muse Glimmer 30B vs Qwen 3.6 27b vs Deepseek V4 Flash 0731)   submitted by   /u/ShadyShroomz [link]   [comments] 34 r/LocalLLaMA community 19d ago Achievable 253 t/s - unsloth/Muse Glimmer 30B UD-Q5_K_M on a 5090 Benchmarked Muse Glimmer 30B on my RTX 5090 (32GB), 262k context, UD-Q5_K_M + dflash-kquant + mmproj. Workload Stock master + DFlash ngram-simple PR #26842 + DFlash Code patch 78 t/s 57 t/s 220-253 t/s Mixed agent turn 77 t/s 68 t/s 188-213 t/s Tool-call JSON 71 t/s 75 t/s… 15 r/LocalLLaMA community 19d ago I compared GGUF quants of Qwen3.6 27B to NVFP4, AWQ, AutoRound, and FP8 There's an interactive chart and some extra data in the blog post if you're interested. There are plenty of KL-divergence benchmarks for GGUF models, but most of them compare one GGUF quant against another. I wanted to know how those quants stack up against other commonly used… 28 Hugging Face Daily Papers research 20d ago Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination Abstract Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized. Contamination mitigation evaluation intervenes in the decoding process to suppress memorization and restore a contaminated model's genuine capability,… 14 Hugging Face Daily Papers research 20d ago Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events Abstract Personality-conditioned LLM agents (PC-Agents) are increasingly used in emotional support, social simulation, and role-playing, motivating the development of lifelong agents that remain coherent over extended interactions. A key component of such coherence is… 7 Hugging Face Daily Papers research 20d ago PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say Abstract LLM-based agents are rapidly advancing, autonomously invoking external tools to complete multi-step tasks for users. However, agents often acquire more sensitive information than the task requires. Existing privacy benchmarks audit what the agent's response or outgoing… 9 arXiv — Machine Learning research 20d ago Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks arXiv:2608.07335v1 Announce Type: new Abstract: Recent advancements in deep reinforcement learning have increasingly favored simplified, highly parallelized paradigms. Notably, the Parallelized Q-Network (PQN) algorithm achieves stable off-policy learning without relying on… 9 arXiv — Machine Learning research 20d ago UAV3DCrop: Benchmarking 3D Reconstruction in Repeated Multi-Angle UAV Crop Surveys arXiv:2608.06404v1 Announce Type: cross Abstract: Accurate 3D crop monitoring underpins data-driven precision agriculture by enabling field-scale analysis of plant structure, growth dynamics, and management response. Modern 3D reconstruction methods perform strongly on generic… 36 arXiv — NLP / Computation & Language research 20d ago Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events arXiv:2608.06485v1 Announce Type: new Abstract: Personality-conditioned LLM agents (PC-Agents) are increasingly used in emotional support, social simulation, and role-playing, motivating the development of lifelong agents that remain coherent over extended interactions. A key… 21 arXiv — NLP / Computation & Language research 20d ago Measuring the Cross-Lingual Comprehension Gap: How the language of the evidence shapes what language models understand arXiv:2608.06506v1 Announce Type: new Abstract: Language models are often evaluated as though capabilities demonstrated in English remain equally available when the same content is presented in other languages. Traditional multilingual benchmarks rarely isolate language while… 16 arXiv — NLP / Computation & Language research 20d ago TradeVerse: A Longitudinal Benchmark of Political Negotiation in International Trade arXiv:2608.06549v1 Announce Type: new Abstract: LLMs are increasingly being applied to tasks involving institutional and political texts, but existing benchmarks evaluate them on isolated documents or single tasks. In realpolitik, negotiations are longitudinal data, where… 11 arXiv — NLP / Computation & Language research 20d ago Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination arXiv:2608.07341v1 Announce Type: new Abstract: Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized. \textbf{Contamination mitigation evaluation} intervenes in the decoding process to suppress memorization and… 7 arXiv — NLP / Computation & Language research 20d ago LitTraceQA: A Benchmark for Multi-Stage Grounding and Verification in Scientific Question Answering arXiv:2608.07370v1 Announce Type: new Abstract: Scientific literature is increasingly used as a knowledge source for language models, retrieval-augmented generation systems, and research assistants, but answering research questions from papers requires more than fluent… 15 arXiv — NLP / Computation & Language research 20d ago StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt Injection arXiv:2608.06477v1 Announce Type: cross Abstract: Computer-use agents (CUAs) face a growing threat from indirect prompt injection, where adversarial instructions are planted in the environment such as web pages. In this paper, we introduce multi-step indirect prompt injection, a… 5 arXiv — NLP / Computation & Language research 20d ago How Should I Pick a Foundation Model for My Robot? In Favor of a Community Evaluation Framework for Social Robots arXiv:2608.06898v1 Announce Type: cross Abstract: Researchers who seek to build social robot applications on foundation models are faced with a difficult question: how should we pick a model? Public leaderboards offer little guidance: the demands of real-time, embodied social… 14 arXiv — NLP / Computation & Language research 20d ago GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks arXiv:2608.07411v1 Announce Type: cross Abstract: In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited insights into their generalization capabilities. In this paper, we present \benchName, a… 36 arXiv — NLP / Computation & Language research 20d ago SABRE: Scalable and Automated Benchmarking of VLMs under Stress arXiv:2608.07435v1 Announce Type: cross Abstract: Vision-language models (VLMs) are improving rapidly, but benchmark development lags behind, making weaknesses hard to identify. Building stress tests is costly: samples must satisfy controlled conditions, remain answerable, and… 22 arXiv — NLP / Computation & Language research 20d ago Dependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource Languages arXiv:2605.02608v2 Announce Type: replace Abstract: Transformer-based models achieve state-of-the-art dependency parsing for high-resource languages, yet their advantage over simpler architectures in low-resource settings remains poorly understood. We evaluate four parsers---the… 32 r/LocalLLaMA community 21d ago DeepSeek v4 Flash 0731 locally on CPU After seeing the benchmark results for the full release of DS v4 Flash 0731, I replaced my 2 x 16GB DDR4 ram sticks with 2 x 32GB DDR4 ram sticks to get a max supported of 128 GB RAM, in hope to be able to run GLM 5.2 equivalent model locally i.e. DS v4 Flash 0731 I also have… 25 r/LocalLLaMA community 21d ago Updated benchmark: Deepseek V4 Flash on SlopCodeBench (local) Howdy - I posted a benchmark here - https://www.reddit.com/r/LocalLLaMA/comments/1vbtiy7/deepseek_v4_flash_on_slopcodebench/ This was using the hosted API - since then I've been playing around with quants Here is the lastest benchmark -… 20 r/LocalLLaMA community 21d ago any reasonably fast public benchmarks I should run quants of deepseek flash 0731 on? I have various quants of this model and am curious how they perform. can anyone recommend which benchmark would be a good test case for quantization effects? Maybe that can be completed with about 1 million tokens?   submitted by   /u/nomorebuttsplz [link]  … 35 r/LocalLLaMA community 22d ago DeepSeek V4 Flash 0731 appreciation post I’m running DSV4F 0731 on dual spark, and honestly… wow. It’s an absolute workhorse, and the benchmarks are real. Everyday tasks with Hermes agent? Effortless. Coding tasks with OpenCode? I’m genuinely amazed at what it can handle. I can throw a two-hour coding session at it,… 12 r/LocalLLaMA community 22d ago Is anyone else finding DeepSeek-V4-Flash unreliable for non-coding tasks? (I am not a native speaker, written by myself, so please bear with me) I really want to like DeepSeek-V4-Flash-0731. But it has serious flaws that don't align with the high score on intelligence benchmarks. And those flaws render it useless unfortunetely for anything else… 21 r/LocalLLaMA community 22d ago DeepSeek V4 Flash 0731 - ARC-AGI Results   submitted by   /u/johnnyApplePRNG [link]   [comments] 20 r/LocalLLaMA community 23d ago LFM2.5-2.6B model+KV cache quantization report LFM2.5-2.6B is a new tiny model by LiquidAI, with benchmarks that put it head to head with much larger models. I've run llama-perplexity on many model GGUF quants, crossed with many KV cache quants, to understand the model's best overall quantization for any given amount of… 32 r/LocalLLaMA community 23d ago A llama.cpp PR makes Q2_0 3.0–3.6x faster on x86 CPUs, 8B decode goes 2.39 → 8.20 tok/s I was going through the current llama.cpp CPU PRs and #26348 stood out because this isn't the usual +5% kernel optimization. It adds an x86 VNNI implementation for the Q2_0 × Q8_0 dot product, and the author's controlled CPU-only benchmarks show roughly 3–3.6x higher throughput… 24 Hugging Face Daily Papers research 23d ago DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces Abstract Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured files, long documents, and multimedia. Existing benchmarks largely isolate structured querying, retrieval, or open-ended… 31 r/LocalLLaMA community 23d ago LabyrinthBench: a local-focused, judge-free LLM benchmark that measures context recall under interference for multi-step agentic tasks. LabyrinthBench measures the thing that actually kills long agent runs — whether a model can still use what it learned twenty turns ago — deterministically, with no LLM judge, on your own hardware, with a swappable harness for testing whatever context-management strategy you… 36 Hugging Face Daily Papers research 23d ago MameLoshnLM: Yiddish Language Model and Evaluation Benchmark Abstract We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have constrained progress in Yiddish… 5 r/LocalLLaMA community 23d ago Gemma 4 QAT could be improved further by Google aligning the QAT model to modern q4_k instead of q4_0 Hello, For the past few days I have been benchmarking Gemma 4 26b QAT UD Q4_K_XL extensively versus Bartowski's Q4_K_L. While QAT is certainly very effective and reducing memory consumption versus the highest q4 quant from him, I also have noticed some regressions in my own… 4 Page 8 of 10 · 500 articles ← Newer Older →