News / #benchmark Tag Benchmark 500 articles archived under #benchmark · RSS Sign in to follow arXiv — NLP / Computation & Language research 25d ago JudgeArena: A Unified Framework for Reproducible LLM-Judge Evaluation arXiv:2608.02620v1 Announce Type: new Abstract: LLM-as-a-judge evaluation has become a dominant paradigm for ranking language models, yet the ecosystem remains fragmented: most benchmarks ship their own code base, hardcode a specific closed-model judge, and support a single… 36 arXiv — NLP / Computation & Language research 25d ago Knowing the Form, Not the Function: Automatically Auditing Answer--Authority Decoupling in Legal Benchmarks arXiv:2608.02621v1 Announce Type: new Abstract: Legal benchmarks typically score final answers even when models also state legal authority. We test whether answer correctness can serve as a proxy for authority grounding. Under ordinary reasoning prompts that did not request… 33 arXiv — NLP / Computation & Language research 25d ago FLARE: Few-shot Learning-based Adaptive Reflective Engine arXiv:2608.02919v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in complex, compound AI systems where performance hinges on the quality of prompts. Recent state-of-the-art optimizers like GEPA (Genetic-Pareto) have argued that reflective… 32 arXiv — NLP / Computation & Language research 25d ago Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks arXiv:2608.02966v1 Announce Type: new Abstract: Most multiple-choice question (MCQ) benchmarks evaluate Large Language Models (LLMs) only by whether they select the correct answers. This binary scoring treats all incorrect responses alike, even though an LLM's preferences among… 22 arXiv — NLP / Computation & Language research 25d ago VIVID: A Culturally Grounded Benchmark Exposing the Figurative Language Gap in Vietnamese NLP arXiv:2608.03095v1 Announce Type: new Abstract: We present VIVID (Vietnamese Idioms for Validation and Interpretation Depth), the first systematic benchmark for evaluating culturally grounded figurative language understanding in Vietnamese. VIVID comprises 1,636 idioms and… 24 arXiv — NLP / Computation & Language research 25d ago Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks arXiv:2608.03340v1 Announce Type: new Abstract: Predicting LLM's capabilities on real-world tasks is essential, yet the extent to which performance on commonsense benchmarks predicts downstream performance remains underspecified. To establish the practical usability of widely… 7 arXiv — NLP / Computation & Language research 25d ago ArtECulture: Benchmarking Culture-Conditioned Visual Emotion Understanding in Multimodal Large Language Models arXiv:2608.03358v1 Announce Type: new Abstract: Existing visual emotion understanding methods typically ignore cultural variations in emotional perception. We introduce culture-conditioned visual emotion understanding, a task that predicts the culture-specific emotional… 10 arXiv — NLP / Computation & Language research 25d ago M-GATE: Multilingual Grammar, Accuracy in Translation, and Efficiency Benchmark for Large Language Models arXiv:2608.03803v1 Announce Type: new Abstract: Multilingual language models are deployed across a hundred or more languages, yet most benchmarks test whether a model can perform a task _in_ a language rather than whether it commands the language itself, conflating fluency with… 14 arXiv — NLP / Computation & Language research 25d ago VIBE: A VAD-Informed Benchmark for Entity-Centered Affective Profiling of Large Language Model Outputs arXiv:2608.03810v1 Announce Type: new Abstract: Large language models routinely describe socially salient targets, including political figures, countries, religions, organizations, historical events, and social groups, encoding affective framing alongside factual content: a… 11 arXiv — NLP / Computation & Language research 25d ago MultiGlobeQA: A Multilingual and Globally Diverse Benchmark for Geospatial Reasoning arXiv:2608.03882v1 Announce Type: new Abstract: Geospatial reasoning, i.e., computing distances, containment, and other spatial relations over real-world entities, is central to navigation and logistics, yet large language models (LLMs) struggle with the required geometric and… 17 arXiv — NLP / Computation & Language research 25d ago PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents arXiv:2608.04003v1 Announce Type: new Abstract: Recursive self-improvement requires agents to turn accumulated experience into better future behavior. Personal AI agents offer a concrete setting for studying this capability because they retain preferences, task histories, tool… 36 arXiv — NLP / Computation & Language research 25d ago WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament arXiv:2608.04008v1 Announce Type: new Abstract: Benchmarks that measure the forecasting ability of large language models are almost always retrospective: the event has happened, the answer is somewhere on the Web, and the evaluation must defend itself against memorisation. We… 23 arXiv — NLP / Computation & Language research 25d ago SocietyBench: Forecasting Counterfactual Social-World Evolution arXiv:2608.04009v1 Announce Type: new Abstract: Large language models (LLMs), and the agents built on top of them, are now benchmarked heavily on whether they can finish a task -- fix a bug, drive a browser, operate a GUI. A complementary social ability, namely how well a model… 11 arXiv — NLP / Computation & Language research 25d ago Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity arXiv:2608.02665v1 Announce Type: cross Abstract: A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that reading is faithful: when an item's intent is held fixed and only its meaning-preserving… 38 r/LocalLLaMA community 25d ago GPT-X2.5-135M scores 3rd place on Open SLM Leaderboard on Huggingface, Beating Facebook's MobileLLM-R1-140M   submitted by   /u/Megneous [link]   [comments] 28 r/LocalLLaMA community 25d ago Local LLM 35B MoE — Real-world coding benchmarks (Qwen vs Ornith vs KAT) I’ve been running a fairly opinionated evaluation loop on ~35B A3B/MoE-class models for coding over the past few months. Not synthetic benchmarks: actual dev workflows, iterative debugging, refactoring passes, and failure recovery. Here’s where things stand for me: Qwen 3.6 (35B… 9 r/LocalLLaMA community 26d ago Design systems from code alone - Without external images, Ling-3.0-flash generated webpages across Bauhaus, Bohemian, acid design, and more—using CSS gradients, SVG paths, typography, and layout to preserve each visual language. Weights went up today so this is downloadable now, MIT, ~128GB for the official FP8. I ran these on the API before that landed, so treat it as a preview of what you'd be pulling rather than a local benchmark   submitted by   /u/AcanthisittaOk1699 [link]   [comments] 32 r/LocalLLaMA community 26d ago inclusionAI/Ling-3.0-flash · Hugging Face The Ling-3.0-flash MoE is now open-weighted at 124B A5B params. I know the original announcements were before the Kimi K3, DeepSeek-V4-Flash and Qwen3.8 hype, but this model might still have a good niche for itself due to its sizing. Discussion on the benchmarks are here:… 12 r/LocalLLaMA community 26d ago Deepseek V4 Flash 2-bit quant is the first model I can run locally that achieves 100% in this SQL benchmark I really like to use this one SQL benchmark when testing new models. I had another post some time ago with my benchmarks, but I decided to post a new one because of how well Deepseek did. I like the benchmark because it's quick to run, is pretty "real-world" and requires good… 37 Hugging Face Daily Papers research 26d ago MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations Abstract Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across… 13 Hugging Face Daily Papers research 26d ago ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures Abstract Scientific figure comprehension and reasoning using multimodal AI requires integrating visual perception with domain-specific reasoning to extract meaningful knowledge, often not presented in the text of a research publication. The Sci-ImageMiner benchmark dataset,… 6 Hugging Face Daily Papers research 26d ago GEOID-Flood: A Large-Scale Multi-Modal Benchmark Dataset for Flood Segmentation Abstract Geospatial foundation models aim to learn representations that transfer across regions and sensors, yet evaluating them on specific tasks requires large, high-quality, multi-modal benchmarks that measure how well such models extract value from data. Concerning flood… 19 arXiv — Machine Learning research 26d ago Learning Compositional Meta-Routing for Agentic Workflows: An Executable Benchmark arXiv:2608.00106v1 Announce Type: new Abstract: Agentic systems must decide not only what answer to produce, but which reasoning and execution operations should precede it. A controller may answer directly, decompose a request, retrieve evidence, execute code, delegate to a… 17 arXiv — Machine Learning research 26d ago Rethinking Pretraining for Specialized Design Data: Evidence from the JONES-19 Cultural Design Dataset arXiv:2608.00135v1 Announce Type: new Abstract: Design and architectural archives encode expert human knowledge in graphical formats, providing a critical testbed for design-inspired Machine Learning (ML) challenges absent with typical computer vision benchmarks. Building on… 33 arXiv — Machine Learning research 26d ago UpliftBench: Revealing Outcome-Regime and Objective Mismatch in Uplift Evaluation arXiv:2608.00915v1 Announce Type: new Abstract: Uplift modeling (conditional-average-treatment-effect estimation) drives personalized targeting, yet published uplift benchmarks frequently disagree on which estimator performs best; we show the disagreement is substantially about… 33 arXiv — Machine Learning research 26d ago Measuring in-context algorithmic reasoning in language models against an exact Bayes-optimal standard arXiv:2608.01575v1 Announce Type: new Abstract: Whether large language models perform genuine algorithmic reasoning or mere pattern completion is hard to test, because most benchmarks lack a ground truth for correct inductive inference. We introduce F-ICL, an in-context-learning… 20 arXiv — NLP / Computation & Language research 26d ago AgentMemBench: A Systematic Benchmark for Evaluating Long-Term Memory Management Strategies in Conversational AI Agents arXiv:2608.00009v1 Announce Type: new Abstract: Long-term memory remains a critical bottleneck for conversational AI agents, whose finite context windows cannot support coherent recall across thousands of turns. We present AgentMemBench, a unified, reproducible benchmark… 32 arXiv — NLP / Computation & Language research 26d ago Obshazard-bench: Benchmarking Multimodal Foundation Models for Real-Time Disaster Intelligence from Raw Earth Observation Streams arXiv:2608.00012v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) are increasingly used to interpret Earth observation data, yet their capability to support real-world disaster emergency response remains insufficiently evaluated. Existing remote sensing… 26 arXiv — NLP / Computation & Language research 26d ago XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding arXiv:2608.00036v1 Announce Type: new Abstract: Real-world document tasks often ask professionals to answer questions from annual reports, regulations, clinical guidelines, and technical manuals that span hundreds or thousands of pages. Some questions also require comparing… 23 arXiv — NLP / Computation & Language research 26d ago CurveShift: Is Agent Progress Scalar? Separating Level from Shape arXiv:2608.00355v1 Announce Type: new Abstract: Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score. These summaries capture the overall performance, but they do… 6 arXiv — NLP / Computation & Language research 26d ago TreeProbe : A Tibetan Medicine Benchmark for Cultural Bias in LLMs arXiv:2608.00640v1 Announce Type: new Abstract: Large language models are increasingly viewed as a potential means of mitigating global health inequities, yet their outputs often reflect dominant high-resource medical traditions and provide limited coverage of traditional… 38 arXiv — NLP / Computation & Language research 26d ago ArabicDialectSafety: A Dialect-Aware Benchmark for Arabic Content Safety Classification arXiv:2608.01291v1 Announce Type: new Abstract: We present ArabicDialectSafety, a human-curated Arabic safety dataset of 25,071 prompts covering six Arabic varieties: Modern Standard Arabic, Syrian, Egyptian, Algerian, Palestinian, and Moroccan. The dataset is annotated with… 19 arXiv — NLP / Computation & Language research 26d ago CrossLex: A Source-Grounded Benchmark for Cross-Jurisdictional Legal Reasoning in Large Language Models arXiv:2608.01292v1 Announce Type: new Abstract: Legal reasoning is inherently jurisdiction-dependent: the same facts can call for different legal rules and yield different conclusions across legal systems. Yet existing benchmarks rarely evaluate whether large language models… 34 arXiv — NLP / Computation & Language research 26d ago LongChart VQA: A Comprehensive Benchmark for MLLMs with Complex Multi-Chart Reasoning arXiv:2608.01328v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are rapidly evolving with expanded context windows and stronger reasoning capabilities, enabling multi-chart understanding and multi-step inference. These abilities are increasingly… 10 arXiv — NLP / Computation & Language research 26d ago Prompt-Induced Waste in Large Reasoning Models: A Preregistered Two-Harness Benchmark of Coding Agents arXiv:2608.01347v1 Announce Type: new Abstract: Large reasoning models used as coding agents incur costs from deliberation, tool calls, and repeated agent turns, yet the causal effect of prompt wording on this spend has not been measured systematically. We present a… 36 arXiv — NLP / Computation & Language research 26d ago Semantic Alignment of AI Models: Concept Collapse, Checkpoint Dynamics, and Cross-Lingual Transfer arXiv:2608.01585v1 Announce Type: new Abstract: Language model benchmarking is a difficult task. Outcome reasoning alone does not test the model's conceptualization of language and popular open-source benchmarks are quickly saturated or ingested as training data. It is important… 7 Hugging Face Daily Papers research 26d ago WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity Abstract Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the apparent appearance of generated videos to the inherent reactivity of the worlds they depict: the ability to infer from… 6 Hugging Face Daily Papers research 26d ago SWE-Touch: Benchmarking Coding Agents When Users Touch the Code Abstract Real-world software development requires coding agents to operate in shared workspaces where users may inspect and modify code during an ongoing task, yet existing repository-level benchmarks typically evaluate agents working alone or restrict user participation to… 38 Hugging Face Daily Papers research 26d ago ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step Abstract To operate robustly in open-world environments, autonomous agents should be able to infer the behavior of unfamiliar systems through interaction alone, even in the absence of documentation. However, existing tool-use benchmarks expose semantic tool schemas in static… 24 r/MachineLearning community 26d ago I created an autonomous boxing benchmark [D] I created an AI boxing match to test the decision speed, adaptability and strategy. I fed the LLMs with data about the current match and if they have vision, they will get even more data. The match has street rules, anything goes and an AI is not defeated until the ref counts to… 24 r/LocalLLaMA community 26d ago 'I ran my own benchmarks on it' seems to be pretty common comment around here. How about dedicating a thread for this and sharing? Of course, the concern is that in the end, this thread will be fed into the models' training data, but I feel benchmarking isn't so open and very fragmented.   submitted by   /u/jinnyjuice [link]   [comments] 38 r/LocalLLaMA community 27d ago Qwen3.8-Max matches Kimi K3 and DeepSeek V4 Flash Qwen3.8-Max (2.4T) is another massive contribution to the open weight community. On benchmarks, it performs closely to Kimi K3 and DeepSeek V4 flash across all categories and is better at coding and software tasks. Qwen3.8-27B will also be open weight soon too. Weights are being… 11 NVIDIA Developer Blog official-blog 27d ago NVIDIA Vera Storage Benchmarks: Faster Encryption, Compression, Integrity Checking, and Recovery for AI-Native Storage Storage is an active part of every agentic AI workflow. As agents retrieve enterprise knowledge, access persistent memory, reuse key-value (KV) cache data,... 17 r/LocalLLaMA community 27d ago "Data center in a Box (on Wheels)" 256Gb VRAM/512Gb RAM AI Server 6-8 Month Operational Review, Stability Write Up, Benchmarks I've been out of these forums for awhile but I figured I would provide a formal update on how this has been going now that it has some operation time under its belt, just to put the information out there and share knowledge if there is any interest. I also wasn't satisfied with… 32 r/LocalLLaMA community 27d ago V4-Flash-0731 - vibes after first weekend of use Spent way too much time with V4-Flash-0731 this weekend and wanted to share my vibes as briefly as possible. I sent it through a bit of real-work and some of my personal benchmarks. My quick thoughts are: Quantization hits this thing like a truck - I've tried a bunch of the Q2… 14 r/LocalLLaMA community 27d ago [RELEASE] SupraBrain-50M-v0.1 Hey there! So today we're releasing SupraBrain-50M, a hybrid language model that combines Gated DeltaNet linear recurrence with Sliding-Window Attention and Surprise-Gated update mechanisms to deliver very strong performance. Here are the benchmarks:… 19 Hugging Face Daily Papers research 27d ago Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants Abstract AI-assisted coding increasingly translates informal user intent into executable software, yet coding requests often contain ambiguities that recur in user-specific ways across tasks and sessions. Existing disambiguation methods typically address each ambiguous request… 15 Hugging Face Daily Papers research 27d ago ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction Abstract Enterprise workflows increasingly rely on agents for schema-guided extraction: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with source evidence as grounding metadata. We present ExtractBench, a… 4 r/LocalLLaMA community 27d ago I benchmarked classic vector RAG vs Google's new OKF format vs both combined — same corpus, same 7 questions, all local (Ollama + ChromaDB) Google Cloud published OKF (Open Knowledge Format) on June 12th — a spec for storing curated knowledge as a directory of markdown files with YAML frontmatter. One concept per file, linked to each other, with an index.md for progressive disclosure. The only required field is… 7 Hugging Face Daily Papers research 27d ago SULAND v2: A Refined RGB Dataset and Deep Learning Object Detection Benchmark for UAV/UGV-Based SUrface LANDmine Detection Under Domain Shift Abstract RGB imagery offers a practical, low-cost option for Unmanned Aerial/Ground Vehicle (UAV/UGV) survey support in surface-landmine detection, but object detectors remain underexplored in this safety-critical domain. Limited cross-architecture benchmarking and insufficient… 32 Page 10 of 10 · 500 articles ← Newer