News / #benchmark Tag Benchmark 500 articles archived under #benchmark · RSS Sign in to follow arXiv — NLP / Computation & Language research 11d ago Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See arXiv:2608.17744v1 Announce Type: new Abstract: Take three frontier mixture-of-experts models (Alibaba, OpenAI, NVIDIA; 3.6-4.0B active parameters each) and fine-tune them to reason in a low-resource language. On accuracy benchmarks almost nothing happens, and the benchmark… 17 arXiv — NLP / Computation & Language research 11d ago From Global Benchmarks to Local Evaluations: Benchmarking LLMs for the German Public Sector arXiv:2608.17827v1 Announce Type: new Abstract: Public institutions face a persistent challenge in selecting LLMs suited to their specific context. Existing benchmarks, however, are of limited use as they primarily reflect English-language and US-centric settings, and often only… 9 arXiv — NLP / Computation & Language research 11d ago BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models arXiv:2608.17895v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) have made significant strides in visual comprehension, their ability to reason about text-dense, professional documents remains incompletely evaluated. Existing benchmarks emphasize… 11 arXiv — NLP / Computation & Language research 11d ago When Writing Style Drifts: Benchmarking Authorship Verification under Distribution Shifts in Genre, Time and the AI-Era arXiv:2608.17979v1 Announce Type: new Abstract: Authorship verification (AV) assumes that an author's writing style remains sufficiently stable to distinguish it from that of other writers. In practice, however, this assumption is challenged by distribution shifts caused by… 30 arXiv — NLP / Computation & Language research 11d ago What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations arXiv:2608.17719v1 Announce Type: cross Abstract: Context: Software systems that depend on commercial large language model APIs must migrate to successor versions when vendors deprecate older models. Migration decisions typically rely on aggregate benchmark scores, which… 25 Hugging Face Daily Papers research 11d ago Energy-Guided Flow Matching Abstract Energy-Guided Flow Matching improves generative quality by progressively revealing high-frequency details through a moving endpoint and adaptive scheduling, reducing training cost and achieving state-of-the-art FID scores. Generated by thinkingmachines/Inkling-Small… 7 r/LocalLLaMA community 11d ago GLM5.3 Artificial Analysis Benchmarks   submitted by   /u/anderspitman [link]   [comments] 33 r/LocalLLaMA community 11d ago Ling-3.0 (BailingMoE3) lands in llama.cpp mainline - Quick benchmarks on Intel Arc B580 Finally llama.cpp now officially supports Ling-3.0! (Starting from build b10472 +) If you want to run them locally, bartowski has already released the GGUF imatrix quantizations for both models: - Ling-3.0-tiny (8B) - Ling-3.0-flash (127B) After quite a while, PR #26608 has… 33 r/LocalLLaMA community 11d ago Is Ling 3 tiny underrated for its size? I was checking out benchmarks of this model and apparantly better than Qwen3.5 9b reasoning across the bench on artificial analysis. I have used the 9b model for variety of stuff and it has been amazing, but if this is better then why not switch. I am downloading it rn to test… 6 Hugging Face Daily Papers research 12d ago Advancing Open and Reproducible Relational Learning: RelArena-α, TabPFN-Rel and RPI Abstract Prior Labs released open-source tools including a unified relational benchmark framework, a TabPFN-based relational model, and a model-agnostic predictive interface to advance reproducible relational learning. Generated by thinkingmachines/Inkling-Small This first… 21 Hugging Face Daily Papers research 12d ago How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks Abstract Autonomous research agents evaluated across the full scientific lifecycle reveal a pervasive lack of metacognitive self-correction, motivating a new benchmark and failure taxonomy. Generated by thinkingmachines/Inkling-Small AI has long assisted scientific research, but… 8 r/LocalLLaMA community 12d ago AA is the reason for Qwen3.8 27B shipped with xhigh I know why Qwen3.8 27B shipped with xhigh reasoning as default, it's to do its best in benchmarks. Models from top labs often get benchmarked at multiple reasoning levels, but that same treatment doesn't apply to other labs. Open models are lucky to even be benchmarked at all.… 16 arXiv — Machine Learning research 12d ago WANDR: A Benchmark for Wide and Deep Research arXiv:2608.14747v1 Announce Type: new Abstract: WANDR (Wide ANd Deep Research) is a benchmark of 500 realistic, challenging data-collection tasks for research agents. Each task requires a system to discover a large set of entities that satisfy specified criteria (breadth),… 19 arXiv — Machine Learning research 12d ago ER-KANs: Efficient and Robust Kolmogorov-Arnold Networks for Data-Scarce Scientific Machine Learning arXiv:2608.14773v1 Announce Type: new Abstract: The efficient-KAN literature---covering Chebyshev, wavelet, and radial-basis-function variants of the original Kolmogorov-Arnold Network---has been benchmarked almost entirely on clean data. We show that this choice conceals a… 32 arXiv — Machine Learning research 12d ago Certifying Compressed Language Models: An Audit and a Statistical Toolkit arXiv:2608.15046v1 Announce Type: new Abstract: A fraction of a point of benchmark accuracy is the usual evidence that a compressed model is equivalent to its original. That quantity is least informative when two models are most alike: a net delta is what survives cancellation… 16 arXiv — Machine Learning research 12d ago FinFraudBench: A Heterogeneous Graph Benchmark for Financial Fraud Detection arXiv:2608.15177v1 Announce Type: new Abstract: The increasing complexity of digital financial systems has reshaped financial fraud detection from isolated transaction classification into relational risk reasoning over interconnected financial entities. This shift has motivated… 17 arXiv — Machine Learning research 12d ago Learning reshapes power-law anisotropy in internal representations arXiv:2608.15239v1 Announce Type: new Abstract: Power-law anisotropy in internal representations has been observed across a wide range of biological and artificial neural systems, from state-of-the-art language models to the mouse cerebral cortex. This anisotropy is a key… 35 arXiv — NLP / Computation & Language research 12d ago Auxiliary uncertainty signals for LLM-assisted systematic review screening: a benchmark across eight Cohen drug-class reviews arXiv:2608.14551v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for title-abstract screening in systematic reviews, but their decisions lack calibrated uncertainty. We show that an auxiliary BERT+GCN classifier supplies a structured uncertainty… 12 arXiv — NLP / Computation & Language research 12d ago Interpretable Cross-Lingual Alignment in Small Language Models: Probing Cultural and Pragmatic Reasoning in Japanese-English Bilingual LLMs arXiv:2608.14896v1 Announce Type: new Abstract: Large language models work well on English and behave in poorly understood ways on languages typologically far from it. Japanese is a clean example, where evaluation still leans on translation quality and JGLUE-style benchmarks,… 21 arXiv — NLP / Computation & Language research 12d ago Gated Against One Model, Open to the Next: Option-Only Solvability in Legal Multiple-Choice Benchmarks arXiv:2608.15428v1 Announce Type: new Abstract: Multiple-choice benchmarks are graded on whether a model picks the right option, not on whether it needed the question. Measuring that gap takes care: a model answering A to most items scores above chance wherever the key sits at… 25 arXiv — NLP / Computation & Language research 12d ago L3Cube-IndicQuest v2: A Large-Scale Multilingual Benchmark for Evaluating Factual Knowledge of Large Language Models Across Indic Languages arXiv:2608.15535v1 Announce Type: new Abstract: We present L3Cube-IndicQuest v2, a large-scale gold-standard multilingual question-answering benchmark for evaluating the India-specific factual knowledge of Large Language Models (LLMs). The benchmark comprises 3,471… 16 arXiv — NLP / Computation & Language research 12d ago When Stories Evolve: Benchmarking LLM Storytelling Across Agent Architectures in Open-Ended World Simulations arXiv:2608.15654v1 Announce Type: new Abstract: Large language models can write fluent stories, but open-ended storytelling requires more than local fluency. In evolving world simulations and AI-native games, models must preserve facts, relationships, causal dependencies, and… 21 arXiv — NLP / Computation & Language research 12d ago PLSQLBench: Benchmarking LLM Systems for Executable Procedural Database Programming arXiv:2608.15931v1 Announce Type: new Abstract: We present PLSQLBench, to our knowledge the first benchmark for evaluating whether LLMs can write executable PL/SQL programs, with correctness measured through execution-based tests. Existing LLM evaluations largely target… 26 arXiv — NLP / Computation & Language research 12d ago Whose Gold? Annotator-Pool Disagreement Is Large at the Item Level, and Hidden by Small Leaderboards arXiv:2608.15980v1 Announce Type: new Abstract: Preference benchmarks are built by hiring annotators, and the identity of those annotators is treated as an implementation detail. We measure what that detail buys. On the 2,885 MultiPref items where both pools are internally… 37 arXiv — NLP / Computation & Language research 12d ago ReRef-3D: A Benchmark for Spatial Referring Expression-Guided 3D Scene Rearrangement arXiv:2608.16011v1 Announce Type: new Abstract: We introduce ReRef-3D, a benchmark for language-guided placement in 3D scenes. It contains 33,826 instructions across 998 CLEVR-derived scenes, spanning 16 placement families and direct, one-hop, and two-hop references. Each… 18 arXiv — NLP / Computation & Language research 12d ago $R^3$-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets arXiv:2608.16033v1 Announce Type: new Abstract: In cognitive science, resource rationality asks how an agent should allocate limited computation to maximize expected value. Most reasoning and agent benchmarks use independent per-task budgets; existing shared-budget studies do… 12 arXiv — NLP / Computation & Language research 12d ago IndicQE-APE: A Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages arXiv:2608.16344v1 Announce Type: new Abstract: Indic quality estimation (QE) and automatic post-editing (APE) data is spread across separate releases, so no single resource supports training and evaluation across tasks and language pairs on one footing. We consolidate the WMT… 18 arXiv — NLP / Computation & Language research 12d ago Toward Better Assessment of LLMs' Performance in Clinical Error Detection arXiv:2608.16643v1 Announce Type: new Abstract: Automated detection of errors in clinical documentation is a promising application of large language models (LLMs), yet decisions to deploy such models rest on benchmarks that evaluate each clinical note in isolation.… 7 Hugging Face Daily Papers research 12d ago HarnessEval-W: Agentifying the Evaluation of Visual Worlds Abstract HarnessEval-W uses hierarchical sub-agents to decompose world-model evaluations into verifiable reasoning chains that justify scores with transparent evidence. Generated by thinkingmachines/Inkling-Small A benchmark should deliver more than a scalar score: what makes an… 29 Hugging Face Daily Papers research 12d ago PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments Abstract PACE-Bench evaluates self-evolving agents on physics adaptation tasks requiring iterative code redesign after environmental mutations, revealing that simulator-grounded reflection outperforms unverified self-revision but mechanism redesign remains a major bottleneck.… 30 Hugging Face Daily Papers research 12d ago VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding Abstract VideoGAIA introduces a multi-turn, tool-augmented benchmark that evaluates agentic video understanding for advanced multimodal models through complex real-world tasks. Generated by thinkingmachines/Inkling-Small Video understanding is a fundamental task for evaluating… 4 Hugging Face Daily Papers research 12d ago UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations Abstract UI-Mate is a foundation GUI agent that uses environment-grounded training and in-context demonstration learning to improve reliability on long-horizon office tasks, achieving state-of-the-art results on computer-use benchmarks. Generated by… 30 Hugging Face Daily Papers research 12d ago ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering Abstract ENTLORE is a benchmark framework that evaluates enterprise question answering by requiring recovery of implicit organizational relations across routine documents, revealing that even with gold sources many latent reasoning questions remain unanswered. Generated by… 7 r/LocalLLaMA community 12d ago Optimizing Qwen3.6 / Qwen3.8-27B on 16GB VRAM: Complete Benchmark Results and Setup Guide (~30-50tps at 32k to 72k context) This post was made with AI. I tried to remove as much slop as possible and keep it straight to the point to save your time as I know how annoying AI slop posts can be, but I still wanted to retain all the details so it can be used as a resource for comparison with other future… 38 Hugging Face Daily Papers research 12d ago VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End? Abstract A unified framework benchmarks and trains multimodal agents that infer intent, plan 3D scenes, invoke tools, and reflect on feedback, revealing that reinforcement learning improves open-source models beyond closed-source frontiers. Generated by… 14 r/MachineLearning community 12d ago We’ve got a workshop on production retrieval-augmented generation with open models, benchmarked end to end, thought it’d be relevant here [D] There’s a hands-on workshop on August 29 that builds and benchmarks this properly, end to end, using entirely open models, no API calls involved. Led by Ben Auffarth, AI Consultant and Founder of Chelsea AI Ventures. What it covers: • Hybrid retrieval (vector + keyword, not… 6 r/LocalLLaMA community 12d ago we benchmark models nobody actually runs qwen3.8-27b looks genuinely impressive on the benchmark tables - beating models many times its size on some of them. but those numbers come from bf16 weights, and nobody here is running a 27b at bf16. we're running the 4-bit at ~17gb because that's what fits on a 4090 or a 24gb… 16 r/LocalLLaMA community 12d ago Local agentic coding Benchmark : Qwen 3.8 27B (in many weights quants / cache quants / engine / reasoning effort) vs others. In medium reasoning mode, it both scores higher than the 3.6 version, AND is very much more efficient (almost half requests needed, and a third less tokens generated) - at DeepSeek v4 Flash 3107 MXFP4 level The xhigh mode is advertised to be the best one for hard tasks. In this… 32 r/LocalLLaMA community 12d ago Artificial Analysis' Qwen3.8-27B benchmarks put it neck and neck with DeepSeek V4 and GPT-5.6 Luna Max   submitted by   /u/anderspitman [link]   [comments] 30 r/LocalLLaMA community 13d ago tencent/EVIE-Preview-4.5B · Hugging Face Overview EVIE-Preview-4.5B is a state-of-the-art multilingual Visual Document Retrieval (VDR) model built upon Qwen3.5-4B . It employs ColBERT-style late interaction with native 128-dimensional multi-vector token embeddings (4.54B parameters, BF16). By combining native… 37 Hugging Face Daily Papers research 13d ago DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data Abstract Mimir v1 is a 1-billion-parameter Hierarchical Reasoning Model trained solely on permissible data that achieves competitive English results and state-of-the-art Danish performance across multiple benchmarks. Generated by thinkingmachines/Inkling-Small Current large… 20 Hugging Face Daily Papers research 13d ago Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems Abstract Multi-agent clinical committees are vulnerable to socially plausible shortcuts rather than isolated cues, and only independent referee oversight reliably detects adoption. Generated by thinkingmachines/Inkling-Small Clinical decision support is moving toward committees… 35 Hugging Face Daily Papers research 13d ago A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images Abstract The ALD/E-ImageMiner benchmark and ICDAR 2026 competition advance machine interpretation of scientific figures through tasks spanning visual reading, domain reasoning, and evidential justification, proposing long-term goals for verifiable multimodal scientific AI.… 30 Hugging Face Daily Papers research 13d ago SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning Abstract On-policy distillation from a long-context reasoning teacher to short-context students improves mathematical proof reasoning and generalizes to science benchmarks by aligning token spans, constraining length growth, and stabilizing training. Generated by… 26 Hugging Face Daily Papers research 13d ago MobileMem: Learning from a Year of Mobile Experiences Abstract MobileMem is a benchmark and framework for evaluating on-device long-term memory through year-scale, multimodal mobile experience trajectories that require temporal reasoning, knowledge updating, and preference inference. Generated by thinkingmachines/Inkling-Small The… 5 arXiv — Machine Learning research 13d ago Don't Claim Benchmark-Oriented Optimization Improves General Coding Capability -- Diverse Evaluation Is Required arXiv:2608.13566v1 Announce Type: new Abstract: Post-training papers, model cards, and blog posts often treat scores on a small set of coding benchmarks (e.g., SWE-bench and LiveCodeBench) as evidence of broad coding capability, both for research artifacts and user-facing… 22 arXiv — Machine Learning research 13d ago PPAPlace: Differentiable Cross-Stage Objectives for Chip Placement Optimization arXiv:2608.13790v1 Announce Type: new Abstract: Macro placement significantly affects a chip's post-route performance, power, and area (PPA). Most placement methods optimize half-perimeter wirelength (HPWL) as the primary objective. However, recent benchmarking shows a near-zero… 13 arXiv — Machine Learning research 13d ago Dynamic Multi-Depot Vehicle Routing with Online Requests: Event-Driven Transformer--DRL and Rolling-Horizon Benchmarking arXiv:2608.13799v1 Announce Type: new Abstract: This paper presents an event-driven learning and benchmarking framework for the Dynamic Multi-Depot Vehicle Routing Problem with progressively revealed requests and evolving vehicle states. Masked MLP and Transformer policies are… 21 arXiv — Machine Learning research 13d ago Generating Benchmark Health Data Using a Tabular Diffusion Transformer arXiv:2608.14496v1 Announce Type: new Abstract: Cross-Tabular Data Generation (CTDG) seeks to learn a generative model from multiple heterogeneous tables and produce new synthetic tabular datasets. However, existing synthetic tabular data generation methods are largely… 38 arXiv — Machine Learning research 13d ago Language-Specific Gaps in AI Safety Training Datasets arXiv:2608.13695v1 Announce Type: cross Abstract: Large language model providers routinely cite multilingual safety benchmarks spanning a dozen or more languages as evidence that their models are safe for non-English-speaking users. We show that these collection-level coverage… 32 Page 5 of 10 · 500 articles ← Newer Older →