News / #benchmark Tag Benchmark 500 articles archived under #benchmark · RSS Sign in to follow arXiv — NLP / Computation & Language research 17d ago Benchmarking Trustworthiness of SLMs: Pre-trained vs. Compressed arXiv:2608.11981v1 Announce Type: new Abstract: Small Language Models (SLMs) have emerged as a more efficient alternative to traditional Large Language Models (LLMs), offering promising potential in resource-constrained scenarios. Existing approaches to building SLMs typically… 34 arXiv — NLP / Computation & Language research 17d ago A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench arXiv:2608.12138v1 Announce Type: new Abstract: General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed… 10 arXiv — NLP / Computation & Language research 17d ago When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs arXiv:2608.11403v1 Announce Type: cross Abstract: Self-consistency (SC) via majority vote is a widely used way to spend inference-time compute: sample N chains of thought, return the plurality answer. On the full GPQA Diamond benchmark (198 graduate-level science questions),… 30 arXiv — NLP / Computation & Language research 17d ago Benchmarking LLM Judges for Mobile Agent Evaluation arXiv:2608.11434v1 Announce Type: cross Abstract: Mobile agent benchmarks increasingly rely on LLM-based judges to evaluate task completion, yet the reliability of these judges on mobile agent trajectories remains largely unexamined. We introduce MobileJudgeBench, a benchmark… 17 arXiv — NLP / Computation & Language research 17d ago FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents arXiv:2608.11683v1 Announce Type: cross Abstract: AI agents are increasingly deployed for professional investment research, yet no benchmark captures the complexity of the full investor workflow. Existing benchmarks mainly target financial data extraction, a narrow slice that… 25 arXiv — NLP / Computation & Language research 17d ago VICBench: A Multi-Language Benchmark for Code Vulnerability Detection arXiv:2608.12246v1 Announce Type: cross Abstract: Evaluating security vulnerability detection tools requires benchmark datasets with vulnerability-inducing commits (VICs) - the commits that first introduce vulnerabilities into codebases. VICs are essential for determining the… 7 Hugging Face Daily Papers research 17d ago From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection Abstract A closed-loop framework combining physics-based video synthesis, diffusion-based video dereflection, and a new benchmark achieves state-of-the-art video reflection removal with fast inference. Generated by thinkingmachines/Inkling-Small Videos captured through glass… 29 Hugging Face Daily Papers research 17d ago MBA: Multimodal Benchmark and Agents for Real-World Business Ideation Abstract Researchers introduce MBA-Bench, a multimodal benchmark for business ideation agents, and propose MBA-b and MBA-k models trained with creativity and feasibility rewards via LoRA fine-tuning and group relative policy optimization, significantly outperforming text-only… 21 r/LocalLLaMA community 17d ago I ran DeepSeek V4 Flash 284B + DSpark on one RTX PRO 6000. The drafter was faster in RAM than VRAM. Hey guys, Just finished benchmarking DeepSeek V4 Flash 284B + DSpark on a single RTX PRO 6000 96GB . Short version: DSpark: ~15–17% faster generation on my coding workload On this setup, the DSpark drafter was faster in system RAM than VRAM q8_0 KV cache: 256K → 768K context… 18 r/LocalLLaMA community 17d ago LFM2.5-VL-3B recognizes Steve from Minecraft running locally on an iPhone 17 Liquid AI put out LFM2.5-VL-3B today, which is a 3.1B vision model that weighs roughly 2GB and fits well on a phone Benchmarks are benchmarks so I tried something sillier. Took a photo of a little Steve toy I have, gave it to the model and asked it what it was looking at It… 29 Hacker News — AI on Front Page community 17d ago Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index Article URL: https://artificialanalysis.ai/articles/grok-4-6-benchmarks-and-analysis Comments URL: https://news.ycombinator.com/item?id=49275385 Points: 236 # Comments: 229 23 r/LocalLLaMA community 17d ago DeepSeek V4-Pro-0813 Benchmarks   submitted by   /u/MagicZhang [link]   [comments] 14 r/LocalLLaMA community 18d ago Gemma 4 QAT handles KV cache quantization MUCH better, KLD benchmarks show Link to the article: KV Cache Quantization on Gemma 4 31B: Non-QAT vs QAT KLD benchmarks with BeeLlama.cpp v0.4.3 , fork of llama.cpp with more KV cache quantization options, comparing Gemma Q4_0 non-QAT vs Gemma Q4_0 QAT. Long story short: QAT is much more friendly to KV cache… 4 Hugging Face Daily Papers research 18d ago 360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents Abstract A new photorealistic urban benchmark reveals large performance gaps for embodied agents in city-scale navigation and spatial reasoning. Generated by thinkingmachines/Inkling-Small We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of… 8 Hugging Face Daily Papers research 18d ago Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Abstract The study introduces a benchmark and formalizes narrative commitment preservation to evaluate long-horizon logical consistency in interactive storytelling with large language models. Generated by thinkingmachines/Inkling-Small The rapid advancement of Large Language… 16 arXiv — Machine Learning research 18d ago Uncertainty-Aware Ensemble Deep Randomized Neural Networks for Classification arXiv:2608.10007v1 Announce Type: new Abstract: The current state-of-the-art (SOTA) deep randomized neural networks, such as deep Random Vector Functional Link (dRVFL) and ensemble deep RVFL (edRVFL), treat all training samples uniformly, which limits their robustness and… 30 arXiv — Machine Learning research 18d ago UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs arXiv:2608.10042v1 Announce Type: new Abstract: Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool use, or response-level personalization. We introduce UserToolBench , a benchmark… 29 arXiv — Machine Learning research 18d ago Toward Human Rights Benchmarking for LLMs: A Pilot Methodology arXiv:2608.10268v1 Announce Type: new Abstract: Large language models (LLMs) increasingly mediate legal determinations over what human rights are realized, and how. Yet, no evaluation benchmark exists to assess whether they can reason correctly about human rights law. To this… 14 arXiv — Machine Learning research 18d ago Benchmarking Time Series Generation Methods for Privacy-Preserving Forecasting arXiv:2608.10891v1 Announce Type: new Abstract: Time series forecasting in privacy-sensitive domains often requires training models on released data rather than original observations. Synthetic time series generation has been developed primarily for data augmentation, where… 6 arXiv — Machine Learning research 18d ago Derivative Computation in PINNs: Automatic Differentiation, Finite Differences and Beyond arXiv:2608.11020v1 Announce Type: new Abstract: We systematically investigate finite-difference (FD) derivative computation in Physics-Informed Neural Networks (PINNs) as an alternative to automatic differentiation (AD). On three benchmark PDEs we show that, with a properly… 31 arXiv — NLP / Computation & Language research 18d ago Mapping and Measuring the Behavioral Evolution of Large Language Models arXiv:2608.11027v1 Announce Type: cross Abstract: Benchmark leaderboards summarize how well a language model performs, but not how its behavior relates to that of other models or changes across generations. We characterize the output behavior of 32 models from six families using… 27 arXiv — Machine Learning research 18d ago Cross-View Feature Matching: Survey, Benchmarking, and Foundation-Model Perspectives arXiv:2608.11093v1 Announce Type: new Abstract: Cross-view feature matching aims to establish reliable correspondences across images with large viewpoint variations. Over the past decade, the field has evolved from task-specific models toward increasingly unified and… 35 arXiv — Machine Learning research 18d ago DACRI: Decision-Aware Causal Intervention Ranking for Critical Supply Chains arXiv:2608.11154v1 Announce Type: new Abstract: Detecting or attributing a supply-chain disruption is not the same as selecting the intervention that maximizes recoverable net value. We present CriticalSCM-Bench v1, a controlled synthetic benchmark with causal ground truth,… 35 arXiv — Machine Learning research 18d ago HyperShape: Hyperelasticity Across Diverse Shapes arXiv:2608.09938v1 Announce Type: cross Abstract: Hyperelastic deformations are highly sensitive to domain geometry and boundary conditions, making generalization across both a critical capability for neural operators applied to these problems. However, existing benchmarks for… 34 arXiv — Machine Learning research 18d ago Energy and Performance Benchmarking of Deep Learning Models for Breast Cancer Detection arXiv:2608.09996v1 Announce Type: cross Abstract: Recent advances in machine learning have greatly improved breast cancer detection, enabling more accurate and timely diagnosis. Deep learning (DL) models show strong potential for medical image analysis; however, as their… 25 arXiv — NLP / Computation & Language research 18d ago TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent arXiv:2608.10258v1 Announce Type: new Abstract: Large language models (LLMs) increasingly provide conversational health information that may influence treatment decisions, yet existing benchmarks do not isolate whether medication-safety boundaries persist across follow-ups after… 29 arXiv — NLP / Computation & Language research 18d ago Locally Deployable Small Language Models for Emergency Department Decision Support: A Systematic Benchmark of Fine-Tuning Strategies arXiv:2608.10273v1 Announce Type: new Abstract: Deploying large language models (LLMs) for decision support in emergency departments (EDs) faces two major challenges: privacy risks of transmitting patient data to closed-source commercial LLMs and the lack of systematic… 4 arXiv — NLP / Computation & Language research 18d ago Every Token Counts: Exact Likert-Scale Distributions for Measuring LLM Attitudes and Biases arXiv:2608.10503v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed as autonomous agents, accurately evaluating their latent values and biases is critical. The NLP community typically evaluates models using large, unstructured benchmarks.… 21 arXiv — NLP / Computation & Language research 18d ago Surfacing the Unsaid: CUE-Bench for Affective Stance in Chinese Discourse arXiv:2608.10810v1 Announce Type: new Abstract: Emotion understanding in discourse requires reasoning beyond surface sentiment because speakers often convey affect through indirect, implicit, polite, ironic, or deliberately mismatched expressions. Existing emotion benchmarks… 5 arXiv — NLP / Computation & Language research 18d ago FaithformBench: Benchmarking Faithfulness of Mathematical Chain-of-Thought Autoformalisation arXiv:2608.10916v1 Announce Type: new Abstract: Autoformalisation (AF) systems map natural language reasoning steps into formal statements in a proof assistant such as Lean. We consider how to assess the faithfulness of these systems. Existing approaches require expensive… 37 arXiv — NLP / Computation & Language research 18d ago ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering arXiv:2608.10679v1 Announce Type: cross Abstract: Enterprise question answering is framed as retrieving internal documents and generating grounded answers. Routine enterprise records, however, are work by-products in which required organizational relations remain implicit across… 7 arXiv — NLP / Computation & Language research 18d ago HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models arXiv:2506.03922v4 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have demonstrated significant potential to advance a broad range of domains. However, current benchmarks for evaluating MLLMs primarily emphasize general knowledge and vertical… 15 Hugging Face Daily Papers research 18d ago JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles Abstract A new jigsaw benchmark with interlocking pieces reveals that vision-language models fail at geometric reasoning and suffer a sharp performance drop as puzzle size increases. Generated by thinkingmachines/Inkling-Small Jigsaw puzzle solving requires jointly reasoning… 7 Hugging Face Daily Papers research 18d ago SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Abstract SPIEval benchmarks mobile assistant LLMs on scattered personal data tasks, revealing major gaps in information retrieval and verification. Generated by thinkingmachines/Inkling-Small Large language models (LLMs) are increasingly deployed as mobile assistants, where a… 36 r/LocalLLaMA community 18d ago New Muse-Glimmer-30B SoTA Quants - hopefully a new lineup :) Hey Folks, I've been making quants for a while - recently I took a short break to get into hardcore research (submitted my first EMNLP paper during it!). Along the way, I built up a little arsenal of quant-optim techniques: everything from novel, paper-pending tricks to some… 19 Hugging Face Daily Papers research 18d ago VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World? Abstract A new benchmark called VibeLifeBench evaluates long-horizon proactive agents across simulated multi-week everyday tasks, revealing that current frontier models perform poorly. Generated by thinkingmachines/Inkling-Small Large language model (LLM) agents are increasingly… 22 r/LocalLLaMA community 18d ago We quantized DeepSeek V4 0731 and benchmarked it against popular quants on 8× RTX 5090 We converted the model from the original safetensors and found two issues. The first one made our quantization fail several times, the second one does not fail at all, it just quietly ruins the base 1) You must use the --no-lazy option, otherwise token_embd.weight will take on… 23 r/LocalLLaMA community 18d ago Local Benchmark : Muse Glimmer 30B vs Qwen 3.6 27B vs Gemma4 31B (and many other models and finetunes) Needs a lot of requests compared to Qwen (almost twice) and Gemma (almost x3). Final score is fine, even though it is "not a coding model" https://wonderrico.github.io/local_llm_benchmark/benchmark-main.html more details on… 31 Hugging Face Daily Papers research 18d ago Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness Abstract Researchers propose source-contrastive evaluation via a localized benchmark to detect data contamination and assess localization robustness in multilingual translation models. Generated by thinkingmachines/Inkling-Small Multilingual translation benchmarks are typically… 4 Hugging Face Daily Papers research 18d ago Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure Abstract Optimized GPU kernel benchmarks reveal that evolutionary LLM proposals exploit evaluation configurations, causing widespread failure to generalize to held-out settings. Generated by thinkingmachines/Inkling-Small Benchmarks for systems that are optimized against the… 11 r/LocalLLaMA community 18d ago DeepSeek V4 Flash 0731 at 27+ t/s decode on Strix Halo — Vulkan + DSpark full guide Been benchmarking DSv4 Flash 0731 on a Flow Z13 (Ryzen AI MAX+ 395, Radeon 8060S / gfx1151, 128GB LPDDR5X) for the past week. Figured I'd share what actually works and what doesn't — there are a lot of gotchas on this hardware. Results Best client-side observation (bench-kv.sh… 7 Hugging Face Daily Papers research 19d ago MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models Abstract MMOOC is a large-scale benchmark assessing whether multimodal language models can correctly refuse out-of-context questions while answering shifted in-context questions, revealing that current models struggle to balance these abilities. Generated by… 37 Hugging Face Daily Papers research 19d ago BDH-CQ: In-Context Learning with Recurrent Latent Reasoning Abstract A 150M-parameter reasoning model using recurrent latent reasoning and in-context learning achieves a new cost-accuracy frontier on ARC-AGI-1. Generated by thinkingmachines/Inkling-Small We introduce BDH-CQ, a reasoning model that combines in-context learning with… 20 Hugging Face Daily Papers research 19d ago WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks Abstract WeClawArena is an auditable benchmark and sandbox for evaluating multi-party agent collaboration across personal workspaces, measuring both task utility and security attack success. Generated by thinkingmachines/Inkling-Small Recent advances in persistent personal-agent… 10 r/LocalLLaMA community 19d ago Luth-2: New State-of-the-Art French Small Language Models Hey everyone, Today we release Luth-2-0.8B and Luth2-2-2B , two non-reasoning models that set a new state of the art for French across a wide variety of tasks for their size 🚀 A few notable scores on French benchmarks compared to models 〜3 times their size: - Luth-2-2B scores… 31 arXiv — Machine Learning research 19d ago From Benchmark Performance to Tool Deployment: Human-in-the-Loop Anomaly Detection arXiv:2608.07770v1 Announce Type: new Abstract: Automated anomaly detection methods often report strong performance on curated academic benchmarks, but their behavior under real-world industrial conditions is less clear. In this work, we evaluate 19 unsupervised anomaly… 7 arXiv — Machine Learning research 19d ago Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training arXiv:2608.08224v1 Announce Type: new Abstract: Reinforcement learning post-training unlocks complex reasoning in LLMs. Yet benchmark scores reveal only whether a model improved, not what changed inside it, nor how it splits finite capability across tasks. A representative… 32 arXiv — Machine Learning research 19d ago When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs arXiv:2608.08542v1 Announce Type: new Abstract: Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, code, or domain specialists into a safety-aligned base using task arithmetic,… 12 arXiv — Machine Learning research 19d ago Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure arXiv:2608.08722v1 Announce Type: new Abstract: Benchmarks for systems that are optimized against the evaluation signal measure something different from what they claim. We document this concretely in two GPU-kernel-optimization suites with held-out generalization gates:… 24 arXiv — NLP / Computation & Language research 19d ago Unified Hallucination Fuzzing for Multimodal Large Language Models arXiv:2608.07525v1 Announce Type: new Abstract: Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high-stakes applications. Existing evaluations, predominantly based on static benchmarks, suffer from… 16 Page 7 of 10 · 500 articles ← Newer Older →