News / #benchmark Tag Benchmark 500 articles archived under #benchmark · RSS Sign in to follow arXiv — NLP / Computation & Language research 13d ago AnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMs arXiv:2608.14320v1 Announce Type: cross Abstract: The anchoring effect is a cognitive bias in which an initial reference value shifts a later judgment toward itself. This effect is well established in human judgment and decision-making, and recent work suggests that large… 27 arXiv — NLP / Computation & Language research 13d ago A Four-Axis Trustworthiness Benchmark for LLM-as-Judge in Principle-Based Regulation arXiv:2608.14329v1 Announce Type: cross Abstract: Principle-based regulation, with evaluative standards such as "fair, clear, and not misleading" or "deliver good outcomes", cannot be reduced to binary predicates, and LLM-as-judge is increasingly used as the substitute. Our… 15 arXiv — NLP / Computation & Language research 13d ago Research-Oriented Human-Centric Evaluation for Foundation Models arXiv:2506.01793v2 Announce Type: replace Abstract: Most current evaluations of foundation models focus on objective benchmarks, such as knowledge coverage and reasoning accuracy, often overlooking users' subjective experiences in human-AI collaboration. To address this gap, we… 14 Hugging Face Daily Papers research 13d ago CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing Abstract CPI-Bench is a comprehensive benchmark for real-world image editing that evaluates multi-image tasks, practical applications, and reasoning-based editing to better differentiate model performance. Generated by thinkingmachines/Inkling-Small With the rapid advancement of… 36 Hugging Face Daily Papers research 13d ago HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark Abstract HumanTracker introduces a large-scale benchmark and preference-aligned metric to evaluate humanoid motion tracking based on perceptual quality and physical contact stability. Generated by thinkingmachines/Inkling-Small Humanoid motion tracking is central to… 5 Hugging Face Daily Papers research 13d ago Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Abstract A new benchmark for AI-generated video detection reveals that current detectors fail to generalize across realistic crisis-related videos and become less reliable as content spreads socially. Generated by thinkingmachines/Inkling-Small Recent video generators can… 27 r/MachineLearning community 13d ago Input 4-5x Reduction with sentence and keyword based trie on chat. [P] Currently struggling with an automatic budget selection, at 25% it’s very similar to benchmarks accuracy and seems even better on actual chat input however it many times retrieves too much. It would be nice to add an algorithm that actually can determine better retrieval other… 11 r/LocalLLaMA community 14d ago Is ternary (1.58-bit) LLMs making a come back? I'm just thinking, ever since microsoft announced bitnet, this sub (and myself) has been hoping for massive ternary models. In the last month alone, prismML dropped 27B ternary (though I've read community experience suggested it sometimes didn't hold up to it's benchmarks),… 33 r/LocalLLaMA community 14d ago Qwen3.8-27B abliterated FP8: refusal 64–99% → 0–6%, and MMLU/GSM8K move less than 1.3 points Been reading the eval table on the abliterated Qwen3.8-27B FP8 build instead of the release notes. It's published as red-team material, disclaimer and all, so the numbers are the interesting part. Refusal across the usual harmful-instruction sets (AdvBench, HarmBench,… 12 r/LocalLLaMA community 14d ago SOTA Apple Silicon Inference (August 15, 2026) This is a HANDWRITTEN post. I spent way too much time trying to get fast inference on Apple Silicon. This post is for people who want to know what's the latest on running local models on their mac, and why they may not be seeing the performance others in the community claim.… 34 r/LocalLLaMA community 14d ago club-5060ti refresh: tested RTX 5060 Ti presets, a proper high-context harness, and Qwen3.8 27B Quick update on the RTX 5060 Ti local LLM repo. It has changed quite a bit since my previous posts. The project started as a collection of practical notes and benchmark results. That was useful, but as the dataset grew it became harder to answer the question most people actually… 10 r/LocalLLaMA community 15d ago [Megathread] Qwen 3.8 27B Release Day Megathread to help with the influx of duplicate / similar posts around the release of the Qwen 3.8 27B release. Quants Fine-Tunes & Abliterations Chat Templates Inference Server Support & Configuration Experiences, Benchmarks & Model Comparisons Official:… 10 r/LocalLLaMA community 15d ago A 150M param recurrent model scores 29.5% on ARC-AGI-1 at $0.0007 per task Not a transformer. It's a recurrent latent reasoning setup that keeps "thinking" in latent space before answering. Sits completely outside the published cost/accuracy frontier for ARC-AGI, and something this size runs on basically anything. Paper is from the Pathway team,… 25 r/LocalLLaMA community 15d ago Qwen3.8 Benchmarks Converted to Charts https://preview.redd.it/uvdb7g4o7djh1.png?width=1600&format=png&auto=webp&s=1db6176144bcf8f40efb15cc04e8370ad52bfe2f https://preview.redd.it/rj4l66eq7djh1.png?width=1257&format=png&auto=webp&s=692b5f76c96e5910fff22ebf5765945af8ef7ec5… 17 TechCrunch — AI news-outlet 15d ago Google will now allow users to remove visible watermark from its AI generations Turning off this setting won't affect invisible benchmarks used to identify an AI generated file. 27 r/LocalLLaMA community 16d ago A preliminary Qwen3.8-27B model card is live! If you scroll down from the countdown at https://huggingface.co/Qwen/Qwen3.8-27B , you see a big model card with a bunch of sections: Highlights, Model Overview, Quickstart, Best Practices, Citation, etc! No benchmarks on this yet as far as I can tell. We'll still need to wait… 15 Smol AI News news-outlet 16d ago not much happened today **Z.ai launched GLM-5.3**, a coding- and cyber-focused model with significant gains on agentic and security benchmarks, achieved through scaled post-training rather than a larger base model. **Alibaba released Qwen3.8-27B**, a native multimodal dense model under Apache 2.0 with… 5 Hugging Face Daily Papers research 16d ago H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models Abstract H2R-Bench evaluates video generation models on transforming human manipulation videos into robot-centric demonstrations across embodiment constraints and interaction fidelity. Generated by thinkingmachines/Inkling-Small Large-scale manipulation data is essential for… 22 arXiv — Machine Learning research 16d ago When Can You Trust Offline Evaluation of Equal-Cost Top-k Allocation? A Controlled, Reproducible Benchmark and Practitioner's Guide arXiv:2608.12489v1 Announce Type: new Abstract: Organizations decide whom to treat under a budget and want to know what a targeting rule would have earned before deploying it. Off-policy evaluation promises this from logged data, but the deployable rule is a deterministic top-k… 18 arXiv — Machine Learning research 16d ago CoMedBench: A Multi-Source Benchmark of Synthetic Medical Data Fidelity and Downstream Utility arXiv:2608.12805v1 Announce Type: new Abstract: Access to clinical data is essential for developing reliable healthcare machine learning systems, but direct use of electronic health records is constrained by privacy regulation, institutional review, data-use agreements, and the… 38 arXiv — Machine Learning research 16d ago Beyond Simulated Benchmarks: Evaluating Motion Representations for Fall Detection Under Real-World Data Scarcity arXiv:2608.13197v1 Announce Type: new Abstract: Falls are a major health concern for older adults, and wearable sensors have been widely explored for detecting falls and enabling timely intervention. However, real-world falls are extremely rare: collecting 100 of them requires… 10 arXiv — Machine Learning research 16d ago Large-scale Testing Global Optimization Methods with Black-box Adversarial Attacks arXiv:2608.13296v1 Announce Type: new Abstract: Existing global optimization benchmark suites are of a moderate size and are based on a small number of analytical functions that date back even to the 1970s. This causes a risk of biasing the development of global optimization… 24 arXiv — NLP / Computation & Language research 16d ago Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition arXiv:2608.12327v1 Announce Type: new Abstract: Multilingual pretrained models nominally support Nepali, yet no controlled benchmark has compared them under a single fine-tuning protocol. We fine-tune six pretrained models (XLSR-53, IndicWav2Vec, MMS-1B, Whisper-Medium,… 34 arXiv — NLP / Computation & Language research 16d ago Are Large Language Models Reliable Reviewers? A Benchmark for Error Detection in Financial Documents arXiv:2608.12342v1 Announce Type: new Abstract: Ensuring the accuracy of financial documents is critical for economic analysis, regulatory compliance, and corporate decision-making. Several studies have shown that Large Language Models (LLMs) perform well in many financial… 7 arXiv — NLP / Computation & Language research 16d ago Vision-Language Models are Fragile Multilingual Associators arXiv:2608.12333v1 Announce Type: new Abstract: Vision-language models must associate visual entities with textual attributes. Whether these associations or concept bindings remain stable when the language of the input changes is unexplored. We introduce M$^2$BIND, a benchmark… 37 arXiv — NLP / Computation & Language research 16d ago Class-Structure Preservation Beats Diversity: A Comprehensive Benchmark of Text Augmentation Methods for Imbalanced Text Classification arXiv:2608.12340v1 Announce Type: new Abstract: With the rapid advancement of large language models (LLMs), generative data augmentation has attracted considerable attention for imbalanced text classification in natural language processing. However, no empirical benchmark to… 7 arXiv — NLP / Computation & Language research 16d ago The "Knowledge-Behavior Gap" in Cultural Taboo Safety of Large Language Models arXiv:2608.12341v1 Announce Type: new Abstract: Cultural taboo safety is essential for deploying large language models (LLMs), as culturally insensitive outputs may cause offense or even social harm. However, existing cultural benchmarks primarily assess cultural knowledge or… 11 arXiv — NLP / Computation & Language research 16d ago Large Language Models Pass the History Exam But Miss the <<History>>: A Polish High School Exit Exam Matura Benchmark arXiv:2608.12343v1 Announce Type: new Abstract: AI chatbots are widely used by students as knowledge sources, yet LLM benchmarks rarely assess interpretative historical reasoning. We evaluate eight leading LLMs on the Polish high school exit exams (Matura) in history - three… 12 arXiv — NLP / Computation & Language research 16d ago Unified Multi-Dimensional Benchmark for Complex Graph Reasoning in Large Language Models arXiv:2608.12391v1 Announce Type: new Abstract: Graph reasoning provides a promising testbed for evaluating the reasoning ability of large language models (LLMs), as graph instances can be programmatically generated, structurally controlled, and naturally scaled to long-input… 36 arXiv — NLP / Computation & Language research 16d ago Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark Contamination Detection arXiv:2608.12652v1 Announce Type: new Abstract: Benchmark contamination is diagnosed today with n-gram overlap, with likelihood-based membership inference, or with canary strings, and each needs something usually unavailable: the training corpus, a well-chosen test statistic, or… 7 arXiv — NLP / Computation & Language research 16d ago BavGround: A Benchmark for Regional Cultural Grounding and Dialect Competence in Bavarian arXiv:2608.12894v1 Announce Type: new Abstract: Cultural evaluation of large language models (LLMs) often focuses on high-resource standard languages, leaving regional culture and dialect communities underrepresented. We introduce BavGround, a benchmark for evaluating Bavarian… 35 arXiv — NLP / Computation & Language research 16d ago LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation arXiv:2608.13136v1 Announce Type: new Abstract: With the rapid advancement of large language models (LLMs), research idea generation has attracted increasing attention. Existing approaches enable LLMs to retrieve relevant literature and propose novel ideas for research areas.… 32 arXiv — NLP / Computation & Language research 16d ago How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures arXiv:2608.13267v1 Announce Type: new Abstract: Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty… 20 arXiv — NLP / Computation & Language research 16d ago Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation arXiv:2608.13326v1 Announce Type: new Abstract: LLM benchmark scores can be precise even when the observation protocol does not identify the behavioral property they are intended to measure. In a controlled, solver-grounded setting, we formalize a protocol-level identifiability… 20 arXiv — NLP / Computation & Language research 16d ago When AI Is Your Pastor: A Benchmark for Theological Triage and Pastoral Guidance in Large Language Models arXiv:2608.12324v1 Announce Type: cross Abstract: People increasingly ask large language models (LLMs) for counsel on questions of faith, doctrine, and pastoral care. These questions are not ordinary information requests. Some ask about core Christian beliefs, some ask about… 37 arXiv — NLP / Computation & Language research 16d ago Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists arXiv:2608.12345v1 Announce Type: cross Abstract: Language models are increasingly deployed as co-scientists, yet their ability to uphold research integrity under institutional pressure remains unmeasured. We introduce IntegrityBench, a benchmark evaluating misconduct… 9 arXiv — NLP / Computation & Language research 16d ago SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries arXiv:2608.12654v1 Announce Type: cross Abstract: Long-running LLM agents act through tools, and a single step can send an email, merge a pull request, or wire a payment. The steering decision is the pre-commit choice at that boundary: proceed, or hold for human or policy… 38 Hugging Face Daily Papers research 16d ago LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers Abstract LLM routing is formalized as a sequential decision process with a unified benchmark and modular infrastructure to compare and improve cost-effective model selection. Generated by thinkingmachines/Inkling-Small No single large language model (LLM) is optimal across all… 6 Hugging Face Daily Papers research 16d ago DarwinX: Evolving Agent Harnesses Through Natural Selection Abstract DarwinX evolves agent harnesses via population selection with frozen models, improving verified performance across benchmarks without benchmark-specific patches. Generated by thinkingmachines/Inkling-Small An LLM agent's capability depends not only on model weights but… 30 Hugging Face Daily Papers research 16d ago AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design Abstract AutoDesign uses a meta-harness optimizer to recursively improve a code agent for structured media generation, achieving state-of-the-art results on paper-to-poster synthesis. Generated by thinkingmachines/Inkling-Small Transforming multimodal sources into condensed and… 11 Hugging Face Daily Papers research 16d ago PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives Abstract PlayWorld benchmarks interactive video world models by using multi-modal agents to pursue long-horizon objectives, evaluating geometry consistency, interaction fidelity, and state evolution. Generated by thinkingmachines/Inkling-Small Video world models simulate future… 23 Hugging Face Daily Papers research 17d ago AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research Abstract The benchmark evaluates autonomous coding agents on open-ended world-model research by having them iteratively improve a starter model across game environments using a shared structured-state format. Generated by thinkingmachines/Inkling-Small World modeling is an… 20 Smol AI News news-outlet 17d ago not much happened today **Google** rapidly released **Gemini 3.7 Flash** just three weeks after 3.6 Flash, targeting coding, web development, knowledge work, and agentic workflows with a 50% introductory price cut and improved benchmark scores like **DeepSWE 65.3%** and **Code Arena Elo 1588**. The… 17 arXiv — Machine Learning research 17d ago Analysis of Federated Aggregation under Model Poisoning and Backdoor Attacks: A Reconstructed Cross-Dataset and Cross-Architecture Benchmark arXiv:2608.11423v1 Announce Type: new Abstract: Robust comparisons of federated aggregation methods require joint consideration of predictive performance, threat definitions, metric semantics, and execution provenance. A 500-cell seed-1 evaluation matrix was reconstructed across… 5 arXiv — NLP / Computation & Language research 17d ago Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs arXiv:2608.11232v1 Announce Type: new Abstract: Evaluating LLM coding agents in algorithmic trading is difficult because static benchmarks risk data contamination and numerical backtest outputs require ground truth from actual code execution. We present Backtrader-Bench, a… 18 arXiv — NLP / Computation & Language research 17d ago CT-$\Delta$Bench: A Benchmark for Longitudinal 3D Medical Imaging Difference Reporting with Vision-Language Models arXiv:2608.11534v1 Announce Type: new Abstract: In medical imaging, the clinical value of Computed Tomography (CT) lies not only in depicting current disease status, but crucially in enabling longitudinal comparison of serial scans to determine disease evolution, a process that… 10 arXiv — NLP / Computation & Language research 17d ago The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance arXiv:2608.11694v1 Announce Type: new Abstract: A benchmark score comes from a single phrasing of each problem. That single phrasing is treated as if it stood for the whole space of ways the same problem could be asked, but it does not. We show that rephrasing a problem while… 11 arXiv — NLP / Computation & Language research 17d ago Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems arXiv:2608.11879v1 Announce Type: new Abstract: Long-running conversational agents increasingly rely on a memory system to avoid resending the whole conversation each turn, yet how much that costs to serve has received little systematic benchmarking. We compare three memory… 15 arXiv — NLP / Computation & Language research 17d ago LODESTAR: Trustworthy Entropy Is Navigated, Not Merely Measured -- Reinforced Polarizer Keeps a Frozen LLM from Being Confidently Misled by the Wrong Evidence arXiv:2608.11922v1 Announce Type: new Abstract: Predictive-distribution entropy makes a strong selection rule in retrieval-augmented question answering: across five QA benchmarks, keeping the candidate answer that a frozen respondent LLM produces with the lowest answer-token… 8 arXiv — NLP / Computation & Language research 17d ago Accuracy and Order Sensitivity Diverge Under Label-Free Strategies arXiv:2608.11947v1 Announce Type: new Abstract: Multiple-choice benchmarks are widely used to evaluate large language models, but MCQ scores conflate knowledge with sensitivity to option order, which makes them unreliable measures of model knowledge. In this paper, we test… 35 Page 6 of 10 · 500 articles ← Newer Older →