News / #benchmark Tag Benchmark 500 articles archived under #benchmark · RSS Sign in to follow arXiv — NLP / Computation & Language research 9d ago ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models arXiv:2608.20338v1 Announce Type: new Abstract: Large Language Models (LLMs) increasingly require selective removal of harmful or sensitive knowledge, called unlearning, yet existing methods and benchmarks fail to evaluate this capability completely. Current approaches rely on… 20 arXiv — NLP / Computation & Language research 9d ago Measuring What a Specification Determines: A Formal Semantic-Block Model and an Execution-Judged Benchmark arXiv:2608.19475v1 Announce Type: cross Abstract: This work introduces a formal semantic-block model for specifications and an execution-judged benchmark for evaluating specification quality independently of model capability. A specification is represented as a structure… 27 arXiv — NLP / Computation & Language research 9d ago Can Agent Memory Systems Track Evolving State? arXiv:2608.19652v1 Announce Type: cross Abstract: As LLM-based agents are deployed for longer and higher-stakes tasks, their memory systems continue to have crucial gaps. While existing memory benchmarks focus largely on recall-shaped tasks, we argue an effective memory system… 7 arXiv — NLP / Computation & Language research 9d ago MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use arXiv:2608.20202v1 Announce Type: cross Abstract: Memory has become a key component of large language models, enabling them to retain information and learn from long-term interactions. However, existing memory benchmarks mainly evaluate whether information is correctly… 35 arXiv — NLP / Computation & Language research 9d ago ContractScrub: A benchmark for final review of legal contracts arXiv:2608.20204v1 Announce Type: cross Abstract: Legal work, with its heavy reliance on processing large amounts of text, is often considered one of the domains most exposed to the use of LLMs. Contract ``scrubbing,'' the final review of transactional agreements for errors and… 17 arXiv — NLP / Computation & Language research 9d ago AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement arXiv:2608.20318v1 Announce Type: cross Abstract: Recursive self-improvement (RSI) asks whether an AI system can improve the process that produces AI systems, so that the next system inherits the improvement. That process is the training algorithm: a better objective or update… 29 arXiv — NLP / Computation & Language research 9d ago DiverValue-Bench: A Benchmark and Fine-Tuning Framework for Aligning Large Language Models with Diverse Human Values arXiv:2509.08022v3 Announce Type: replace Abstract: Aligning large language models (LLMs) with diverse human values is essential for safe and effective deployment, yet existing benchmarks often overlook cultural and demographic variation. We introduce DiverValue-Bench, a… 13 r/LocalLLaMA community 9d ago SenseNova U1.5-Lite full release: expert training, OPD distillation, one model at inference Benchmarks: Benchmark U1 Preview Full Qwen-Image-Bench 47.14 55.20 (PE) 60.18 (PE) ImgEdit 3.9 4.37 4.59 GEdit-Bench-EN 7.47 8.14 8.26 Instead of just scaling up, they train task-specialized expert models for text rendering and infographics, aesthetic quality, and image editing.… 15 Hugging Face Daily Papers research 9d ago MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use Abstract Retrieved memories can induce reasoning errors and belief distortions in large language models, and an inference-time strategy helps avoid these cognitive traps while maintaining benchmark performance. Generated by thinkingmachines/Inkling-Small Memory has become a key… 12 Hugging Face official-blog 9d ago Measuring benchmark optimization in speech recognition Back to Articles a]:hidden"> Measuring benchmark optimization in speech recognition Published August 21, 2026 Update on GitHub Upvote 1 Theo Lebryk tlebryk02 HumeAI Eric Bezzam bezzam Alice aliceebaird HumeAI David Ayllon dayllon HumeAI Jakub Piotr Cłapa jpc HumeAI Jens Madsen… 22 r/LocalLLaMA community 9d ago Been tweaking my Qwen 3.8 setup, up to 45+ steady T/ps at 8bit quant. Realised I'm now top T/ps for this model+ctx across all benchmarked M-series chips. Full args linked below, happy to discuss as this was a pain of trial and error. https://omlx.ai/benchmarks/performance/2pko3m1k - you can expand the raw args, but I have full annotations of what worked and what didn't. I'm now testing the model on acutal coding and haven't seen any issues with performance vs default suggested vals for the vanilla model.… 35 r/LocalLLaMA community 9d ago Fine-tuning Cactus Needle 2 can match DeepSeek v4 on the specific task Hey LocalLlama, Henry from Cactus here! When we trained Needle 2, I had a strict rule to not expose the model to any data sample that remotely felt like these benchmarks. It seemed over-the-top, but benchmarks are easy to overfit around, yet struggle in the wild, especially… 9 r/LocalLLaMA community 9d ago Qwen3.8-27B scored 29/30 on AIME 2026 with FP8 + xhigh reasoning — BF16 vs FP8 results I benchmarked Qwen3.8-27B on MathArena/aime_2026 dataset, comparing BF16 and FP8 weights at medium and xhigh reasoning effort. Interesting findings are: quantized FP8 xhigh is better than BF 16 medium equally good as 16 BF xhigh with better speed. On problem 7, both BF16 xhigh… 28 r/LocalLLaMA community 9d ago [MASSIVE TINY RELEASE] - Supra2-Medium-Base - a tiny 25M parameters model competing heavily with our previous 50M model! Hey guys! Supra2-Medium is finally out! It's a 25M parameters qwen3 architecture model trained entirely from scratch (on our new rig: RTX 5060 Ti 16GB + the new RTX 5060 8GB!). Here's how it competes in benchmarks with Supra-50M-Base (which is double as large!!):… 8 r/LocalLLaMA community 10d ago Aurora-80K releases! A modern tiny language model. I'm introducing Aurora-80K, a small language model with exactly 80 thousand parameters. It uses a factorized 4,096-token vocabulary despite having only 80K parameters. The benchmarks: Wikitext-2 BPB: 3.2902 BLiMP: 52.31% Arc-Easy: 26.05% More information about the model is… 7 r/LocalLLaMA community 10d ago [Draft - Open PR] AVX2: Speed up large batch size prompt processing of IQ models by bartowski1182 · Pull Request #27402 · ggml-org/llama.cpp IQ quants are particularly slow on CPU at large batch sizes (what you'd see for imatrix and perplexity) Benchmark numbers I ran PPL against master and this PR to get speed and numbers on --chunks 50 for Qwen3.6-27B and Qwen3.6-35B-A3B on EPYC 9654 using 24 threads Created pure… 23 r/LocalLLaMA community 10d ago New benchmark just dropped! The pelican on a bicycle is sooo outdated, so I came up with a new, improved version. Qwen3.8-27b medium (UD-Q4_K_XL) vs. Sol 5.6 high vs. Qwen3.6-35B (UD-Q6_K_XL) Prompt (only real with typo!): "Create a svg of a horse on a blue bycicle in the desert, with a camel in the… 15 Hugging Face Daily Papers research 10d ago FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents Abstract FM-Bench evaluates long-horizon decision-making of LLM agents managing a football club over 20 years, revealing that managerial behavior rather than scale or token spend drives performance. Generated by thinkingmachines/Inkling-Small Language model agents now execute… 31 r/LocalLLaMA community 10d ago 3 days benchmarking most llama.cpp flags on my weird 40gb vram laptop + tb4 egpu setup. Got +70% generation, +40% prefill, 60k more context, and filed a bug in llama around MTP. What I learned. tldr: went from 16~ t/s to 27~ t/s generation. got my usable context up from 220k to the full 262k without sacrificing anything. prefill also increased from 376 to 573 command I ended up with, fwiw: llama-server -m Qwen3.8-27B-UD-Q6_K_XL.gguf -c 262144 -ngl 999 -fa on \ -ctk… 5 arXiv — Machine Learning research 10d ago ClosureBench: A Constructive Benchmark for Compositional Graph Reasoning arXiv:2608.18242v1 Announce Type: new Abstract: We introduce ClosureBench, a constructive benchmark for compositional graph-relational reasoning with programmatically verified ground truth. Unlike fixed-test-set benchmarks vulnerable to data contamination, ClosureBench generates… 13 arXiv — Machine Learning research 10d ago An Empirical Benchmark of Deep Time-Series Models for Smart Meter Energy Forecasting arXiv:2608.18675v1 Announce Type: new Abstract: Accurate forecasting of energy consumption is important for the efficient operation of power systems, with direct implications for operational costs, energy management, and system maintenance. Due to the availability of extensive… 14 arXiv — Machine Learning research 10d ago A FEM-Based Surrogate Modelling and Optimization Framework for Physics-Constrained Electromagnetic Coil Design arXiv:2608.18903v1 Announce Type: new Abstract: This work evaluates surrogate-assisted optimization of a seven-parameter current-excited coil--core benchmark subject to geometric, manufacturing, and separate core and copper mass constraints. A Python--MPh--COMSOL workflow… 32 arXiv — Machine Learning research 10d ago Lost in Aggregation: How Benchmarks Overlook Irreplaceable Model Strengths arXiv:2608.18919v1 Announce Type: new Abstract: Tabular machine learning benchmarks typically summarize performance by averaging scores, ranks, or pairwise wins across datasets. Such aggregates are useful for selecting robust default models, but they can obscure a different… 6 arXiv — NLP / Computation & Language research 10d ago Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis arXiv:2608.18940v1 Announce Type: cross Abstract: Single-step retrosynthesis is a central component of computer-aided synthesis planning, yet its intrinsically one-to-many nature is poorly captured by single-answer evaluation and benchmarking protocols. To address this, we… 11 arXiv — NLP / Computation & Language research 10d ago LongNovel: A Multi-Scale Benchmark for Hallucination Detection in Long-Context Novel Summarization arXiv:2608.18082v1 Announce Type: new Abstract: Although context windows have expanded significantly in recent years, hallucinations in long-context summarization remain a challenge. Long novels are better suited than news or papers for researching these hallucinations, due to… 25 arXiv — NLP / Computation & Language research 10d ago FrenchNews-7: Benchmarking Cross-Publisher French News Editorial Desk Classification arXiv:2608.18097v1 Announce Type: new Abstract: We present FrenchNews-7, a cross-publisher France-based French-language news editorial desk classification benchmark combining a large multi-outlet corpus, a URL-derived seven-class taxonomy, and a fine-tuned CamemBERT classifier.… 8 arXiv — NLP / Computation & Language research 10d ago MemFuse: Multi-Source Memory Fusion from Fragmented Observations arXiv:2608.18704v1 Announce Type: new Abstract: Long-term memory is essential for agents that operate across extended interactions, yet existing memory systems and benchmarks predominantly focus on single-source textual histories. In realistic settings, however, relevant… 10 arXiv — NLP / Computation & Language research 10d ago What Makes Software Issue Resolution Tasks Difficult for Agents? arXiv:2608.18280v1 Announce Type: cross Abstract: Background. Advances in agentic systems are simultaneously, and rapidly, saturating benchmarks. Despite this often discussed phenomena, benchmark scores remain difficult to interpret due to the lack of control and… 37 arXiv — NLP / Computation & Language research 10d ago ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents arXiv:2608.18307v1 Announce Type: cross Abstract: Current evaluation of computer-use agents is split between long-horizon workflow benchmarks and atomic GUI-grounding tests. This leaves an under-instrumented middle layer: realistic component-centered interactions (e.g., toggle a… 26 arXiv — NLP / Computation & Language research 10d ago GreekBarRetrieval: A Benchmark for Greek Statutory Retrieval arXiv:2608.18752v1 Announce Type: cross Abstract: Statutory retrieval is necessary for citation-grounded legal question answering, but remains underexplored for Greek. We introduce GreekBarRetrieval, a public retrieval benchmark derived from, and complementing GreekBarBench,… 25 Hugging Face Daily Papers research 10d ago Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence Abstract Zetta is a closed-loop embodied harness that evolves runtime critics and recovery skills online to govern physical execution at action frequency, achieving high success on robot benchmarks with faster inference and scaling self-exploration. Generated by… 36 Hugging Face Daily Papers research 10d ago SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation Abstract SoftVTBench introduces a synchronized visuo-tactile dataset and deformation-aware benchmark for evaluating physical interaction quality during deformable-object manipulation. Generated by thinkingmachines/Inkling-Small Physical interaction quality is central to… 24 Hugging Face Daily Papers research 10d ago SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Abstract Semantic task completion video generation evaluates whether generated videos achieve intended outcomes with semantic grounding, supported by a curated dataset and vision-language model-based benchmark. Generated by thinkingmachines/Inkling-Small We introduce Semantic… 22 r/LocalLLaMA community 10d ago Qwen3.8-27B took a serious hit to *knowledge* vs 3.6 Like many of you I've spent the last few days throwing Qwen3.8-27B against all of my usual use-cases and personal tasks/harnesses and workflows. It's great, phenomenal sometimes, but that's not what this post is about. One of my little personal benchmarks is a little set of… 10 Hugging Face Daily Papers research 10d ago Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis Abstract Top-K prompting and plausibility-aware training improve diverse reaction prediction in single-step retrosynthesis, yielding state-of-the-art results on a large verified reaction dataset and motivating ensemble systems. Generated by thinkingmachines/Inkling-Small… 34 r/LocalLLaMA community 10d ago Qwen 3.8 27B SlopCodeBench results Howdy, I'm back again - running my favorite benchmark (it's still unsaturated for the time being so might as well!) previous runs a b https://github.com/michaelasper/benchmarks/blob/main/qwen3.8-27b-pi-on-slop-code-bench.md I ran this via OpenRouter because my mac would cry… 10 r/LocalLLaMA community 10d ago Introducing Qwen3.8-27B Dynamic v3 Unsloth GGUFs Hey everyone! We’re releasing new Qwen3.8-27B GGUFs with 10% higher accuracy for the same size. This uses a new version of Dynamic v3.0 Unsloth Dynamic V3 outperforms others by >10% on Div-300, KLD & more benchmarks. We also release 1-bit quants that retain 77% accuracy. Run on… 34 Hugging Face Daily Papers research 10d ago PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX Abstract PTXBench evaluates large language models on architecture-specific GPU kernel optimization, revealing uneven success and performance gaps that supervised fine-tuning only partially addresses. Generated by thinkingmachines/Inkling-Small We introduce PTXBench, a benchmark… 27 r/LocalLLaMA community 10d ago Ornith-1.5 (397B [DeepSWE 56], 35B-A3B, 9B) Aloha! 🌺Introducing Ornith-1.5, a family of open-source LLMs spanning 9B Dense, 35B MoE, and 397B MoE, trained with self-improving strategies. It achieves state-of-the-art performance among open-source models of comparable size and delivers performance comparable to Claude Opus… 28 Hugging Face Daily Papers research 11d ago StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows Abstract StartupBench evaluates end-to-end AI agents on real-world startup workflows and reveals that even top models complete only about 30% of tasks, highlighting gaps in instruction following and domain expertise. Generated by thinkingmachines/Inkling-Small Recent advances in… 7 Hugging Face Daily Papers research 11d ago CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing Abstract A new dataset, benchmark, and 22B model enable compositional instruction-guided video editing with multi-region attention and temporal coherence. Generated by thinkingmachines/Inkling-Small The quality and diversity of instruction-based video editing datasets are… 11 Hugging Face Daily Papers research 11d ago HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety Abstract HarnessRisk evaluates agent harness safety across six operational phases, revealing that configuration vulnerabilities and detection gaps allow high attack success despite preserved utility. Generated by thinkingmachines/Inkling-Small Large language models are… 31 arXiv — Machine Learning research 11d ago Benchmarking Classical and Transformer-Based Models for Document Sensitivity Classification arXiv:2608.16928v1 Announce Type: new Abstract: Automatic sensitivity classification of organizational documents is a critical yet underserved problem, where the consequences of misclassification range from regulatory violations to security breaches. While AI-based approaches… 9 arXiv — Machine Learning research 11d ago Deep Learning for Cross-Border Electricity Price Forecasting: A Comparative Study arXiv:2608.17091v1 Announce Type: new Abstract: While publicly available electricity market data presents a valuable resource for forecasting research, the field lacks established benchmark datasets for standardized comparison. As a result, many studies have relied on different… 12 arXiv — Machine Learning research 11d ago Beyond MSE: Rethinking the Evaluation Metric and Benchmarking for Irregular Time Series Forecasting arXiv:2608.17293v1 Announce Type: new Abstract: Existing research on irregular time-series forecasting has primarily focused on model design, while evaluation metrics remain insufficiently studied. Existing benchmarks typically use mean squared error (MSE) as the evaluation… 30 arXiv — Machine Learning research 11d ago Evaluating RL Explainability Methods by How Much They Help Fix Bugs in Agents arXiv:2608.17524v1 Announce Type: new Abstract: This preliminary paper outlines a planned evaluation benchmark for Explainable Reinforcement Learning (XRL) methods. Current evaluations rely on functionally-grounded metrics like faithfulness and compactness, and on human-grounded… 31 arXiv — NLP / Computation & Language research 11d ago An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning arXiv:2608.17804v1 Announce Type: cross Abstract: Practical LLM unlearning is usually evaluated through two objectives: suppress target-specific knowledge and preserve non-target utility. In generative QA, this leaves a third behavior underspecified: when a target-adjacent… 4 arXiv — NLP / Computation & Language research 11d ago Temporal Leakage in Financial News NLP: A Multi-Architecture Audit with a Regime-Specific M&A Signal arXiv:2608.17223v1 Announce Type: new Abstract: Financial-news direction prediction has become a popular NLP benchmark, yet reported gains depend critically on whether the train-test split is chronological or random, i.e., on temporal leakage. We audit this dependence on a… 8 arXiv — NLP / Computation & Language research 11d ago PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX arXiv:2608.17379v1 Announce Type: new Abstract: We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-specific PTX for GPU kernel optimization. PTXBench measures functional correctness, whether selected target… 17 arXiv — NLP / Computation & Language research 11d ago Effects of Answer Format Variation on Gender Bias in Large Language Models arXiv:2608.17516v1 Announce Type: new Abstract: Gender bias or other social biases in large language models (LLMs) are frequently evaluated with question answering or survey benchmarks where the LLM needs to give a response in a predefined answer format. It is well known in… 6 Page 4 of 10 · 500 articles ← Newer Older →