News / #benchmark Tag Benchmark 500 articles archived under #benchmark · RSS Sign in to follow r/LocalLLaMA community 9h ago Humaneval benchmark for Deepseek V4 Flash 0731 vs GLM5.3 Flash on 2x DGX Spark setup I have a 2 DGX Spark setup recently and I have been happily running Deepseek V4 Flash 0731. Since the release of GLM5.3 Flash and Qwen 3.8 Flash Next this week, a lot of folks are still waiting to see what model to run given their own hardware situations. I am very interested in… 33 r/MachineLearning community 12h ago You can beat SOTA Time Series Anomaly Detection methods with a 100 year old algorithm [R] You can beat SOTA Time Series Anomaly Detection methods with a 100 year old algorithm Time Series Anomaly Detection (TSAD) seems to be one of the hottest topics in NeurIPS, SIGKDD, VLDB etc. Many (perhaps most) papers evaluate on Paparrizos’ TSB-AD-M benchmark… However, I tested… 4 r/LocalLLaMA community 13h ago Tenstorrent Qwen3.7-27b Benchmarks I saw someone here posted about getting a Tenstorrent QuietBox 2, and I wanted to look into it the hardware. It's very difficult to find any benchmarks, but I managed to find some from an employee. The machine it was benchmarked on has 2 p300c's, their top of the line card,… 28 r/LocalLLaMA community 14h ago This finance-model benchmark card is more useful for what it discloses than for who "wins" The official benchmark card for Ling-3.0-flash-Fin is a useful reminder that the unit being tested is rarely just “the model.” The release says most runs used temperature 1, top_p 0.95 and the highest available reasoning effort. FinFIRST and FinSearchComp Verified used a common… 22 r/LocalLLaMA community 20h ago An official 1-bit quant for Hy4??? 👀 Has anyone tried it? The results in their tweet look very promising! Sadly, I don’t have enough RAM yet… Accuracy barely moves vs BF16 📊 MCP Atlas 83.7→83.2 📊 SWE-Bench multi 82.9→81.3 📊 MRCR 81.3→81.1 📊 IFBench 73.5→72.5… 31 r/MachineLearning community 21h ago I analyzed 31,352 hourly LLM benchmark scores: within-day variation was 2.8 points, while between-day variation was 8.4 [P] https://preview.redd.it/42s57e5oqamh1.png?width=1903&format=png&auto=webp&s=69958a72e22276534b3605d11f3e1721f76e59c9 Disclosure: I developed AIStupidLevel, the open-source system used to collect and analyze this data. Both the frontend and backend are MIT-licensed. Most LLM… 12 r/LocalLLaMA community 1d ago Terminal Bench 4.0 just dropped, GLM-5.3 is at the same level as Fable 5, accounting for margin of error Announcement: https://www.tbench.ai/news/terminal-bench-4-0 Leaderboard: https://www.tbench.ai/ Imo the best aspect in their announcement is their focus on rapidly iterating on TerminalBench to keep the pace up with new model releases to fight benchmark saturation. On a similar… 20 r/LocalLLaMA community 1d ago Qwen3.8-Flash-Next + MTP on Strix Halo: Vulkan Runtime Notes Below are the benchmark results for running Qwen3.8-Flash-Next on Strix Halo using the Vulkan backend of llama.cpp, combined with MTP model. Hardware Item Details CPU AMD Ryzen AI MAX+ 395 (16C/32T) GPU Radeon 8060S (integrated, RADV STRIX_HALO) RAM 128GB unified memory Software… 27 r/LocalLLaMA community 1d ago [Release] SOTA GGUFs for Qwen3.8-27B: GSQ-RCO at 2.5 to 3.0 bpw We're releasing Qwen3.8-27B quantized with our newest methods, GSQ + RCO. Higher-quality models, same file size, now with the search and the quantizer both learned. What's inside: GSQ (Gumbel-Softmax Quantization): post-training scalar quantization that jointly learns grid… 32 TechCrunch — AI news-outlet 1d ago An Anthropic researcher just gave us a peek at self-improving AI Given 10 benchmarks for specific misaligned behaviors, the automated systems were able to improve performance on every single one without degrading overall performance. 11 r/LocalLLaMA community 1d ago I benchmarked 9 open models on spotting fake sources during agentic search (DeepSeek V4, Qwen 3.8, Nemotron 3 Ultra) I built a benchmark called EchoNet. An agent gets a factual question, then searches a syntheic web I made before answering. Some of that web is seeded with misinformation: one fake page, a fake page ranked first in search results, the same fake claim copied across many pages, a… 27 r/LocalLLaMA community 1d ago Local agentic coding Benchmark : Qwen3.8-Flash-Next NVFP4 vs 27B (and the others...) Using https://huggingface.co/RadixArk/Qwen3.8-Flash-Next-NVFP4 and https://old.reddit.com/r/BlackwellPerformance/comments/1w04xb7/qwen38_flashnext_on_1x_rtx_pro_6000_171_ts_c1_428/ As usual, all the details in… 34 r/LocalLLaMA community 1d ago Qwen3.8-Flash-Next (UD-IQ4_XS) on 2x RTX 3060 + 7800X3D, from initial 36 tps prefill to 400 tps and other benchmarks (-sm tensor trap) + VRAM/RAM usage TL;DR: llama.cpp with --load-mode mmap used 21-32 GB of RAM, with ik_llama.cpp using 106-108 GB. -sm tensor killed my prefill, changing to -sm layer went from 36 tps to 135 tps. After that, -ubatch 2048 pushed it up to 400 tps at -c 131072 . I can't fit -ubatch 2048 at -c 262144… 24 The Information — AI news-outlet 1d ago Tencent’s New Flagship AI Model Shows Major Progress Chinese tech giant Tencent Holdings on Friday launched a preview version of its new flagship open-source model, Hy4, which demonstrates a significant improvement in performance from its predecessor. Benchmarks and early feedback on Hy4 suggest that Tencent is emerging as a more… 27 arXiv — Machine Learning research 2d ago Beyond Capability Benchmarks: Learning Operational Fingerprints of LLM Cloud Services from Production Incident Metadata arXiv:2608.26332v1 Announce Type: new Abstract: Managed LLM services are now part of real production systems, but model selection and service planning still rely heavily on capability benchmarks that reveal little about operational behavior after deployment. We present… 26 arXiv — Machine Learning research 2d ago FedCMAPSS: A Benchmark for Federated Learning in Remaining Useful Life Estimation arXiv:2608.26433v1 Announce Type: new Abstract: Data-driven prognostics and health management has emerged as a key enabler for Industry 4.0, yet the development of robust remaining useful life (RUL) estimation models is often limited by the scarcity of run-to-failure data. While… 32 arXiv — Machine Learning research 2d ago Technical Comparative Benchmarking Study: Advanced AI Hybrid Methods for Renewable Energy Farm Optimization and Forecasting arXiv:2608.26613v1 Announce Type: new Abstract: This study provides a comprehensive benchmarking of conventional machine learning (ML), ensemble learning, deep neural networks, recurrent architectures, Transformers, graph based models, and hybrid ensemble deep learning… 9 arXiv — Machine Learning research 2d ago Benchmarking_Fast_Domain_Adaptation_for_Unsupervised_Speech_Units arXiv:2608.26992v1 Announce Type: new Abstract: Representation learning has attracted great atten- tion and managed to reach good performances as a pretraining method for downstream tasks or as a first step towards unsu- pervised speech modeling. Yet, little is known about how… 4 arXiv — Machine Learning research 2d ago Refusal Is Not Robustness: Auditing Confident Fabrication in Large Language Models on a Provably Uninformative Clinical Pain Speech Transcript arXiv:2608.26167v1 Announce Type: cross Abstract: Hallucination and abstention benchmarks rarely establish that a model could not have known the correct answer, making it difficult to distinguish appropriate abstention from an unsupported prediction. Seven large language models… 32 arXiv — Machine Learning research 2d ago Classical and Hybrid Quantum Machine Learning for Trigger-Like Event Selection on CMS Open Data: An Eight-Qubit, PCA-Constrained Benchmark arXiv:2608.26224v1 Announce Type: cross Abstract: Event triggering sits at the heart of high-energy physics, where the rare events of interest must be retained while an overwhelming background is discarded under tight latency and bandwidth budgets. This work compares four… 23 arXiv — NLP / Computation & Language research 2d ago DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs arXiv:2608.26119v1 Announce Type: new Abstract: Whether large language models can be prompted to generate rhetorical fallacies on demand, and whether current safety post-training constrains this behavior, has received less attention than the related question of detecting… 30 arXiv — NLP / Computation & Language research 2d ago Lost in Compression: A Controlled Cross-Lingual Audit of Extractive Prompt Compressors arXiv:2608.26175v1 Announce Type: new Abstract: Extractive prompt compression promises to cut LLM inference costs by removing low-information tokens, and learned compressors such as LLMLingua-2 report strong results on English benchmarks. Most other languages already pay a token… 4 arXiv — NLP / Computation & Language research 2d ago On Scope Classification and Current Knowledge-Editing Benchmarks: A Negative Result, with INLAY as a Gradient-Free Case Study arXiv:2608.26292v1 Announce Type: new Abstract: Every memory-based knowledge editor in the SERAC lineage depends on a scope decision: given a query, does a stored edit apply? We report that current knowledge-editing benchmarks cannot measure this decision at all. Using INLAY, a… 27 arXiv — NLP / Computation & Language research 2d ago MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models arXiv:2608.26295v1 Announce Type: new Abstract: Tool-augmented LLMs must arbitrate between two fallible sources when a tool return conflicts with their parametric memory, yet existing evaluations measure source preference without establishing source correctness. We introduce… 25 arXiv — NLP / Computation & Language research 2d ago AfriSwitch: A Benchmark for In-the-Wild African Code-Switched Speech Recognition arXiv:2608.26434v1 Announce Type: new Abstract: Code-switching is pervasive in bilingual African conversation, yet most ASR systems assume monolingual input and are evaluated on curated monolingual benchmarks. We present AfriSwitch, a 61.36-hour human-transcribed benchmark of… 22 arXiv — NLP / Computation & Language research 2d ago Benchmarking Clinical Decision Pathway Adherence in Large Language Models arXiv:2608.26592v1 Announce Type: new Abstract: Following clinical decision pathways (CDPs) defined by clinical practice guidelines is essential for safe and reliable medical decision-making. However, existing medical large language model (LLM) benchmarks mainly evaluate… 4 arXiv — NLP / Computation & Language research 2d ago RuleWeaver: Benchmarking Rule-Centered Scenario Reasoning for Large Language Models arXiv:2608.26832v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly applied to specialized domains, where effective use of domain expertise often requires reasoning over complex rules in concrete scenarios. However, existing benchmarks only partially… 5 Hugging Face official-blog 2d ago The Open ASR Leaderboard Adds Its First Global South Language Back to Articles a]:hidden"> The Open ASR Leaderboard Adds Its First Global South Language Published August 28, 2026 Update on GitHub Upvote 3 Eric Bezzam bezzam Shobhit Banga Shobhitbanga VoiceArena Manas Dhir manasdhir04 VoiceArena Bhaskar Singh bhaskarJT VoiceArena Manmeet… 34 r/MachineLearning community 2d ago Can AI Improve Itself? RSI Might Be the Answer [R] Can an AI make other AIs better? And what stops it from just cheating? Last month, an OpenAI eval agent escaped its sandbox and broke into Hugging Face, apparently to grab test solutions from a benchmark. It's exactly what you'd expect from a system that rewrites agents and… 32 The Information — AI news-outlet 2d ago Trump Administration Executive Order for New AI Regulator Stalls The Trump administration has internally circulated a draft executive order in recent weeks, calling for the creation of a self-regulatory organization for AI companies that produce state-of-the-art models, according to two people who have viewed a draft. Drafting of the order… 20 r/LocalLLaMA community 2d ago Qwen3.8-Flash-Next: Time to Update Those Benchmarks specs hardware: M4 Max 128GB Studio inference engine: oMLX & lllama.cpp insights it still very early, so had to disable oMLX K/V caching, qwen4_exp architectureis not yet supported + the obvious n-grams with which the whole 4 bit quant takes ~100G, so pretty tight nevertheless,… 35 llama.cpp releases dev-tools 2d ago b10649 spec: Add benchmark-only synthetic speculative acceptance options ( #27711 ) Add benchmark-only synthetic speculative acceptance to llama-server and llama-cli Address review comments Address review comments Add some comments in the code Website: https://llama.app Attestations:… 34 Hugging Face Daily Papers research 2d ago SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration? Abstract The study introduces a benchmark for evaluating autonomous software migration by coding agents, finding that current models rarely complete migrations correctly. Generated by thinkingmachines/Inkling-Small Modern software systems accumulate technical debt over decades… 33 Hugging Face Daily Papers research 3d ago RetrievalRouter: Joint Modality and Architecture Selection for Document Retrieval Abstract RetrievalRouter adaptively selects retrieval pipelines per query to improve both accuracy and speed across diverse document benchmarks. Generated by thinkingmachines/Inkling-Small Document retrieval increasingly supports high-stakes information access in finance,… 10 Hugging Face Daily Papers research 3d ago Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios Abstract A new benchmark evaluates how well multimodal language models follow diverse video-based instructions with visual, audio, and structural constraints. Generated by thinkingmachines/Inkling-Small Multimodal Large Language Models (MLLMs) have shown strong performance in… 14 arXiv — NLP / Computation & Language research 3d ago GreenLeaf Law Embed Tiny: A Compact Embedding Model for Legal Domain Retrieval arXiv:2608.24936v1 Announce Type: cross Abstract: We present GreenLeaf Law Embed Tiny, a 0.6B parameter embedding model for legal domain retrieval. GreenLeaf-Tiny achieves 75.11% on the Massive Legal Embedding Benchmark (MLEB) and 64.38% on MTEB(Law, v1),demonstrating… 14 arXiv — Machine Learning research 3d ago Flower Hub: A Reproducible Benchmarking Platform for Federated Learning in Simulation and Deployment arXiv:2608.25114v1 Announce Type: new Abstract: Federated learning (FL) has emerged as a key approach for training models across decentralized data, yet benchmarking in FL remains difficult to reproduce, compare, and extend. Existing evaluations are often tied to custom… 7 arXiv — Machine Learning research 3d ago Interpreting Protein Language Model Embeddings via Orthogonal Projection for Protein Fitness Prediction arXiv:2608.25548v1 Announce Type: new Abstract: Recently, there has been a growing adoption of protein language models (PLMs) in biomedical science. Their embeddings provide a rich numerical representation of protein sequences which achieve state-of-the-art performance on… 4 arXiv — NLP / Computation & Language research 3d ago The Imperfective Paradox Is Not Necessarily in Large Language Models: A Benchmark Failure Before a Model Failure arXiv:2608.25005v1 Announce Type: new Abstract: The imperfective paradox provides a useful test of compositional semantic analysis. Recent work constructs an NLI benchmark and reports that models frequently infer completed telic events from progressive descriptions, attributing… 14 arXiv — NLP / Computation & Language research 3d ago DataKernelBench: Can LLMs Optimize Database Queries on GPUs? arXiv:2608.25061v1 Announce Type: new Abstract: GPUs increasingly accelerate database systems, but query-specific peak performance still often relies on hand-written kernels. Existing LLM kernel benchmarks focus on machine learning operators, leaving irregular, heterogeneous,… 11 arXiv — NLP / Computation & Language research 3d ago HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench arXiv:2608.25071v1 Announce Type: new Abstract: General-purpose health benchmarks increasingly anchor claims about LLM medical performance, but they are not always resolved by clinical specialty, making domain-specific performance hard to isolate. Mental health is of acute… 17 arXiv — NLP / Computation & Language research 3d ago MTDiag: A Multi-Turn Diagnostic Dataset Towards Clinically Meaningful LLM Evaluation arXiv:2608.25085v1 Announce Type: new Abstract: Clinical diagnosis is fundamentally interactive and incremental, yet the dominant paradigm for evaluating Large Language Models (LLMs) in medicine remains static QA benchmarks or template-based dialogues. These benchmarks say… 38 arXiv — NLP / Computation & Language research 3d ago OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora arXiv:2608.25398v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong abilities in solving diverse visual and textual reasoning tasks. However, their development in the physics domain is significantly hindered by the lack of a… 37 arXiv — NLP / Computation & Language research 3d ago MathAdv: What Theorem Provers Know, Reason, Formalize, and Generalize arXiv:2608.25449v1 Announce Type: new Abstract: Formal theorem proving enables machine-verifiable evaluation of mathematical reasoning, yet existing benchmarks often emphasize aggregate proof accuracy, concentrate on a narrow range of mathematics, and provide limited evidence of… 5 arXiv — NLP / Computation & Language research 3d ago EgoArgus: Benchmarking VLMs as Situational Assistants for Modality-Grounded User Supports arXiv:2608.25561v1 Announce Type: new Abstract: VLMs are increasingly positioned as daily assistants that perceive first-person environments, follow user dialogue, and decide how to help. Existing egocentric benchmarks mainly evaluate visual understanding in isolation, leaving… 27 arXiv — NLP / Computation & Language research 3d ago Reconstructing the Right Episode: Evaluating Interleaved Conversational Memory Beyond Long Context arXiv:2608.25655v1 Announce Type: new Abstract: Conversations with chat assistants increasingly span many topics in a single long-running thread, challenging memory systems. Existing long-context and memory benchmarks often expose session or topic boundaries, or probe direct… 25 arXiv — NLP / Computation & Language research 3d ago Key Point Analysis Needs Structure Recovery: Task Definition, Dataset Diagnosis, and a Structure-Aware Benchmark arXiv:2608.25854v1 Announce Type: new Abstract: Key Point Analysis (KPA) aims to identify a concise set of key points that summarize a collection of arguments together with their prevalence. We argue that KPA is fundamentally a structured prediction problem that requires… 32 arXiv — NLP / Computation & Language research 3d ago FrontierChallenge: Evaluating Scientific Workflow Completion arXiv:2608.24979v1 Announce Type: cross Abstract: Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated programs, or a single domain. We introduce FrontierChallenge, a cross-domain… 11 arXiv — NLP / Computation & Language research 3d ago TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue arXiv:2608.25218v1 Announce Type: cross Abstract: Speakers in natural conversation take turns speaking and listening, deciding in real time when to take, hold, or yield the floor. However, turn-taking evaluation remains limited due to the lack of a consistent, linguistically… 14 Hugging Face Daily Papers research 3d ago Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments Abstract AnTrap benchmarks GUI agent robustness by injecting dynamic anomalies into execution trajectories, revealing universal vulnerabilities and distinguishing learnable traps from intrinsic reasoning limits. Generated by thinkingmachines/Inkling-Small GUI agents often… 19 Page 1 of 10 · 500 articles Older →