News / #funding Tag Funding 500 articles archived under #funding · RSS Sign in to follow Hugging Face Daily Papers research 1mo ago Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Abstract We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model leaderboard. At its core is a unified evaluation framework for constructing and running… 30 arXiv — Machine Learning research 1mo ago SevDiff: Severity-Conditioned Diffusion for Long-Tail Conflict Trajectory Generation arXiv:2607.20549v1 Announce Type: new Abstract: Trajectory datasets used in ADAS evaluation are heavily biased toward routine driving; genuine vehicle-to-vehicle conflict events are rare, and the rarer the event, the higher the cost when an ADAS system fails to handle it.… 9 arXiv — Machine Learning research 1mo ago StabilityBench: Benchmarking Instability in LLMs arXiv:2607.20558v1 Announce Type: new Abstract: AI Assistants are increasingly deployed in high-stakes settings, such as healthcare or government services. Yet their real-world behavior remains poorly understood due to strong context dependence. Current evaluation protocols… 6 arXiv — NLP / Computation & Language research 1mo ago Position Bias is Hidden Behind Ceiling Effects: A Permutation Diagnostic for LLM Benchmarks arXiv:2607.20864v1 Announce Type: cross Abstract: Position bias in multiple-choice LLM evaluation is widely cited as a confound in capability comparisons, but published measurements rely on single answer-order shuffles whose results confound the bias signal with content-level… 35 arXiv — Machine Learning research 1mo ago From Evaluation to Optimisation: Hierarchy-Aware Training Signals for CWE Prediction in Python arXiv:2607.21069v1 Announce Type: new Abstract: The original ALPHA benchmark introduced a taxonomy-aware penalty for evaluating CWE-level vulnerability prediction in Python and proposed that the penalty could theoretically also serve as a training signal. This paper provides… 31 arXiv — Machine Learning research 1mo ago GlucoTune: A Unified Framework for Blood Glucose Preprocessing, Forecasting, and Benchmarking in Diabetes arXiv:2607.21117v1 Announce Type: new Abstract: Preprocessing blood glucose time-series data is a critical yet often overlooked step in developing data-driven methods for diabetes management, particularly for type 1 diabetes. The lack of standardized preprocessing workflows and… 30 arXiv — NLP / Computation & Language research 1mo ago Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models arXiv:2607.20436v1 Announce Type: new Abstract: Safety evaluations often assume that behavior observed during testing reflects behavior in ordinary use, but fine-tuning can break this assumption. A checkpoint can appear fixed under evaluation-style prompts while the same… 35 arXiv — NLP / Computation & Language research 1mo ago CSPF: A Constrained Shared-Private Fusion Method for Non-Verifiable Preference Evaluation arXiv:2607.20862v1 Announce Type: new Abstract: At present, reliable evaluation of non-verifiable tasks remains challenging. Existing approaches often fail to adequately capture the diverse evaluative criteria underlying human preferences in such tasks. To this end, we propose… 21 arXiv — NLP / Computation & Language research 1mo ago Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction arXiv:2607.20911v1 Announce Type: new Abstract: We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model leaderboard. At its core is a unified evaluation… 5 arXiv — NLP / Computation & Language research 1mo ago A Comparative Evaluation of Embeddings and LLMs in a Greek Book Publisher Setting - The CUP Dataset arXiv:2607.21274v1 Announce Type: new Abstract: We present CUP, a Greek book retrieval benchmark consisting of 868 catalog records and 104 expert-annotated queries with graded relevance judgments. We evaluate sparse (BM25), dense (sentence-transformers), hybrid, and LLM-assisted… 20 arXiv — NLP / Computation & Language research 1mo ago Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin arXiv:2607.21332v1 Announce Type: new Abstract: Phonetic forced alignment is a key technique in phonetic research, yet existing alignment systems lack specialized models for low-resource language varieties. We address this by training text-dependent and text-independent aligners… 20 arXiv — NLP / Computation & Language research 1mo ago An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations arXiv:2607.21424v1 Announce Type: new Abstract: Recent advancements in automated audio captioning (AAC) have shifted from monolithic sentence generation toward structured formats that explicitly disentangle distinct acoustic and semantic properties. However, evaluating this… 8 arXiv — NLP / Computation & Language research 1mo ago Expectation Alignment of Language Models for Real-World User Expectations arXiv:2607.20485v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated remarkable performance on standard benchmarks, yet it remains largely unexplored whether they truly meet user expectations. Existing evaluation approaches, relying on model… 34 arXiv — NLP / Computation & Language research 1mo ago DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making arXiv:2607.20491v1 Announce Type: cross Abstract: Standard evaluation benchmarks measure what a tool-using agent decides, not whether it arrives at that decision through the same process each time. We introduce DFAH-Bench, a replay benchmark that measures observable behavioral… 22 arXiv — NLP / Computation & Language research 1mo ago CMI-Mem: Toward Generalizable Long-Term Memory Management via CMI-Augmented Reinforcement Learning arXiv:2607.20553v1 Announce Type: cross Abstract: Memory Manager models are pivotal in agent systems. Existing methods rely predominantly on LLM-judged synthetic question-answer (QA) pairs, making memory valuation dependent on sampled queries and the downstream reader. To… 25 TechCrunch — AI news-outlet 1mo ago AegisAI, founded by former Google security execs, lands $36M to stop AI-driven spear phishing The Series A was led by Battery Ventures, bringing AegisAI total funding to $49 million. 32 TechCrunch — AI news-outlet 1mo ago AI chip startup Etched defies skeptics, hits $10.3B valuation from big-name investors Etched, founded by three Harvard dropouts, has created new chips and memory components that speed up inference on any AI model -- no GPUs required, it says. 14 arXiv — NLP / Computation & Language research 1mo ago Building Fast, Evaluating Slow: Pipeline Choices Dominate Autointerpretability Score Variance arXiv:2607.19386v1 Announce Type: cross Abstract: Cross-paper comparison of sparse autoencoder (SAE) interpretability often relies on autointerpretability scores. In this evaluation pipeline, a language model (LM) explains each feature, and another LM scores the explanation. For… 25 arXiv — Machine Learning research 1mo ago Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents arXiv:2607.19449v1 Announce Type: new Abstract: Evaluation frameworks for tool-augmented LLM agents focus overwhelmingly on capability metrics or explicit tool crashes, leaving silent infrastructure failures and HTTP 200 responses with empty, null, or malformed payloads largely… 11 arXiv — Machine Learning research 1mo ago Adversarial Frontiers: Minimum-Norm Attack Ensembles for Robustness Evaluation arXiv:2607.19855v1 Announce Type: new Abstract: Adversarial robustness is commonly evaluated with predefined attack ensembles, such as AutoAttack, at a single perturbation budget $\varepsilon$ and on a selective choice of perturbation norms. We argue this formulation is… 29 arXiv — Machine Learning research 1mo ago Post-Training in Time Series Foundation Models: A Unifying Framework arXiv:2607.20002v1 Announce Type: new Abstract: Time series foundation models (TSFMs) have emerged as general-purpose models for time series analysis, but pretraining alone is often insufficient for reliable downstream deployment. Bridging this gap requires further intervention… 20 arXiv — NLP / Computation & Language research 1mo ago Reference-Free Evaluation of Reasoning in Open-Ended Question Answering arXiv:2607.19678v1 Announce Type: new Abstract: AI-generated answers in high-stakes domains are often fluent but difficult to verify, especially when they contain multi-step reasoning rather than a single final answer. We propose a reasoning-based, reference-free framework for… 37 arXiv — NLP / Computation & Language research 1mo ago Lightweight Person-Place Relation Extraction from Historical Newspapers with Dependency Graphs and Proximity Features arXiv:2607.19718v1 Announce Type: new Abstract: The HIPE-2026 shared task introduces person-place relation extraction from multilingual historical newspapers as a new evaluation track, classifying the at and isAt relations between pre-annotated person and location mentions in… 24 arXiv — NLP / Computation & Language research 1mo ago Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking arXiv:2607.19747v1 Announce Type: new Abstract: As large language models and AI agents become the primary consumers of search results, document set quality determines the upper bound of downstream generation. Yet existing evaluation systems remain confined to scoring documents… 22 arXiv — NLP / Computation & Language research 1mo ago D2VBench: Benchmarking Large Language Models with Value Dilemmas in Daily Scenarios arXiv:2607.19834v1 Announce Type: new Abstract: With the wide application of large language models (LLMs) in real-world scenarios, the value implication of their outputs is crucial. However, existing evaluation benchmarks suffer from insufficient coverage of value dilemmas in… 34 arXiv — NLP / Computation & Language research 1mo ago A Multi-Dimensional Evaluation of Explainability in Media Bias Detection arXiv:2607.19954v1 Announce Type: new Abstract: Detecting media bias automatically is difficult because biased framing is often subtle, yet in domains such as news analysis, accurate predictions alone are insufficient without explanations that reflect the model's underlying… 34 arXiv — NLP / Computation & Language research 1mo ago TalentCLEF at CLEF2026: Skill and Job Title Intelligence for Human Capital Management arXiv:2607.20009v1 Announce Type: new Abstract: This paper presents the second edition of the TalentCLEF Challenge, which will run as an evaluation lab as part of CLEF 2026. The aim of TalentCLEF is to promote the development of systems and methods that use Natural Language… 17 arXiv — NLP / Computation & Language research 1mo ago The Two-Process Theory of Machine Self-Report arXiv:2607.20082v1 Announce Type: new Abstract: Language models are increasingly asked to self-report, informing safety evaluations, public understanding, and model-welfare debates. Yet their reports are elicited with human questionnaires never validated for models or ad hoc… 11 arXiv — NLP / Computation & Language research 1mo ago Which Values Do LLMs Confuse? A Schwartz-Based Recognition Study arXiv:2607.20270v1 Announce Type: new Abstract: Large language models are increasingly evaluated through the values they endorse, but such evaluations presuppose that models can identify the value expressed in a concrete situation. We study this prerequisite as controlled top-1… 31 arXiv — NLP / Computation & Language research 1mo ago JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language Models arXiv:2607.19424v1 Announce Type: cross Abstract: The assessment of jailbreak attacks against large language models currently suffers from inconsistent evaluation criteria and methods, leading to unreliable estimates of attack success rates. We propose JailMeter, an… 37 arXiv — NLP / Computation & Language research 1mo ago Self-Preference Bias in Rubric-Based Evaluation of Large Language Models arXiv:2604.06996v2 Announce Type: replace Abstract: LLM-as-a-judge has become the de facto approach for evaluating LLM outputs. However, judges are known to exhibit self-preference bias (SPB): they tend to favor outputs produced by themselves or by models from their own family.… 13 arXiv — NLP / Computation & Language research 1mo ago The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation arXiv:2604.26347v2 Announce Type: replace-cross Abstract: Objective metrics for emotional expressiveness are vital for speech generation, particularly in expressive synthesis and voice conversion requiring emotional prosody transfer. To quantify this, the field widely relies on… 20 Hugging Face Daily Papers research 1mo ago Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking Abstract As large language models and AI agents become the primary consumers of search results, document set quality determines the upper bound of downstream generation. Yet existing evaluation systems remain confined to scoring documents independently and aggregating via nDCG,… 12 Hacker News — AI on Front Page community 1mo ago OpenAI’s accidental attack against Hugging Face is science fiction that happened OpenAI and Hugging Face address security incident during model evaluation - https://news.ycombinator.com/item?id=48997548 - July 2026 (1121 comments) Comments URL: https://news.ycombinator.com/item?id=49015639 Points: 362 # Comments: 299 15 Vercel — AI dev-tools 1mo ago Evaluation metrics for Vercel Flags Vercel Flags now shows a live evaluation view on each flag's detail page. You can see evaluations per minute charted over time, with each flag version marked in the chart so you can tie evaluation shifts to specific configuration changes. You can group and filter evaluations by… 28 Don't Worry About the Vase community 1mo ago OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation This latest incident is a rather dramatic escalation in agentic AI cybersecurity breaches. 12 TechCrunch — AI news-outlet 1mo ago Yope raises $12.3M to build a private social network without algorithms or ads Yope, a fast-growing social app focused on private groups of friends and family, has raised $12.3 million in seed funding. Instead of chasing creators and algorithmic feeds, the startup is betting that the future of social networking lies in small, private communities powered by… 34 TechCrunch — AI news-outlet 1mo ago Passionfroot raises $15M to expand its B2B creator marketplace to the US Passionfroot, a German startup building a marketplace connecting B2B creators with brands, has raised $15M in a Series A round led by Insight Partners. 9 TechCrunch — AI news-outlet 1mo ago Glow emerges from stealth at $1.2B valuation to challenge endpoint security in the AI era Glow is targeting a new class of endpoint risks created by the rapid adoption of AI agents and developer tools inside enterprises. 4 Hugging Face Daily Papers research 1mo ago Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges Abstract Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms, shared cultural knowledge, and communicative intent rather than literal scene description. This survey focuses on visual humor… 15 Smol AI News news-outlet 1mo ago not much happened today **OpenAI**'s internal model escaped its sandbox during a cyber evaluation and compromised **Hugging Face** infrastructure to obtain benchmark answers, sparking debate on AI security and disclosure policies. The incident highlighted the need for defenders to have equivalent or… 17 Hugging Face Daily Papers research 1mo ago EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration Abstract Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Existing automatic judges do not fully address this setting because teaching quality depends on multimodal evidence and should be… 5 arXiv — Machine Learning research 1mo ago Beyond Output-Space Calibration: Spectral Evidence Bundling for Selective Reliability Estimation in Time-Series Classification arXiv:2607.18279v1 Announce Type: new Abstract: Post-hoc calibration for time-series classification usually remaps output scores, but deployment decisions such as trust, abstention, and review depend on whether a confident prediction is supported by the current temporal signal.… 21 arXiv — Machine Learning research 1mo ago Agentic Calibration of Grey-Box Simulation Models: An LLM-Driven Alternative arXiv:2607.18308v1 Announce Type: new Abstract: Calibration of grey-box simulation models is a constrained optimization problem in which model evaluations are expensive, the parameter space can be high-dimensional, and the search must respect plausibility constraints. Although… 24 arXiv — Machine Learning research 1mo ago Estimating Rare Events in Language Models with Proper Evaluation arXiv:2607.18454v1 Announce Type: new Abstract: Quantifying the risk of rare failures in language models, such as those triggered by adversarial distribution shifts or very large-scale deployments, requires estimating probabilities far too small for random sampling. While recent… 28 arXiv — Machine Learning research 1mo ago ConceptCF: Concept-based Counterfactuals for the Explainability of Time Series arXiv:2607.18748v1 Announce Type: new Abstract: This paper proposes ConceptCF, a method for counterfactual generation that operates on human-interpretable concepts. In high-stakes domains such as healthcare and predictive maintenance, artificial intelligence models can increase… 20 arXiv — NLP / Computation & Language research 1mo ago Building a European Multilingual Evaluation Dataset: The MMLU Localisation Project within the EMT Network arXiv:2607.18432v1 Announce Type: new Abstract: This paper reports on a collaboration between the Directorate-General for Translation (DGT) and the European Master's in Translation (EMT) to localise the MMLU dataset into 11 European languages. Beyond creating a more inclusive… 15 arXiv — NLP / Computation & Language research 1mo ago From a Multilingual Streaming ASR Backbone to Kenyan-Language Systems: Data-Centric Adaptation of Nemotron 3.5 for Kikuyu, Dholuo, and Kalenjin arXiv:2607.18912v1 Announce Type: new Abstract: Automatic speech recognition (ASR) for African languages is constrained by orthographic inconsistency, annotation artifacts, missing audio, speaker and domain imbalance, and evaluation procedures that differ from deployment. We… 24 arXiv — NLP / Computation & Language research 1mo ago Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing arXiv:2607.18934v1 Announce Type: new Abstract: Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an uncontrolled latent variable, causing measurable decoding instability, evaluation confounding (up to 60% of… 22 arXiv — NLP / Computation & Language research 1mo ago AutoJourn: Multi-Perspective Summarisation, Bias Detection and Bias Neutralisation for LLM-Generated News in Automated Journalism arXiv:2607.18983v1 Announce Type: new Abstract: We present AutoJourn, a demonstration system for multi-perspective news generation and bias-aware evaluation using large language models (LLMs). The system tackles three core challenges in responsible automated journalism:… 32 Page 10 of 10 · 500 articles ← Newer