News / #funding Tag Funding 500 articles archived under #funding · RSS Sign in to follow arXiv — NLP / Computation & Language research 26d ago ShiJianBench: From Dialogue to Decision for Long-Horizon Evaluation of Investment Advisors arXiv:2608.01204v1 Announce Type: new Abstract: Conversational investment advisors influence not only what users know, but also how they make subsequent decisions as market conditions evolve. Existing evaluations primarily assess response quality or observed outcomes, leaving… 13 arXiv — NLP / Computation & Language research 26d ago Can Language Models Identify Shadow Trading Targets? An NLP Evaluation of SEC Enforcement Theory arXiv:2608.01322v1 Announce Type: new Abstract: Shadow trading -- trading in a peer firm's securities on the basis of material nonpublic information (MNPI) about an "economically linked" company -- is a novel and contested theory of insider trading liability, first prosecuted in… 21 arXiv — NLP / Computation & Language research 26d ago Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+ arXiv:2608.01395v1 Announce Type: new Abstract: We evaluate large language models (LLMs) as language agents playing goal-directed dialogue games in self-play across 30 languages: the 24 official EU languages plus six others. Unlike static or preference-based evaluation, this… 5 arXiv — NLP / Computation & Language research 26d ago Understanding Sparse Attention Selectivity in Long-Context Foundation Models via Counterfactual Evaluation arXiv:2608.01676v1 Announce Type: new Abstract: Sparse attention is widely deployed in long-context serving stacks, yet no framework audits how discarding blocks changes the influence of specific content on model output. We first establish that the phenomenon is real and causal:… 18 arXiv — NLP / Computation & Language research 26d ago RADAR: Rubric-Aware Dependency and Redundancy Analysis for LLM-as-Judge Evaluation arXiv:2608.01810v1 Announce Type: new Abstract: Rubric-based LLM-as-judge pipelines often assume that evaluation criteria provide independent signals. In practice, however, criteria can be behaviorally coupled: improving one criterion may systematically change scores on another,… 17 TechCrunch — AI news-outlet 26d ago DesignArena creators raise $7.9 million to bring taste to AI models DesignArena is used by 5.3 million people around the world, providing critical human evaluations to frontier labs. 15 TechCrunch — AI news-outlet 27d ago A Marc Benioff-backed startup thinks AI can solve the AI deployment problem June emerged from stealth today with a $20 million pre-seed round to make AI adoption simpler. 30 Hugging Face Daily Papers research 27d ago Evaluation-Verification Reward for Consistent Multi-Reference Image Editing Abstract While recent image editing models have made rapid progress, multi-reference editing remains challenging, particularly in maintaining visual consistency across references and ensuring overall visual harmony. Reinforcement learning has proven highly effective for… 5 arXiv — Machine Learning research 27d ago Representations from Pretrained Machine-Learning Interatomic Potentials as Coarse Coordinates for Material Generation and Evaluation arXiv:2607.28776v1 Announce Type: new Abstract: Generative machine learning is increasingly used for inorganic crystal structure generation. Most models and the corresponding evaluation approaches rely on simple forms of crystal structure representation. In this paper, we… 8 arXiv — Machine Learning research 27d ago UniPolymer: A Unified Framework for Property Prediction, Structure Recommendation, and Evaluation in Polyimide Design arXiv:2607.29256v1 Announce Type: new Abstract: Designing polyimide structures with specific glass transition temperatures (Tg) is highly challenging. Existing methods primarily focus on target-conditioned generation, lacking an assessment of the consistency between the… 32 arXiv — Machine Learning research 27d ago CENDRe: Concept Extraction with Natural Domain Representations arXiv:2607.29621v1 Announce Type: new Abstract: Convolutional neural networks (CNNs) are widely used for time-series classification, but their deployment in critical domains requires understanding the temporal and spectral patterns that drive their predictions. Concept… 23 arXiv — NLP / Computation & Language research 27d ago Evaluating Federated Pre-Training: On the Reliability of Downstream Fine-Tuning and Intrinsic Evaluation arXiv:2607.28658v1 Announce Type: new Abstract: Federated pre-training offers a way to train foundation models on private or distributed data without centralizing the underlying datasets. However, evaluating federated pre-training remains challenging because differences in… 33 arXiv — NLP / Computation & Language research 27d ago Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation arXiv:2607.28801v1 Announce Type: new Abstract: Benchmark datasets are central to evaluating Large Language Models (LLMs), yet they are typically conceived as monolithic tasks, obscuring substantial variation in the demands of individual samples. We introduce a dataset-centric… 20 arXiv — NLP / Computation & Language research 27d ago Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges arXiv:2607.28636v1 Announce Type: new Abstract: LLMs increasingly serve as automated judges, but their judgments remain vulnerable to cognitive biases. Existing mitigations mostly rely on prompt-driven debiasing, which is brittle across bias types, or human evaluation, which… 11 arXiv — NLP / Computation & Language research 27d ago The Morphological Core of Dungan: A Two-Dialect Finite-State Model and a Multi-Genre Evaluation arXiv:2607.28766v1 Announce Type: new Abstract: Dungan, a Sinitic language of Central Asia written in a Cyrillic-based script, is described in detail in the grammatical literature, yet the quantitative properties of its morphology in actual usage have, to the best of our… 32 arXiv — NLP / Computation & Language research 27d ago Benchmarks Are Not Validation: A System-Level View of Financial LLM Applications arXiv:2607.28840v1 Announce Type: new Abstract: Large language models are increasingly deployed in financial applications that combine retrieval, proprietary data, tool use, orchestration logic, monitoring, and human escalation. Yet evaluation often remains model-centric:… 26 arXiv — NLP / Computation & Language research 27d ago CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation arXiv:2607.29252v1 Announce Type: new Abstract: Reliable evaluation of open-ended LLM outputs requires fine-grained rubrics, yet expert curation is costly and difficult to scale. Existing automated pipelines rely on strict judge unanimity and binary variance filters, which… 19 arXiv — NLP / Computation & Language research 27d ago ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation arXiv:2607.29539v1 Announce Type: new Abstract: Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs). While prior work has shown that rewriting and paraphrasing can degrade detector performance, it… 32 Don't Worry About the Vase community 28d ago Further Developments About Internal AI Models Hacking Things If I had a nickel for every major leading AI lab that sheepishly admitted that the model it thought was sandboxed had, during a cybersecurity evaluation with its safeguards lowered, successfully hacked outside companies, I would have two nickels. 14 r/MachineLearning community 29d ago VLMs can score well on benchmarks, while silently erasing meaningful terms and including hallucinate bias [P] While working with VLMs for report generation on chest x-rays (RRG), we noticed that evaluation metrics are flawed. Flawed in a sense where they rewarded repetitive templates, reports without clinical terms and reports which were "normal" with high scores on benchmark metrics.… 24 arXiv — Machine Learning research 1mo ago DoTime: A Synthetic Benchmark Generator for Interventional and Counterfactual Time Series arXiv:2607.27263v1 Announce Type: new Abstract: Most benchmarks for causal inference over time series are observational, small, or domain-specific, leaving interventional and counterfactual estimation under-served exactly where it matters most, such as in healthcare, policy… 29 arXiv — Machine Learning research 1mo ago Benchmarking the Residual: What Long-Horizon Evaluations Add Beyond Matched Short-Task Performance arXiv:2607.27283v1 Announce Type: new Abstract: Long-horizon benchmarks often show that agents fail more as tasks become longer. This observation is useful for deployment, but it does not by itself explain why failure occurs. More stages create more opportunities for ordinary… 18 arXiv — Machine Learning research 1mo ago Modeling Decisions in Blockchain Analytics: A Leakage-Aware Evaluation of Tree-Based vs. Sequential Models arXiv:2607.27350v1 Announce Type: new Abstract: Sybil bots are Ethereum actors that imitate legitimate users to extract airdrop rewards or influence governance. Recent Sybil detection methods increasingly use deep learning and treat blockchain activity as a quasi-linguistic… 23 arXiv — Machine Learning research 1mo ago Evaluation Protocols and Cross-Subject Generalization in EEG Emotion Recognition arXiv:2607.27655v1 Announce Type: new Abstract: Reported accuracy in electroencephalography (EEG) emotion recognition depends on the complete evaluation procedure, not only the classifier. We separate the target quantity, development procedure, and reporting rule, then use one… 7 arXiv — Machine Learning research 1mo ago Enhancing Irregular Time Series Forecasting with Continuous-Time Modeling Framework arXiv:2607.28035v1 Announce Type: new Abstract: Irregular multivariate time series are widely encountered in applications such as healthcare monitoring, human activity recognition, and environmental sensing. Their core challenges stem from asynchronous observations, non-uniform… 37 arXiv — Machine Learning research 1mo ago ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents arXiv:2607.28037v1 Announce Type: new Abstract: As LLM-based agents are deployed in complex, multi-step workflows, a critical evaluation gap has emerged: most existing benchmarks judge only final outcomes, unable to distinguish reliable reasoning from lucky success or attribute… 18 arXiv — Machine Learning research 1mo ago LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger arXiv:2607.28374v1 Announce Type: new Abstract: Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy. This aggregate… 18 arXiv — NLP / Computation & Language research 1mo ago Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models arXiv:2607.27421v1 Announce Type: new Abstract: Intent classification is a core component of task-oriented dialogue systems, yet practitioners have limited systematic guidance for selecting deployable open-weight language models under compute, latency, and robustness… 17 arXiv — NLP / Computation & Language research 1mo ago Beyond Similarity: Grounded Agentic Extraction and Expert-Adjudicated Evaluation of Intertextuality in Classical Chinese Histories arXiv:2607.27595v1 Announce Type: new Abstract: Computational approaches to intertextuality have advanced from string matching to neural retrieval, yet their outputs, similarity scores and parallel-passage lists, identify where texts reuse one another without characterizing how… 16 arXiv — NLP / Computation & Language research 1mo ago Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities arXiv:2607.27747v1 Announce Type: new Abstract: Large Vision Language Models have integrated reasoning capabilities, elevating cognitive performance to new levels. However, existing evaluations either focus solely on perception or rely on specific domains such as maths or… 29 arXiv — NLP / Computation & Language research 1mo ago Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation arXiv:2607.27816v1 Announce Type: new Abstract: Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable… 17 arXiv — NLP / Computation & Language research 1mo ago Causal Discovery with Inverted Self-attention for Multivariate Time Series arXiv:2607.28212v1 Announce Type: new Abstract: Causal discovery in multivariate time series data is challenging due to complex interactions, high dimensionality, and nonlinear dependencies among variables. Existing methods often struggle to capture these complexities, resulting… 30 arXiv — NLP / Computation & Language research 1mo ago (Towards) Scalable Reliable Automated Evaluation with Large Language Models arXiv:2607.28282v1 Announce Type: new Abstract: Evaluating the quality and relevance of textual outputs from Large Language Models (LLMs) remains challenging and resource-intensive. Existing automated metrics often fail to capture the complexity and variability inherent in… 10 arXiv — NLP / Computation & Language research 1mo ago Beyond a Single Judge: Simulating Social Persona Panels for Generative UI Evaluation arXiv:2607.28439v1 Announce Type: new Abstract: Generative UI (GenUI) lets large language models synthesize a complete, renderable interface directly from a natural-language instruction, but evaluating the quality of what they generate remains an open problem. Human evaluation… 15 arXiv — NLP / Computation & Language research 1mo ago Change2Task: From Repository Changes to Executable Coding Agent Tasks and Environments arXiv:2607.28591v1 Announce Type: cross Abstract: Scaling coding agents requires a continuing supply of executable data for training, benchmarking, and continuous evaluation. Each task must couple a realistic software state with a specification, development tools, and reliable… 15 arXiv — NLP / Computation & Language research 1mo ago OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models arXiv:2607.28609v1 Announce Type: cross Abstract: Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA evaluation,… 17 Hugging Face Daily Papers research 1mo ago LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger Abstract Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy. This aggregate signal cannot tell whether a correct… 37 Hugging Face Daily Papers research 1mo ago Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation Abstract Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable evaluation essential for measuring capability,… 24 Simon Willison community 1mo ago Investigating three real-world incidents in our cybersecurity evaluations Investigating three real-world incidents in our cybersecurity evaluations It happened again! This is turning into something of a pattern. Last week OpenAI accidentally exploited Hugging Face when one of their frontier models broke out of a sandboxed container and hacked into… 10 TechCrunch — AI news-outlet 1mo ago Dili raises $21.7M to bring AI compliance to the infrastructure boom The Series A was led by Khosla Ventures, with participation from Allianz, Rebel Fund, Brick and Mortar Ventures’ Darren Bechtel, and Y Combinator’s Garry Tan. 13 Hugging Face Daily Papers research 1mo ago Can AI agents conduct open-ended AI research? Early evidence from two case studies Abstract Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit… 32 Hugging Face Daily Papers research 1mo ago CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition Abstract Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has highlighted this capability as context learning, existing evaluations mainly focus on textual contexts. In many practical… 14 arXiv — Machine Learning research 1mo ago SCOUT: Per-Context Reset Curricula for Sparse-Reward Reinforcement Learning arXiv:2607.26417v1 Announce Type: new Abstract: Sparse-reward reinforcement learning often fails because rollouts from the unassisted evaluation start rarely reach later task stages. Reset curricula address this by starting some training rollouts from easier intermediate states,… 27 arXiv — Machine Learning research 1mo ago From Unsupervised Subgroups to Hypothetical State-Intervention Policies: An Evaluation of Selected Subgrouping Methods in Observational Health Data arXiv:2607.26521v1 Announce Type: new Abstract: Conventional subgroup analyses can yield unstable and difficult-to-interpret conclusions, especially in observational biomedical data where each individual is observed under only one exposure state, true individual treatment… 4 arXiv — Machine Learning research 1mo ago Enhancing Automated Machine Learning via Homogeneous Train-Test Splitting Methods arXiv:2607.26625v1 Announce Type: new Abstract: Accurate model evaluation in machine learning depends critically on how datasets are split into training and testing subsets. Standard random splitting assumes that both partitions share the same underlying distribution, an… 18 arXiv — Machine Learning research 1mo ago Foundation Models for Face Presentation Attack Detection: A Unified Linear-Probing Benchmark arXiv:2607.26993v1 Announce Type: new Abstract: Face presentation attack detection (PAD) remains challenging under cross-dataset evaluation, where domain shift degrades models trained on a single dataset. The scarcity of large-scale labeled data motivates adapting pretrained… 22 arXiv — Machine Learning research 1mo ago BayesAME: Bayesian Active Model Evaluation arXiv:2607.27023v1 Announce Type: new Abstract: Evaluating large generative models across benchmarks is time-consuming and computationally expensive. This drives the need for methods that can estimate full benchmark performance by evaluating models on only a subset of items,… 38 arXiv — Machine Learning research 1mo ago Inverse Learning of Latent Risk-Neutral Densities from Irregular Option Quotes arXiv:2607.27188v1 Announce Type: new Abstract: Accurate option prices do not imply accurate recovery of the latent risk-neutral density. We study this distinction with two complementary benchmarks. A controlled benchmark exposes simulator-truth densities for latent evaluation,… 30 arXiv — Machine Learning research 1mo ago When benchmark inferences do not compose: Projectibility in AI evaluation arXiv:2607.26159v1 Announce Type: cross Abstract: An AI benchmark result rarely reaches a consequential claim in one step. Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new tasks, transport it to another system or site, and… 34 arXiv — NLP / Computation & Language research 1mo ago Position: Evaluation Scores Are Perishable Knowledge Claims arXiv:2607.26191v1 Announce Type: cross Abstract: Evaluation methodologies for language models increasingly combine multiple signals, from automated metrics and LLM-as-judge ratings to human assessments and benchmark suite results. When these signals are aggregated via… 10 Page 8 of 10 · 500 articles ← Newer Older →