News / #funding Tag Funding 500 articles archived under #funding · RSS Sign in to follow r/MachineLearning community 13h ago You can beat SOTA Time Series Anomaly Detection methods with a 100 year old algorithm [R] You can beat SOTA Time Series Anomaly Detection methods with a 100 year old algorithm Time Series Anomaly Detection (TSAD) seems to be one of the hottest topics in NeurIPS, SIGKDD, VLDB etc. Many (perhaps most) papers evaluate on Paparrizos’ TSB-AD-M benchmark… However, I tested… 4 The Information — AI news-outlet 1d ago Andreessen Horowitz Raises $1.1 Billion for AI Hardware Fund Andreessen Horowitz raised $1.1 billion for a fund that will invest in physical AI and infrastructure startups, the firm announced Friday. The money will be used to back companies building chips, memory, networking, storage and other companies involved in the AI buildout.… 15 TechCrunch — AI news-outlet 1d ago Neocloud Lambda secures $1B in debt to buy more chips Neocloud Lambda has raised $1B in private debt to buy Nvidia AI chips and lease them to Microsoft. It's the latest in a string of loans, underscoring the high cost of the AI boom. 17 Hugging Face Daily Papers research 1d ago What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference in Inspect Evals Abstract Evaluation artifacts specify a forward computation: a task, scorer, and reported metric. They do not necessarily license the claim attached to that metric because the historical evidence and alternative semantics needed to replay it may be unbound. We formalize this… 25 arXiv — Machine Learning research 2d ago NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation arXiv:2608.26222v1 Announce Type: new Abstract: Safety evaluation is critical for assessing whether aligned Large Language Models (LLMs) remain robust against jailbreak attacks. Existing automated testing methods, however, largely rely on response-level feedback: each candidate… 12 arXiv — Machine Learning research 2d ago Tabular Deep Learning for Algorithmic Trading: Cross-Regime Bayesian Optimisation for Equity Signal Generation arXiv:2608.27076v1 Announce Type: new Abstract: Algorithmic trading now represents a market exceeding $20 billion, where even marginal gains in signal robustness can translate into economically significant returns. Existing evaluations of equity prediction models do not… 24 arXiv — Machine Learning research 2d ago TraceBench: Controlled Evaluation of LLM Agents for Time-Series Root-Cause Attribution arXiv:2608.27182v1 Announce Type: new Abstract: LLM agents are increasingly applied to anomaly detection and root-cause analysis in time-series observations collected from real-world systems; however, their performance on these tasks has not been systematically evaluated under… 10 arXiv — Machine Learning research 2d ago Profit based evaluation of machine learning for nitrogen recommendations in winter wheat arXiv:2608.27205v1 Announce Type: new Abstract: Nitrogen rates for winter wheat are set before the season, under unknown prices and weather. The standard UK advice does not respond to prices, yet recent price swings moved the most profitable rate by tens of kilograms per… 26 arXiv — NLP / Computation & Language research 2d ago Cross-Platform Generalisation Failure in Mental Health Natural Language Processing: A Five-Axis Fairness Audit of Transformer Models on Social Media arXiv:2608.26138v1 Announce Type: new Abstract: We introduce the Cross-Platform Fairness Evaluation (CPFE) framework -- a five-axis audit protocol covering discriminative performance, calibration, statistical significance, prediction equity, and attribution stability -- and… 34 arXiv — NLP / Computation & Language research 2d ago ElementCheck: Complexity-Aware Long-Form Text Factuality Evaluation via Sentence Elements arXiv:2608.26118v1 Announce Type: new Abstract: Existing long-form factuality evaluation relies on the decompose-retrieve-verify pipeline. However, the pipeline suffers from noise from claim decomposition and fixed verification granularity, resulting in unreliable results. We… 26 arXiv — NLP / Computation & Language research 2d ago Evaluating Language Models in Realistic Conversational Contexts arXiv:2608.26131v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed to serve open-ended, multi-turn interactions, evaluating conversational quality at human scale has become a central challenge. Existing evaluation frameworks built for… 20 arXiv — NLP / Computation & Language research 2d ago Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation arXiv:2608.26159v1 Announce Type: new Abstract: Self-Generated Text Recognition (SGTR)--the ability of an LLM to identify its own outputs--poses risks to AI safeguards that rely on LLMs as evaluators or monitors. Specifically, an LLM may recognize outputs from other copies of… 34 arXiv — NLP / Computation & Language research 2d ago MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models arXiv:2608.26295v1 Announce Type: new Abstract: Tool-augmented LLMs must arbitrate between two fallible sources when a tool return conflicts with their parametric memory, yet existing evaluations measure source preference without establishing source correctness. We introduce… 25 arXiv — NLP / Computation & Language research 2d ago Why RAGs Hallucinate: Penalty-Aware Evaluation of Retrieval-Augmented Generation Systems with Knowledge-Gap Canaries arXiv:2608.26385v1 Announce Type: new Abstract: Volume-based accuracy rewards retrieval-augmented generation (RAG) systems for guessing: a system that answers everything outscores one that declines when its knowledge base cannot support an answer. Building on the… 6 arXiv — NLP / Computation & Language research 2d ago LowRankArena: A Standardized Evaluation Platform for SVD-Based LLM Compression arXiv:2608.26389v1 Announce Type: new Abstract: SVD-based low-rank compression has become a fast-growing direction for reducing the memory and computational cost of large language models (LLMs). However, meaningful comparison across existing studies remains difficult as prior… 30 arXiv — NLP / Computation & Language research 2d ago Multi-Expert Conformal Risk Control for Pairwise LLM Judging in Open-Ended Dialogue arXiv:2608.26529v1 Announce Type: new Abstract: In this paper, we explore multi-expert Conformal Risk Control (CRC) algorithms for pairwise LLM-as-a-Judge evaluation in open-ended dialogue. Our core insight is that multi-expert aggregation offers a complementary remedy to CRC:… 11 arXiv — NLP / Computation & Language research 2d ago Which Metrics Save the Most Human Annotation? Prediction-Powered Evaluation and Meta-Evaluation arXiv:2608.26638v1 Announce Type: new Abstract: Across various non-verifiable tasks, human evaluation is reliable but expensive, while automatic metrics are more scalable but often biased. Building on prediction-powered inference (PPI), we propose prediction-powered evaluation,… 35 arXiv — NLP / Computation & Language research 2d ago Do LLMs Understand Personality? Rethinking Persona Fidelity Evaluation through Structured Behavioral Inference arXiv:2608.26674v1 Announce Type: new Abstract: As large language models are increasingly deployed to simulate diverse human characters, ensuring persona fidelity, defined as the extent to which an agent's behavior consistently reflects the psychological and stylistic… 6 arXiv — NLP / Computation & Language research 2d ago JudgeStealer: Extracting LLM Judging Capabilities across Evaluation Protocols arXiv:2608.26982v1 Announce Type: new Abstract: Large language model (LLM) judges are increasingly used across various evaluation scenarios, making their judgment capabilities valuable intellectual property. However, black-box access exposes these capabilities to model… 34 arXiv — NLP / Computation & Language research 2d ago When Text Misleads: Inconsistent-Aware Reasoning for Audio-Grounded Dialogue arXiv:2608.27176v1 Announce Type: new Abstract: Understanding spoken dialogue requires joint reasoning over lexical content and paralinguistic acoustic signals such as emotion and conversational intent. However, existing evaluations often allow shortcuts based on transcripts or… 32 Google DeepMind official-blog 2d ago Piloting the world's first double-blind AI evaluations Piloting the world's first double-blind AI evaluations 22 Hugging Face Daily Papers research 3d ago FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling Abstract FIRM-Video uses checklist-driven verification of temporal visual evidence to build reliable reward models for text-to-video evaluation and alignment. Generated by thinkingmachines/Inkling-Small Reliable reward models are essential for text-to-video evaluation and… 32 arXiv — Machine Learning research 3d ago Why and When Neural Networks Improve Local Approximation in Optimization arXiv:2608.24963v1 Announce Type: new Abstract: Published experience with neural surrogates in derivative-free optimisation is contradictory: the same family of models that cuts the evaluation count of one solver leaves another unchanged, or makes it worse. We show that the… 17 arXiv — Machine Learning research 3d ago Flower Hub: A Reproducible Benchmarking Platform for Federated Learning in Simulation and Deployment arXiv:2608.25114v1 Announce Type: new Abstract: Federated learning (FL) has emerged as a key approach for training models across decentralized data, yet benchmarking in FL remains difficult to reproduce, compare, and extend. Existing evaluations are often tied to custom… 7 arXiv — NLP / Computation & Language research 3d ago LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale arXiv:2608.25204v1 Announce Type: cross Abstract: We introduce LibriBrain100, a large-scale MEG dataset for speech decoding designed from the ground up for reproducible, standardised evaluation. LibriBrain100 more than doubles the size of the original LibriBrain release,… 15 arXiv — Machine Learning research 3d ago Long-Term Behavioral Evaluation for Trusted Collaborator Selection via Bidirectional Mamba arXiv:2608.25232v1 Announce Type: new Abstract: Effective selection of trustworthy collaborators is crucial to ensuring the successful completion of collaborative tasks, which requires accurate assessments of both long-term device behavior and short-term collaborative dynamics.… 16 arXiv — Machine Learning research 3d ago Modeling spatio-temporal locality in multi-step forecasting of geo-referenced time series arXiv:2608.25698v1 Announce Type: new Abstract: Forecasting future measurements from geographically distributed sensors is essential across many domains. However, the spatial distribution of these sensors raises multiple challenges, primarily due to spatial autocorrelation… 19 arXiv — Machine Learning research 3d ago Towards A Unified Information Bottleneck Framework for Time Series Explanations arXiv:2608.25897v1 Announce Type: new Abstract: Explaining deep learning models operating on time series data is crucial in various applications that require transparent and interpretable insights into model behavior. {Existing explanation methods generally fall into two… 12 arXiv — NLP / Computation & Language research 3d ago MTDiag: A Multi-Turn Diagnostic Dataset Towards Clinically Meaningful LLM Evaluation arXiv:2608.25085v1 Announce Type: new Abstract: Clinical diagnosis is fundamentally interactive and incremental, yet the dominant paradigm for evaluating Large Language Models (LLMs) in medicine remains static QA benchmarks or template-based dialogues. These benchmarks say… 38 arXiv — NLP / Computation & Language research 3d ago Apples to Apples? Towards Comparable Crosslingual Language Model Evaluation arXiv:2608.25089v1 Announce Type: new Abstract: Crosslingual evaluation of language models that enables fair comparisons remains a fundamental challenge in multilingual NLP. Existing studies adopt a variety of downstream tasks and intrinsic metrics with different theoretical… 17 arXiv — NLP / Computation & Language research 3d ago MathAdv: What Theorem Provers Know, Reason, Formalize, and Generalize arXiv:2608.25449v1 Announce Type: new Abstract: Formal theorem proving enables machine-verifiable evaluation of mathematical reasoning, yet existing benchmarks often emphasize aggregate proof accuracy, concentrate on a narrow range of mathematics, and provide limited evidence of… 5 arXiv — NLP / Computation & Language research 3d ago Generative vs. Encoder Large Language Models for ASR Evaluation: A Comparative Study arXiv:2608.25574v1 Announce Type: new Abstract: Automatic Speech Recognition (ASR) is typically evaluated using Word Error Rate (WER), which poorly reflects semantic similarity. While embedding-based metrics correlate better with human judgments, the respective roles of encoder… 23 arXiv — NLP / Computation & Language research 3d ago Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty arXiv:2608.25660v1 Announce Type: new Abstract: Automated novelty judgment can accelerate scientific discovery by enabling efficient evaluation, refinement, and comparison of research ideas. While large language models are increasingly adopted for this task, we investigate a… 17 arXiv — NLP / Computation & Language research 3d ago Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence arXiv:2608.25869v1 Announce Type: new Abstract: Large language models (LLMs) increasingly assess generated content, giving rise to the LLM-as-a-Judge paradigm. These systems now score outputs, filter content, and gate iterative refinement in production pipelines, where each… 21 arXiv — NLP / Computation & Language research 3d ago Query-Side Attacks on GNN-Based KGQA: Tracing Failures from Entity Linking to Answer Generation arXiv:2608.25922v1 Announce Type: new Abstract: GNN-based Knowledge Graph Question Answering (KGQA) pipelines process queries through four discrete stages: entity linking, subgraph retrieval, GNN reasoning, and answer generation. Standard robustness evaluations conflate… 15 arXiv — NLP / Computation & Language research 3d ago TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue arXiv:2608.25218v1 Announce Type: cross Abstract: Speakers in natural conversation take turns speaking and listening, deciding in real time when to take, hold, or yield the floor. However, turn-taking evaluation remains limited due to the lack of a consistent, linguistically… 14 arXiv — NLP / Computation & Language research 3d ago The "Curse of Knowledge" in LLM Query Simulation: Concept Provenance for Tracing Answer-Side Intrusion arXiv:2608.25245v1 Announce Type: cross Abstract: LLM-generated search queries are widely used to augment IR evaluation, yet they may contain concepts that presuppose answer-side document knowledge, violating the information-access boundary of pre-search users. Existing… 25 arXiv — NLP / Computation & Language research 3d ago FinRiskAtlas: Decision-Aligned Evaluation of Large Language Models for Financial Risk Review arXiv:2608.25325v1 Announce Type: cross Abstract: Deploying large language models for professional financial review requires more than measuring general financial competence: models must perform the specific review operation required by a workflow and determine whether available… 38 The Information — AI news-outlet 3d ago SoftBank in Talks to Buy Majority Stake in Humanoid Maker 1X at $6 Billion Valuation SoftBank is in talks to buy a majority stake in 1X Technologies, an OpenAI-backed humanoid robot developer, in a deal that would value the startup at about $6 billion, according to people with knowledge of the deal. The investment would buttress SoftBank’s robotics ambition. The… 21 TechCrunch — AI news-outlet 3d ago Viral AI startup Instinct has raised $350 million at a $2.5 billion valuation The startup is only a year old but it has already generated a massive amount of hype (and money) while also spurring privacy concerns. 30 The Information — AI news-outlet 3d ago Four-month-old AI Assistant Startup Instinct Raises at a $2.5 Billion Valuation Instinct, a startup whose AI assistant connects to users’ applications and performs tasks on their behalf, is raising a Series B at a $2.5 billion valuation in a round led by Index Ventures and Benchmark, said a person with direct knowledge. The round will close in the next… 7 r/MachineLearning community 3d ago A dataset with 52 Text to image model evaluation [P] I created a simple text to image benchmark. I curated 192 prompts that are difficult for T2I models in various ways: text rendering, spatial reasoning, human realism, negations, etc... I then asked a VLM to judge every output against a pre-specified binary question with the… 30 TechCrunch — AI news-outlet 3d ago Arga is building a better way to train enterprise AI agents Arga has raised $10 million in a seed funding round that was led by General Catalyst, with participation from Box Group, Emergence, Gradient and SV Angel. 5 The Information — AI news-outlet 4d ago OpenAI Says Its Jalapeño AI Chip Is Better Than Nvidia’s Blackwell OpenAI on Tuesday released an evaluation of its new Jalapeño AI server chip, claiming it is faster and more powerful than Nvidia’s flagship Blackwell chips. The results appear promising for OpenAI, which hopes to reduce its reliance on Nvidia’s hardware someday, but it’s still… 24 arXiv — Machine Learning research 4d ago ChorusTIC: Training-Free Multivariate Time Series Classification via Chorus In-Context Learning arXiv:2608.24033v1 Announce Type: new Abstract: Time series classification underpins applications in healthcare, sensing, and industrial monitoring. Although time series foundation models support forecasting and transferable representation learning, classification still… 28 arXiv — Machine Learning research 4d ago Structured Frequency-Domain Evidence for LLM-Based Time-Series Anomaly Detection arXiv:2608.24113v1 Announce Type: new Abstract: Time-series anomalies can appear not only as pointwise deviations but also as changes in recurring temporal structure, such as shifted periodicity or localized oscillatory fluctuations. However, existing LLM-based time-series… 23 arXiv — Machine Learning research 4d ago Robust Data-Collection Policy Learning for Low-Variance Online Policy Evaluation arXiv:2608.24146v1 Announce Type: new Abstract: In reinforcement learning policy evaluation, classic on-policy methods often suffer from high variance when estimating policy performance. To mitigate this issue, behavior policy search has been proposed to learn data-collecting… 12 arXiv — Machine Learning research 4d ago Evaluating Deep Multivariate Imputation Models on Wearable Device Data arXiv:2608.24436v1 Announce Type: new Abstract: Wearable device data enables continuous health monitoring, but suffers from structured missingness: features sharing a physical sensor drop out together. Deep imputation methods such as BRITS and SAITS have seen limited evaluation… 33 arXiv — Machine Learning research 4d ago From Numerical Simulators of PDEs to Neural Emulators and Back arXiv:2608.24547v1 Announce Type: new Abstract: Simulation is central to modern engineering and science, but the cost of numerical solvers for partial differential equations (PDEs) remains a bottleneck whenever fast or many-query evaluations are required. Neural emulators… 28 arXiv — Machine Learning research 4d ago Data Leakage Inflates Generalizability of Power Outage Prediction Models arXiv:2608.24665v1 Announce Type: new Abstract: Power outage prediction models are increasingly used in assessments of climate-driven infrastructure risk, yet current evaluation practices obscure whether these models generalize to the novel conditions such applications require.… 6 Page 1 of 10 · 500 articles Older →