News / #funding Tag Funding 500 articles archived under #funding · RSS Sign in to follow arXiv — Machine Learning research 24d ago CAMP: A Cycle-Aware Multi-Scale Patch Mixer for Time Series Forecasting arXiv:2608.04051v1 Announce Type: new Abstract: Real-world time series are often governed by recurring patterns, but their dominant periods may vary across datasets, forecasting settings, and individual input windows. Existing cycle-aware forecasters commonly rely on a single… 35 arXiv — Machine Learning research 24d ago MINT: Tensor Decomposition on Stacked Recurrence Matrices for Time Series Data Mining arXiv:2608.04157v1 Announce Type: new Abstract: Recurrence plots are a time series data mining primitive applied to a variety of domains (e.g. star light curves, sound waveforms, CCT telemetry). This work proposes tensorized self-similarity matrices as a primitive for univariate… 6 arXiv — Machine Learning research 24d ago TS2TabPFN: Time Series Classification and Extrinsic Regression through Feature Extraction and a Tabular Foundation Model arXiv:2608.04174v1 Announce Type: new Abstract: Time series data are ubiquitous in practical applications, where classification (TSC) and extrinsic regression (TSER) have emerged as essential tasks for obtaining value from temporal sequences. While the literature has seen… 6 arXiv — Machine Learning research 24d ago Cost-Aware Multi-Objective Bandits: Theory and Application to Budgeted LLM Configuration Evaluation arXiv:2608.04333v1 Announce Type: new Abstract: Large language model (LLM) configuration evaluation is challenging due to limited evaluation budgets, varying costs, and multiple competing objectives. In this paper, we formulate LLM configuration evaluation as a cost-aware… 7 arXiv — Machine Learning research 24d ago EvtGraph: Event-Adaptive Compression for Sparse Temporal Graph Learning in Multimodal Time Series arXiv:2608.04368v1 Announce Type: new Abstract: Multimodal temporal data are inherently irregular and uneven in information density, yet most models rely on uniform discretization, leading to inefficient representations. We propose \textbf{EvtGraph}, a unified framework that… 37 arXiv — Machine Learning research 24d ago Active Learning Guided Design Space Refinement for Scalable Multi-Objective Bayesian Optimization in Materials Discovery arXiv:2608.04651v1 Announce Type: new Abstract: Advanced materials discovery increasingly relies on machine learning and Bayesian optimization to explore large discrete design spaces under limited evaluation budgets. However, conventional Bayesian optimization (BO) can become… 6 arXiv — Machine Learning research 24d ago Benchmarking Deep Learning Models for Dense Event Classification of Offshore Wind Infrastructure in Sentinel-1 Time Series arXiv:2608.04706v1 Announce Type: new Abstract: Monitoring of offshore wind energy infrastructure life cycles, especially during the deployment phase, is an important contribution for stakeholders to make informed decisions in a phase of increasing deployment activities. ESA's… 31 arXiv — Machine Learning research 24d ago MGSB: Manifold Gated Signature Branch Pressure-Domain Baseline Architecture for Two-Phase Pipeline Flows Under Distributional Shift arXiv:2608.04805v1 Announce Type: new Abstract: Leak detection models for multiphase pipelines often degrade when deployed under flow regimes that differ from training. Existing evaluations typically assess performance under in-distribution operating conditions, masking failures… 15 arXiv — NLP / Computation & Language research 24d ago Mind the Cap: Output-Budget Regimes Change the Measured Multilingual Reasoning Gap arXiv:2608.04160v1 Announce Type: new Abstract: Multilingual evaluations report accuracy at a single output-token cap, but languages need different numbers of tokens to express the same content, so the cap is a hidden experimental variable. We test whether the… 34 arXiv — NLP / Computation & Language research 24d ago Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary arXiv:2608.04240v1 Announce Type: new Abstract: Superhuman game engines in domains like chess have made expert-level evaluations easily accessible, yet they communicate what is true without the natural-language explanations that make such expertise educationally useful to… 17 arXiv — NLP / Computation & Language research 24d ago Towards End-to-End Multilingual Metaphor Processing: Integrating Detection, Translation, and Evaluation arXiv:2608.04260v1 Announce Type: new Abstract: Metaphorical language remains a major challenge for multilingual natural language processing because successful interpretation and translation require reasoning beyond literal lexical meaning. Existing research has largely… 14 arXiv — NLP / Computation & Language research 24d ago STRIVE: Probing Reasoning Limits in Graded Plausibility Generation and Evaluation arXiv:2608.04567v1 Announce Type: new Abstract: Event knowledge concerns who does what to whom. Psycholinguists use event-plausibility judgments to examine how this knowledge supports human language processing. To isolate plausibility effects, these studies require controlled… 21 arXiv — NLP / Computation & Language research 24d ago Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses? arXiv:2608.04828v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on skills, structured documents that specify when to act, which procedure to follow, and which tools are allowed. Existing evaluations mostly judge the quality of a skill or its… 5 arXiv — NLP / Computation & Language research 24d ago Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification arXiv:2608.04899v1 Announce Type: new Abstract: Confidence estimation is essential when LLMs are used for classification, indicating when predictions can be trusted. However, common approaches such as verbalization produce extremely sparse outputs. For instance, Qwen3-32B… 27 arXiv — NLP / Computation & Language research 24d ago Simile Understanding in Text-to-Image Models: An Evaluation Framework arXiv:2608.04750v1 Announce Type: cross Abstract: Similes provide a compact and expressive way to describe visual characteristics in text prompts. Recent text-to-image models (t2i models) can produce visually compelling outputs from simile prompts, yet even frontier models… 12 Simon Willison community 24d ago Third-party cyber evaluations involving OpenAI models Third-party cyber evaluations involving OpenAI models And another one . I had to create a accidental-cyberattacks tag to keep track of them all! This post from OpenAI covers both the UK AI Safety Institute attack (see my previous post ) and another attack enabled by Irregular :… 22 Simon Willison community 24d ago Third-party cyber evaluations involving OpenAI models Third-party cyber evaluations involving OpenAI models And another one . I had to create a accidental-cyberattacks tag to keep track of them all! This post from OpenAI covers both the UK AI Safety Institute attack (see my previous post ) and another attack enabled by Irregular :… 19 Simon Willison community 24d ago Incident Report: unsanctioned agent behaviour during cyber testing Incident Report: unsanctioned agent behaviour during cyber testing It happened again . This time it was the UK government's AI Security Institute who accidentally attacked other companies while running an evaluation with models with the safety filters turned off. From their… 37 Simon Willison community 24d ago Incident Report: unsanctioned agent behaviour during cyber testing Incident Report: unsanctioned agent behaviour during cyber testing It happened again . This time it was the UK government's AI Security Institute who accidentally attacked other companies while running an evaluation with models with the safety filters turned off. From their… 14 TechCrunch — AI news-outlet 25d ago AI makes weather prediction better. Can WindBorne make it lucrative? WindBorne Systems has raised $37 million Series B round to scale its weather balloons and AI forecasts. 9 Hugging Face Daily Papers research 25d ago Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Abstract We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open-web exploration. Preliminary evaluations reveal two critical bottlenecks in… 8 arXiv — Machine Learning research 25d ago Evaluation Blindness: How Silent Measurement Failures Corrupt AI Systems from Training to Deployment arXiv:2608.02786v1 Announce Type: new Abstract: AI systems can fail silently. The failure propagates through training loops, evaluation pipelines, and production monitoring stacks until downstream harm makes it visible. This paper introduces evaluation blindness: a measurement… 30 arXiv — Machine Learning research 25d ago Forecasting Revenue with its Customer-Base Drivers: When and Why Coordination Helps arXiv:2608.02911v1 Announce Type: new Abstract: Revenue forecasts guide acquisition budgets, demand planning, and customer-based valuations, yet an aggregate forecast does not show whether change reflects acquisition, repeat purchasing, spending per order, or offsetting… 8 arXiv — Machine Learning research 25d ago Paired Recipient-based Evaluation of Survival Prediction for Deceased Donor Kidney Transplants arXiv:2608.03017v1 Announce Type: new Abstract: There has been significant interest in using machine learning algorithms to predict kidney transplant outcomes, such as the number of years until a graft inevitably fails. These prediction algorithms could possibly be used for… 6 arXiv — Machine Learning research 25d ago FinVerse: Financial Time-Series Benchmark arXiv:2608.03259v1 Announce Type: new Abstract: As time-series foundation models have emerged, the need for benchmarks that can evaluate their forecasting ability in meaningful ways has become increasingly important. Existing time-series forecasting benchmarks provide useful… 14 arXiv — Machine Learning research 25d ago Noise-Aware Shrinkage for Differentially Private Zeroth-Order Fine-Tuning of Large Language Models arXiv:2608.03277v1 Announce Type: new Abstract: Differentially private zeroth-order optimization (DP-ZO) enables memory-efficient private fine-tuning of large language models using only forward evaluations. Existing aggregation-based DP-ZO methods reconstruct model updates at a… 23 arXiv — Machine Learning research 25d ago TimeRLM: Recursive Language Models Enable Precise Anomaly Localization in Long-Context Time-Series arXiv:2608.03391v1 Announce Type: new Abstract: Precise anomaly localization over long-context time series is a crucial task in monitoring applications across clinical care, industrial operations, financial services, and logistics, where brief evidence may hide inside long spans… 31 arXiv — Machine Learning research 25d ago Shorter Reasoning, Earlier Answers? An Evaluation of Reasoning Interfaces arXiv:2608.03401v1 Announce Type: new Abstract: Large language models often reason at length before answering, increasing cost and latency. Prompts and trained settings can shorten this reasoning, but a shorter trace may only show that the model stopped sooner. Here, we evaluate… 28 arXiv — Machine Learning research 25d ago PRISM: Powerful Time Series to Image (TS2I) Representations for Multivariate Anomaly Detection arXiv:2608.03926v1 Announce Type: new Abstract: Time series anomaly detection (TSAD) underpins applications in predictive maintenance, finance, and cloud computing, however performance remains sensitive to representation choices, especially in multivariate settings. While… 19 arXiv — NLP / Computation & Language research 25d ago Evaluating OpenAI's Privacy Filter: Cross-Lingual, Cross-Domain PII Detection Across 42 Benchmarks arXiv:2608.02616v1 Announce Type: new Abstract: We present the first independent, systematic evaluation of OpenAI's Privacy Filter (OPF), a 1.5B-parameter bidirectional PII detector, across 42 synthetic benchmarks spanning 22 languages and 5 domains. Zero-shot, OPF achieves… 34 arXiv — NLP / Computation & Language research 25d ago Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety arXiv:2608.02617v1 Announce Type: new Abstract: We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert feedback from MOOVE (Massive Open Online Validation and Evaluation), a… 15 arXiv — NLP / Computation & Language research 25d ago JudgeArena: A Unified Framework for Reproducible LLM-Judge Evaluation arXiv:2608.02620v1 Announce Type: new Abstract: LLM-as-a-judge evaluation has become a dominant paradigm for ranking language models, yet the ecosystem remains fragmented: most benchmarks ship their own code base, hardcode a specific closed-model judge, and support a single… 36 arXiv — NLP / Computation & Language research 25d ago TQLite: Multi-LLM Jury Guided Distillation for Real-time MQM Translation Quality Evaluation arXiv:2608.02975v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated impressive performance in MQM-based translation quality (TQ) evaluation, and recent advances in large reasoning models (LRMs) promise even greater improvements. However, both LLMs and… 22 arXiv — NLP / Computation & Language research 25d ago Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models arXiv:2608.03038v1 Announce Type: new Abstract: Statistical reasoning is multidimensional, yet evaluations of large language models (LLMs) typically emphasize response accuracy while overlooking how models construct and communicate statistical explanations. This study… 15 arXiv — NLP / Computation & Language research 25d ago Dynamically Allocating Evaluation Effort for Model Ranking arXiv:2608.03437v1 Announce Type: new Abstract: While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability. When identifying top-performing models, typical evaluation protocols waste effort by exhaustively evaluating all… 11 arXiv — NLP / Computation & Language research 25d ago Cross-Lingual Bias in Large Language Models: A Comparative Analysis of English and Swahili arXiv:2608.03532v1 Announce Type: new Abstract: Large language models are increasingly deployed in multilingual contexts, yet safety alignment and bias evaluation remain overwhelmingly English-centric. We investigate whether social biases generalise across languages by… 5 arXiv — NLP / Computation & Language research 25d ago How Closely Do LLM Reviews Align with Human Peer Review? arXiv:2608.03659v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers align with both conference decisions and human reviewing priorities within the… 18 arXiv — NLP / Computation & Language research 25d ago WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament arXiv:2608.04008v1 Announce Type: new Abstract: Benchmarks that measure the forecasting ability of large language models are almost always retrospective: the event has happened, the answer is somewhere on the Web, and the evaluation must defend itself against memorisation. We… 23 Hugging Face Daily Papers research 25d ago CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Abstract Captions serve as a primary supervision signal for both multimodal understanding and text-to-image generation. However, previous evaluations treat the caption quality as a single scalar objective, which conflates two distinct properties: (1) how much visual information… 22 r/LocalLLaMA community 25d ago Local LLM 35B MoE — Real-world coding benchmarks (Qwen vs Ornith vs KAT) I’ve been running a fairly opinionated evaluation loop on ~35B A3B/MoE-class models for coding over the past few months. Not synthetic benchmarks: actual dev workflows, iterative debugging, refactoring passes, and failure recovery. Here’s where things stand for me: Qwen 3.6 (35B… 9 Hugging Face Daily Papers research 26d ago Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures Abstract Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and which intervention would improve the agent system. This creates a repair-assignment problem: the same visible failure may call for model post-training,… 38 arXiv — Machine Learning research 26d ago Leak It: A Probabilistic Approach to Training-Data Extraction from Black-Box Language Models arXiv:2608.00144v1 Announce Type: new Abstract: Membership inference (MIA) on language models is usually summarised by an aggregate ROC-AUC, but such evaluations are confounded: model-free blind baselines separate members from non-members from surface text alone. We study… 13 arXiv — Machine Learning research 26d ago AutoCause: A Python framework that automates expert decisions in environmental time-series causal discovery arXiv:2608.00198v1 Announce Type: new Abstract: Environmental time-series causal discovery requires expert decisions about method choice, conditional-independence tests, lag horizons, sample-size adequacy, multiple-testing control, and evidence interpretation. Applied… 17 arXiv — Machine Learning research 26d ago An Embedded RISC-V Evaluation of Kolmogorov--Arnold Networks in Hard-Constrained Recurrent Physics-Informed Models arXiv:2608.00737v1 Announce Type: new Abstract: Hard-constrained recurrent physics-informed networks (HRPINNs) embed known dynamics inside a recurrent numerical integrator and restrict a neural branch to learning only the residual dynamics that the first-principles model does… 18 arXiv — Machine Learning research 26d ago AdvPlan-Bench: Adversarial Evaluation of Structured Plan-Generation Agents arXiv:2608.00832v1 Announce Type: new Abstract: Structured plan-generation agents are often evaluated as if a plan has quality in isolation, yet many realistic planning tasks require asking how a candidate behaves when another agent can search for responses. We introduce… 27 arXiv — Machine Learning research 26d ago UpliftBench: Revealing Outcome-Regime and Objective Mismatch in Uplift Evaluation arXiv:2608.00915v1 Announce Type: new Abstract: Uplift modeling (conditional-average-treatment-effect estimation) drives personalized targeting, yet published uplift benchmarks frequently disagree on which estimator performs best; we show the disagreement is substantially about… 33 arXiv — Machine Learning research 26d ago Who Belongs in the Eval Set? A Capability-Taxonomy-Driven Pipeline for Curating Regression Eval Sets in Agent-Extensibility Platforms arXiv:2608.01004v1 Announce Type: new Abstract: Platform teams hosting agent-extensibility surfaces face a regression-economics paradox: every onboarding customer ships an evaluation set tuned to their domain, but the platform's regression set must live under a hard query-count… 28 arXiv — Machine Learning research 26d ago When May a Model Replace the Experiment? Audits, Licenses, and the Price of Trust in Surrogate-Driven Design arXiv:2608.01378v1 Announce Type: new Abstract: Design campaigns in chemistry, materials science, and machine learning share a bottleneck: determining how good a candidate truly is requires an expensive evaluation - an experiment, a first-principles simulation, or a full… 20 arXiv — NLP / Computation & Language research 26d ago Averaging Bias: Human Faithfulness Annotations are not Locally Faithful arXiv:2608.00205v1 Announce Type: new Abstract: Evaluation of faithfulness of text summarization treats a model generated summary as faithful only if every of its sentences is supported by the source document: a strict conjunctive rule under which a single unsupported sentence… 29 arXiv — NLP / Computation & Language research 26d ago Deep Research Pretraining via Predictive Navigation arXiv:2608.00432v1 Announce Type: new Abstract: Deep research agents are often trained on expensive, environment-grounded tool-use trajectories that require repeated retrieval, document inspection, and report evaluation. We introduce Deep Research Pretraining (DRP), an offline… 8 Page 7 of 10 · 500 articles ← Newer Older →