News / #funding Tag Funding 500 articles archived under #funding · RSS Sign in to follow arXiv — NLP / Computation & Language research 12d ago When AI Rewrites, Classifiers Relax: Uncertainty-Aware Sentiment Analysis on Sarcastic and AI-Paraphrased Social Text arXiv:2608.15338v1 Announce Type: new Abstract: Sentiment classifiers are increasingly applied to social media content that is either sarcastic or AI-generated --- two distributional regimes where standard evaluations offer little guidance. We present a three-part empirical… 30 arXiv — NLP / Computation & Language research 12d ago A Cognitively Motivated Multidimensional Framework for Evaluating Metaphor Explanations arXiv:2608.15828v1 Announce Type: new Abstract: Current evaluation of metaphor explanations relies mainly on holistic quality ratings, revealing little about how explanation quality is structured or where human judgments agree and diverge. We introduce a cognitively motivated… 22 arXiv — NLP / Computation & Language research 12d ago PLSQLBench: Benchmarking LLM Systems for Executable Procedural Database Programming arXiv:2608.15931v1 Announce Type: new Abstract: We present PLSQLBench, to our knowledge the first benchmark for evaluating whether LLMs can write executable PL/SQL programs, with correctness measured through execution-based tests. Existing LLM evaluations largely target… 26 arXiv — NLP / Computation & Language research 12d ago Clause Encounters of the Third Kind: Can LLMs Replace Language Teachers? arXiv:2608.16286v1 Announce Type: new Abstract: While various organizations now actively encourage LLM use in classrooms, we still lack rigorous, systematic evaluations of how well these models actually perform the fundamental tasks of language pedagogy. This paper examines… 13 arXiv — NLP / Computation & Language research 12d ago IndicQE-APE: A Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages arXiv:2608.16344v1 Announce Type: new Abstract: Indic quality estimation (QE) and automatic post-editing (APE) data is spread across separate releases, so no single resource supports training and evaluation across tasks and language pairs on one footing. We consolidate the WMT… 18 arXiv — NLP / Computation & Language research 12d ago Unadapted Multilingual ASR on a Garrusi Kurdish Evaluation Set: A Common-Reference Staged Normalization Analysis arXiv:2608.16379v1 Announce Type: new Abstract: Evaluating speech recognition for a Kurdish variety written in a Latin field orthography, using a model that outputs Arabic script, creates a measurement problem before a modelling one: direct scoring treats writing-system… 38 arXiv — NLP / Computation & Language research 12d ago Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents arXiv:2608.14606v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as synthetic survey respondents, but existing evaluations ask whether answers look plausible at the individual level. We argue the right question is psychometric: do LLMs… 30 Hugging Face Daily Papers research 12d ago HarnessEval-W: Agentifying the Evaluation of Visual Worlds Abstract HarnessEval-W uses hierarchical sub-agents to decompose world-model evaluations into verifiable reasoning chains that justify scores with transparent evidence. Generated by thinkingmachines/Inkling-Small A benchmark should deliver more than a scalar score: what makes an… 29 TechCrunch — AI news-outlet 12d ago Groq raises $350M to fuel its pivot from AI chips to neocloud Groq raised $350 million at a $3.5 billion valuation as the former AI chipmaker pivots to a neocloud business and expands its Nvidia-powered data center footprint. 33 TechCrunch — AI news-outlet 12d ago Wispr raises $280M at $2B valuation as it looks beyond dictation Wispr’s total funding is now over $361 million. 29 Hugging Face Daily Papers research 13d ago PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment Abstract PRM-as-a-Judge 1.5 provides fine-grained process metrics and reliability tools to evaluate embodied robotic models beyond binary success rates. Generated by thinkingmachines/Inkling-Small Fine-grained robotic evaluation matters for understanding embodied models, going… 33 arXiv — Machine Learning research 13d ago Don't Claim Benchmark-Oriented Optimization Improves General Coding Capability -- Diverse Evaluation Is Required arXiv:2608.13566v1 Announce Type: new Abstract: Post-training papers, model cards, and blog posts often treat scores on a small set of coding benchmarks (e.g., SWE-bench and LiveCodeBench) as evidence of broad coding capability, both for research artifacts and user-facing… 22 arXiv — Machine Learning research 13d ago Model-agnostic Retrieval-Augmented Extended Forecasting for time series arXiv:2608.14054v1 Announce Type: new Abstract: Time series forecasting with pretrained foundation models has demonstrated strong zero-shot capabilities. However, achieving optimal performance on time series with short or negligible historical data in domain-specific… 15 arXiv — Machine Learning research 13d ago Multi-Objective Bayesian Optimization for Model Merging arXiv:2608.14264v1 Announce Type: new Abstract: Model merging combines trained models directly in weight space, offering a compute-efficient alternative to additional fine-tuning. Selecting merge parameters is nevertheless difficult because downstream evaluations are expensive,… 9 arXiv — Machine Learning research 13d ago AI Evaluation Should Work With Humans arXiv:2608.13577v1 Announce Type: cross Abstract: This position paper argues that the dominant paradigm of AI evaluation (which focuses on superhuman autonomous performance and so implicitly targets the goal of replacing humans) is guiding AI development in the wrong direction.… 10 arXiv — NLP / Computation & Language research 13d ago Measuring Fairness in Large Audio Language Models via Semantic-Aware Bias Estimation arXiv:2608.13624v1 Announce Type: new Abstract: Large Audio Language Models (LALMs) have seen increasing use for audio understanding tasks such as speech recognition and audio question answering, raising concerns about fairness across demographic subgroups. Fairness evaluation… 7 arXiv — NLP / Computation & Language research 13d ago When Lexical Change Misleads: Rethinking Dynamic Topic Model Evaluation with Traditional and LLM-Based Metrics arXiv:2608.13835v1 Announce Type: new Abstract: Dynamic topic models capture evolving word distributions, but traditional coherence metrics may fail when vocabulary changes while semantic meaning persists. We evaluate 120 topics from CoNTM and DLDA across NYT, DBLP, and arXiv,… 36 arXiv — NLP / Computation & Language research 13d ago Information Satisfaction: A Reader-Centered Axis for Summarization Evaluation arXiv:2608.14457v1 Announce Type: new Abstract: The majority of work on summarization evaluation focuses on general summary quality (e.g., ROUGE, BERTScore) or specific desired properties (e.g., readability, factuality). However, these metrics fail to measure the utility of a… 10 arXiv — NLP / Computation & Language research 13d ago Reading Between The Lines: Modeling and Evaluating Behavioral Realism in Legal Simulation arXiv:2608.13712v1 Announce Type: cross Abstract: Deposition training requires attorneys to manage dynamic witness behavior, yet legal-AI evaluations largely focus on factual accuracy, reasoning, or response-level plausibility. We introduce WitnessSim, a deposition simulator… 13 arXiv — NLP / Computation & Language research 13d ago Research-Oriented Human-Centric Evaluation for Foundation Models arXiv:2506.01793v2 Announce Type: replace Abstract: Most current evaluations of foundation models focus on objective benchmarks, such as knowledge coverage and reasoning accuracy, often overlooking users' subjective experiences in human-AI collaboration. To address this gap, we… 14 arXiv — NLP / Computation & Language research 13d ago Theory-Grounded Evaluation Exposes the Authorship Gap in LLM Personalization arXiv:2604.26460v2 Announce Type: replace Abstract: Stylistic personalization - making LLMs write in a specific individual's style, rather than merely adapting to task preferences - lacks evaluation grounded in authorship science. We show that grounding evaluation in authorship… 29 arXiv — NLP / Computation & Language research 13d ago Lost in Phonation: Voice Quality Variation as an Evaluation Dimension for Speech Foundation Models arXiv:2510.25577v2 Announce Type: replace-cross Abstract: Recent advances in Speech Foundation Models (SFMs) enable direct processing of raw audio, allowing models to respond to subtle paralinguistic variation. However, how these models interpret non-lexical cues remains largely… 6 Hugging Face Daily Papers research 13d ago Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Abstract A new benchmark for AI-generated video detection reveals that current detectors fail to generalize across realistic crisis-related videos and become less reliable as content spreads socially. Generated by thinkingmachines/Inkling-Small Recent video generators can… 27 Hugging Face Daily Papers research 13d ago Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development Abstract Frontier autonomous agents excel at engineering optimization but show unstable performance, limited novelty, and variable experience reuse across long-horizon tasks. Generated by thinkingmachines/Inkling-Small Autonomous agents are increasingly capable of improving… 14 arXiv — Machine Learning research 16d ago Unifying Generative Models with Path Integrals arXiv:2608.12438v1 Announce Type: new Abstract: We formulate generative modeling as a path integral in which flow-based, diffusion-based, variational, and adversarial models arise as different evaluation principles for a single master action. Its… 36 arXiv — Machine Learning research 16d ago When Can You Trust Offline Evaluation of Equal-Cost Top-k Allocation? A Controlled, Reproducible Benchmark and Practitioner's Guide arXiv:2608.12489v1 Announce Type: new Abstract: Organizations decide whom to treat under a budget and want to know what a targeting rule would have earned before deploying it. Off-policy evaluation promises this from logged data, but the deployable rule is a deterministic top-k… 18 arXiv — Machine Learning research 16d ago GENADA: efficient generative time series adversarial attack framework arXiv:2608.12535v1 Announce Type: new Abstract: Deep learning models are widely used for time series analysis in domains such as healthcare, finance, energy systems, and environmental monitoring. However, these models remain vulnerable to adversarial attacks, where small input… 17 arXiv — Machine Learning research 16d ago H-VAEP and H-xT: Valuing Offensive On-the-Ball Actions in Handball by Estimating Probabilities arXiv:2608.12926v1 Announce Type: new Abstract: Traditional player evaluation in professional handball relies on basic box-score metrics or heuristic indices, which fail to credit the multi-player build-up chain. While football (soccer) analytics has adopted Expected Threat (xT)… 31 arXiv — Machine Learning research 16d ago Incremental Evaluation and Training in Relational Deep Learning arXiv:2608.13023v1 Announce Type: new Abstract: Relational Deep Learning (RDL) models multi-tabular databases as temporal heterogeneous graphs to enable end-to-end representation learning. However, prevailing RDL evaluation practices rely on static, single-episode dataset… 24 arXiv — Machine Learning research 16d ago A Probe Direction Is a Property of Its Prompt arXiv:2608.13329v1 Announce Type: new Abstract: A model that behaves differently when it senses it is being tested would undermine the evaluations we rely on, so recent work has sought to read that sense directly from a model's activations. The standard instrument contrasts… 8 arXiv — Machine Learning research 16d ago Where You Measure Decides What You Measure: Position Selection in Ablation-Based SAE Evaluation arXiv:2608.13337v1 Announce Type: new Abstract: Sparse autoencoders are meant to name the things a language model computes, and the usual way to check that a latent matters is to switch it off and see what changes. But a latent fires at many tokens, and the effect has to be… 14 arXiv — NLP / Computation & Language research 16d ago On Measuring Semantic Preservation in Legal Ontology Learning arXiv:2608.12326v1 Announce Type: new Abstract: Ontology learning transforms unstructured text into structured representations for automated reasoning. Yet structuring information risks losing it, and current evaluation methodologies cannot detect such loss, focusing on… 26 arXiv — NLP / Computation & Language research 16d ago AnchorSIPS: A Synthetic Dataset and Evaluation Resource for Evidence-Supported Psychosis-Risk Symptom Measurement arXiv:2608.12329v1 Announce Type: new Abstract: Progress on AI for psychosis-risk assessment is limited by a data-access bottleneck. Real clinical interviews are difficult to share because of privacy, governance, and consent constraints. We present AnchorSIPS, a synthetic… 32 arXiv — NLP / Computation & Language research 16d ago Query Timing Produces Opposite Positional Biases Between LLMs and Humans arXiv:2608.12387v1 Announce Type: new Abstract: Positional biases such as recency and primacy effects have been documented in large language models (LLMs), yet the underlying mechanism by which these models make their evaluations remains poorly understood. Both primacy and… 11 arXiv — NLP / Computation & Language research 16d ago BavGround: A Benchmark for Regional Cultural Grounding and Dialect Competence in Bavarian arXiv:2608.12894v1 Announce Type: new Abstract: Cultural evaluation of large language models (LLMs) often focuses on high-resource standard languages, leaving regional culture and dialect communities underrepresented. We introduce BavGround, a benchmark for evaluating Bavarian… 35 arXiv — NLP / Computation & Language research 16d ago How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures arXiv:2608.13267v1 Announce Type: new Abstract: Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty… 20 arXiv — NLP / Computation & Language research 16d ago Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation arXiv:2608.13326v1 Announce Type: new Abstract: LLM benchmark scores can be precise even when the observation protocol does not identify the behavioral property they are intended to measure. In a controlled, solver-grounded setting, we formalize a protocol-level identifiability… 20 Hugging Face Daily Papers research 16d ago How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review Abstract Rhetorical framing significantly biases AI scientific review scores in structured ways, with effects shaped by reviewer identity, score range, and evaluation strictness rather than rewriting complexity. Generated by thinkingmachines/Inkling-Small As large language… 4 TechCrunch — AI news-outlet 16d ago Databricks wanted to raise $1B, investors wanted $15B. It settled on $5B at a $190B valuation. AI is expensive, Ali Ghodsi tells TechCrunch. With so many investors wanting into his latest round, he said yes to more than planned. 18 arXiv — Machine Learning research 17d ago Long-Horizon Forecasting of Complete Financial Statements with Forma arXiv:2608.11327v1 Announce Type: new Abstract: Specialist training beats generalist scale when forecasting financial statements. To our knowledge, no prior work jointly forecasts complete financial statements beyond one year, yet in a discounted-cash-flow valuation most firm… 6 arXiv — Machine Learning research 17d ago Unmasking Toxic Mimicry in Medical Offline Reinforcement Learning for ICU Sepsis Management via Counterfactual Clinical Audits arXiv:2608.11410v1 Announce Type: new Abstract: Offline reinforcement learning (RL) offers considerable promise for optimizing ICU treatment decisions, yet standard evaluation metrics Mean Squared Error (MSE) and Fitted Q-Evaluation (FQE) assess only behavioral imitation and… 37 arXiv — Machine Learning research 17d ago Analysis of Federated Aggregation under Model Poisoning and Backdoor Attacks: A Reconstructed Cross-Dataset and Cross-Architecture Benchmark arXiv:2608.11423v1 Announce Type: new Abstract: Robust comparisons of federated aggregation methods require joint consideration of predictive performance, threat definitions, metric semantics, and execution provenance. A 500-cell seed-1 evaluation matrix was reconstructed across… 5 arXiv — Machine Learning research 17d ago When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection in Delayed-Feedback Contextual Bandits arXiv:2608.11560v1 Announce Type: new Abstract: Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online… 4 arXiv — Machine Learning research 17d ago Robust and Efficient Noisy-Label Time-Series Classification via Dynamic Time Warping Based Granular Ball Computing arXiv:2608.11704v1 Announce Type: new Abstract: Dynamic Time Warping (DTW)-based Nearest-Neighbor (NN) classifiers are effective for time-series classification but are vulnerable to mislabeled training samples and require numerous DTW computations during inference. We propose… 6 arXiv — Machine Learning research 17d ago JAPE: Joint Anomaly Prediction and Intrinsic Explanation in Multivariate Time Series arXiv:2608.11801v1 Announce Type: new Abstract: Multivariate time-series anomaly prediction aims to identify whether and when anomalies will occur over a future horizon from historical observations. Existing methods primarily characterize anomalies as deviations in future… 18 arXiv — Machine Learning research 17d ago Towards Truly Unsupervised Evaluation of Feature Selection arXiv:2608.12057v1 Announce Type: new Abstract: Feature selection is one of the most important and fundamental tasks in data mining, tackled by a family of methods with an established set of evaluation techniques to measure the quality of a specific method. Most of the methods… 9 arXiv — NLP / Computation & Language research 17d ago TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation arXiv:2608.11236v1 Announce Type: new Abstract: Roleplay evaluation should do more than assign a single score: it should reveal which role requirements were tested, which failed, and which dialogue evidence supports the judgment. We propose TRACE Bench, a task-driven agentic… 28 arXiv — NLP / Computation & Language research 17d ago Group Alignment-Induced Sycophancy: A Two-Sided Evaluation of Steerable Pluralistic Alignment arXiv:2608.11528v1 Announce Type: new Abstract: Group alignment adapts a language model to a demographic group to produce responses that reflect the group's opinions, values, and preferences. Sycophancy, a well-documented by-product of alignment, causes the model to over-agree… 23 arXiv — NLP / Computation & Language research 17d ago LabelFusion-TS: Fusing Large Language Models, Transformer Encoders, and Financial Time Series for Monetary-Policy Stance Classification arXiv:2608.11753v1 Announce Type: new Abstract: Financial text is produced and interpreted within a market environment, yet financial text classifiers almost always receive text alone. We study whether financial time series are useful as an additional input on the task of… 33 arXiv — NLP / Computation & Language research 17d ago GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation arXiv:2608.11787v1 Announce Type: new Abstract: Generating actionable financial advice from business records demands that models integrate numerical reasoning, domain knowledge, and sound judgment, while avoiding recommendations that could harm the business. Direct supervision… 6 Page 4 of 10 · 500 articles ← Newer Older →