News / #funding Tag Funding 500 articles archived under #funding · RSS Sign in to follow arXiv — NLP / Computation & Language research 9d ago HealMed: Multilingual Evaluation of Large Language Models in Medicine arXiv:2608.19981v1 Announce Type: new Abstract: We present HealMed, an expert-reviewed benchmark for multilingual evaluation of large language models in medicine. HealMed contains 1,000 examples in each of nine languages, drawn from nine datasets and covering three task formats:… 27 arXiv — NLP / Computation & Language research 9d ago OenoBench: A Wine-Domain Benchmark for Knowledge-Grounded Evaluation of Large Language Models arXiv:2608.20106v1 Announce Type: new Abstract: We introduce OenoBench, a wine-domain knowledge benchmark of 3,266 multiple-choice questions across six pillars (regions, grape varieties, viticulture, winemaking, producers, business) and four difficulty tiers. The corpus is built… 12 arXiv — NLP / Computation & Language research 9d ago Qworld: Question-Specific Evaluation Criteria for LLMs arXiv:2603.23522v2 Announce Type: replace Abstract: Evaluating large language models (LLMs) on open-ended questions is difficult because response quality depends on the question's context. Binary scores and static rubrics fail to capture these context-dependent requirements.… 37 r/LocalLLaMA community 9d ago Been tweaking my Qwen 3.8 setup, up to 45+ steady T/ps at 8bit quant. Realised I'm now top T/ps for this model+ctx across all benchmarked M-series chips. Full args linked below, happy to discuss as this was a pain of trial and error. https://omlx.ai/benchmarks/performance/2pko3m1k - you can expand the raw args, but I have full annotations of what worked and what didn't. I'm now testing the model on acutal coding and haven't seen any issues with performance vs default suggested vals for the vanilla model.… 35 arXiv — Machine Learning research 10d ago Multi-Class Electrical and Mechanical Fault Classification Using Random Convolutional Kernels arXiv:2608.18716v1 Announce Type: new Abstract: Diagnosing faults in rotating machinery is essential for ensuring the reliability of industrial processes. Random convolutional kernel-based Time Series Classification (TSC) methods, such as ROCKET and its variants, provide an… 7 arXiv — NLP / Computation & Language research 10d ago Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis arXiv:2608.18940v1 Announce Type: cross Abstract: Single-step retrosynthesis is a central component of computer-aided synthesis planning, yet its intrinsically one-to-many nature is poorly captured by single-answer evaluation and benchmarking protocols. To address this, we… 11 arXiv — Machine Learning research 10d ago Monroe: A Molecular Foundation Model for In-Context Probabilistic Inference arXiv:2608.18982v1 Announce Type: new Abstract: Bioassay activity prediction is often data-limited because drug-discovery datasets rely on time-consuming and expensive wet-lab experiments for data generation and evaluation. This challenge has inspired recent research into… 19 arXiv — Machine Learning research 10d ago Discretizing Continuous Time Series for Imputation with Masked Diffusion Training arXiv:2608.19119v1 Announce Type: new Abstract: Time series imputation is a crucial area for reliable time series analysis, yet it remains challenging due to the complex temporal dynamics and noise of real-world data. Existing approaches, however, exhibit two limitations:… 21 arXiv — Machine Learning research 10d ago L\'evy Attention: Single-Pass Predictive Uncertainty for Continuous-Time Attention arXiv:2608.19171v1 Announce Type: new Abstract: Deep models for irregularly-sampled time series answer queries at arbitrary continuous timestamps, yet report nothing about how far each answer should be trusted. We show the attention layer itself can close that gap: with the… 29 arXiv — NLP / Computation & Language research 10d ago MAVEN: A Macro-Societal Value Evaluation Framework of Multimodal Content with Compact Aligned Evaluators arXiv:2608.18096v1 Announce Type: new Abstract: Assessing whether multimodal content aligns with macro-societal values, such as peace, justice, and freedom, has become an increasingly urgent challenge. Existing frameworks are largely confined to safety-oriented taxonomies,… 38 arXiv — NLP / Computation & Language research 10d ago Self- and Other-Labels Induce Bidirectional Bias in LLM Judges arXiv:2608.18091v1 Announce Type: new Abstract: As LLM-as-a-judge systems become increasingly widespread, self-preference in LLMs -- the tendency to favor one's own outputs -- raises growing concerns about evaluation reliability. However, it has been studied predominantly on… 31 arXiv — NLP / Computation & Language research 10d ago Institutional Prestige as Geographic Bias in Large Language Models: Evidence from Three Factorial Experiments with Bootstrap Confidence Intervals arXiv:2608.18107v1 Announce Type: new Abstract: We investigate whether large language models (LLMs) systematically discriminate in candidate evaluations based on applicant name ethnicity and/or institutional prestige and geographic location. Three factorial experiments are… 23 arXiv — NLP / Computation & Language research 10d ago Are LLMs Safe Beyond Text: Do Emojis Expose Gaps in Safety Evaluation arXiv:2608.18164v1 Announce Type: new Abstract: Safety evaluations of large language models (LLMs) predominantly rely on text-based adversarial prompts, potentially overlooking vulnerabilities arising from alternative input representations. This work examines emoji-augmented… 9 arXiv — NLP / Computation & Language research 10d ago MissDiag: Diagnostic Evaluation of Incomplete-Knowledge Robustness in KGQA and KG-RAG arXiv:2608.18489v1 Announce Type: new Abstract: Knowledge graph question answering (KGQA) and knowledge-graph-based retrieval-augmented generation (KG-RAG) aim to ground answers in explicit graph evidence, but real-world knowledge graphs are often sparse, outdated, and… 28 arXiv — NLP / Computation & Language research 10d ago Execution-grounded evaluation reveals hidden failures in language-model calculations for environmental science arXiv:2608.18726v1 Announce Type: new Abstract: Large language models are increasingly used for quantitative work in the environmental sciences, yet existing evaluations score only final answers, leaving calculation process unobserved. Here we introduce AtmosCoder-Bench, an… 23 arXiv — NLP / Computation & Language research 10d ago Assessing Quality of Experience in Natural Language Generation of German Text arXiv:2608.18888v1 Announce Type: new Abstract: The rapid advancement of Natural Language Generation (NLG) has made the reliable evaluation of generated text increasingly critical, as these systems, such as large language models (LLMs), are now widely deployed in real-world… 8 arXiv — NLP / Computation & Language research 10d ago ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents arXiv:2608.18307v1 Announce Type: cross Abstract: Current evaluation of computer-use agents is split between long-horizon workflow benchmarks and atomic GUI-grounding tests. This leaves an under-instrumented middle layer: realistic component-centered interactions (e.g., toggle a… 26 arXiv — NLP / Computation & Language research 10d ago Building real-time digital twin instances with Function+Data Flow: user evaluation and extension for iterative pipelines arXiv:2608.18480v1 Announce Type: cross Abstract: Digital twins (DTs) increasingly leverage artificial intelligence (AI) and machine learning (ML) pipelines, both to build real-time DTs from high-fidelity simulations and to instantiate them with historical data. However,… 10 The Information — AI news-outlet 10d ago Nvidia Discusses Funding Its AI Data Supplier Mercor at a $20 Billion Valuation Nvidia has discussed an investment in Mercor, a data labeling provider that helps the chip designer develop its open-source AI models, according to a person with knowledge of the process. The investment would be part of a $20 billion-valuation round. Existing investor General… 13 arXiv — Machine Learning research 11d ago Position: Fairness Failure in Generative Models is an Evaluation Problem arXiv:2608.16974v1 Announce Type: new Abstract: Despite groundbreaking advancements in generative models during the last decade, concerns about their lack of fairness, reinforcing societal inequalities and harming marginalized groups, remain under-addressed and difficult to act… 14 arXiv — Machine Learning research 11d ago Iterative tensor network transformations for element-wise evaluation of elementary and filtering functions arXiv:2608.17135v1 Announce Type: new Abstract: Tensor networks are powerful formats for compressing large-scale data. However, their application to general data processing has been limited by the difficulty of performing nonlinear operations. Here, we introduce iterative tensor… 11 arXiv — Machine Learning research 11d ago Beyond MSE: Rethinking the Evaluation Metric and Benchmarking for Irregular Time Series Forecasting arXiv:2608.17293v1 Announce Type: new Abstract: Existing research on irregular time-series forecasting has primarily focused on model design, while evaluation metrics remain insufficiently studied. Existing benchmarks typically use mean squared error (MSE) as the evaluation… 30 arXiv — Machine Learning research 11d ago Evaluating RL Explainability Methods by How Much They Help Fix Bugs in Agents arXiv:2608.17524v1 Announce Type: new Abstract: This preliminary paper outlines a planned evaluation benchmark for Explainable Reinforcement Learning (XRL) methods. Current evaluations rely on functionally-grounded metrics like faithfulness and compactness, and on human-grounded… 31 arXiv — Machine Learning research 11d ago Cross-View Correspondence Is a Measurement Intervention: Two-Sided Validation for Agent Evaluation and Credit Assignment arXiv:2608.17713v1 Announce Type: new Abstract: Agent evaluations and trace-based learning often compare outputs across transformed views through a post-response correspondence treated as neutral preprocessing. We show that this correspondence is a measurement intervention:… 33 arXiv — Machine Learning research 11d ago Revisiting WEASEL 2.0: Reproduction, Sensitivity, and an Adaptive Ensemble-Size Rule arXiv:2608.18021v1 Announce Type: new Abstract: WEASEL 2.0 is a dictionary-based time series classifier that combines dilated sliding windows with a randomised hyperparameter ensemble and a fixed-size dense feature representation. Two of its hyperparameter choices, the maximum… 14 arXiv — Machine Learning research 11d ago SPSA Hyperparameter Tuning for Variational Quantum Natural Language Inference arXiv:2608.16939v1 Announce Type: cross Abstract: Training variational quantum models requires choosing between parameter-shift gradients, which are exact but cost $O(P)$ forward evaluations, and simultaneous perturbation stochastic approximation (SPSA), which uses only two… 7 arXiv — NLP / Computation & Language research 11d ago Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges arXiv:2608.17605v1 Announce Type: new Abstract: Conversational AI is moving beyond isolated text prompts toward sustained, multimodal interaction. In real conversations, users clarify goals, revise requests, interrupt responses, switch topics, and introduce new evidence while… 30 arXiv — NLP / Computation & Language research 11d ago Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses arXiv:2608.17810v1 Announce Type: new Abstract: The evaluation of large language models (LLMs) relies heavily on human-designed assessments, implicitly assuming that AI and humans employ similar underlying cognitive constructs. Challenging this assumption, we investigate whether… 18 arXiv — NLP / Computation & Language research 11d ago From Global Benchmarks to Local Evaluations: Benchmarking LLMs for the German Public Sector arXiv:2608.17827v1 Announce Type: new Abstract: Public institutions face a persistent challenge in selecting LLMs suited to their specific context. Existing benchmarks, however, are of limited use as they primarily reflect English-language and US-centric settings, and often only… 9 arXiv — NLP / Computation & Language research 11d ago Chain-of-Experience for Continual LLM Improvement arXiv:2608.18027v1 Announce Type: new Abstract: Humans continuously learn from experience, whereas conventional large language model (LLM) evaluations ignore the models' ability to improve through inference-time interaction. In this paper, we study how LLMs learn from iterative… 32 arXiv — NLP / Computation & Language research 11d ago TokEval: A Tokenizer Evaluation Suite arXiv:2608.18062v1 Announce Type: new Abstract: Language model tokenizers are typically selected with minimal evaluation, despite the fact that their design choices directly impact model capabilities. This can be partly attributed to a limited understanding of which tokenizer… 24 arXiv — NLP / Computation & Language research 11d ago Multi-Agent AI System for Radiology Report Structuring and Quality Assurance with Independent Radiologist Evaluation arXiv:2608.18072v1 Announce Type: new Abstract: Purpose: To develop and evaluate a locally deployed multi-agent AI system for radiology report structuring and quality assurance. Materials and Methods: This retrospective study included 638 radiology reports from CT examinations… 32 arXiv — NLP / Computation & Language research 11d ago Grounding Healthcare LLMs in a Causal Knowledge Graph: Framework, Metrics, and a Cardiovascular Pilot arXiv:2608.15382v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly proposed for healthcare decision support, but their evaluations still reward single-answer accuracy rather than reasoning about interventions, mechanisms, harms, evidence, and… 16 arXiv — NLP / Computation & Language research 11d ago The Authenticity Gap in Human Evaluation arXiv:2205.11930v3 Announce Type: replace Abstract: Human ratings are the gold standard in NLG evaluation. The standard protocol is to collect ratings of generated text, average across annotators, and rank NLG systems by their average scores. However, little consideration has… 11 arXiv — NLP / Computation & Language research 11d ago SCOPE: Selective Conformal Optimized Pairwise LLM Judging arXiv:2602.13110v4 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used as scalable judges in pairwise evaluation, but they remain prone to miscalibration and biases. We propose \textsc{Scope} (Selective Conformal Optimized Pairwise Evaluation), a… 21 arXiv — NLP / Computation & Language research 11d ago Eval4Sim: An Evaluation Framework for Persona Simulation arXiv:2603.02876v2 Announce Type: replace Abstract: Large Language Model personas, explicit profiles specifying a user's attributes, preferences, and behavioural tendencies, are increasingly used to simulate human conversations for user modelling, social reasoning, and… 13 Hugging Face Daily Papers research 11d ago Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents Abstract Empirical evaluation of diverse memory substrates for long-horizon LLM agents reveals regime-dependent trade-offs, motivating adaptive substrate routing for reliable agent memory. Generated by thinkingmachines/Inkling-Small Memory is becoming core infrastructure for… 13 The Information — AI news-outlet 11d ago QuickBooks Challengers Are Fetching $1 Billion Valuations Fast growth at AI startups like Mercor and Replit is also boosting the fortunes of startups that help them balance their books. Rillet, which makes AI-powered accounting software it sells to other companies, said Tuesday it has raised $100 million in a funding round led by… 29 TechCrunch — AI news-outlet 11d ago Etched’s valuation doubles to $21B in a month Jane Street has installed Etched's first shipped AI cluster system, and was so impressed, it led another massive round, the startup says. 23 Hugging Face Daily Papers research 12d ago How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks Abstract Autonomous research agents evaluated across the full scientific lifecycle reveal a pervasive lack of metacognitive self-correction, motivating a new benchmark and failure taxonomy. Generated by thinkingmachines/Inkling-Small AI has long assisted scientific research, but… 8 arXiv — Machine Learning research 12d ago Geometry Is Not Robustness: A Trajectory-Level Study of PGD Evaluation arXiv:2608.14594v1 Announce Type: new Abstract: Projected Gradient Descent (PGD) is widely used to evaluate adversarial robustness, typically via final adversarial accuracy, which does not capture model behaviour throughout the attack. Recent work proposes trajectory-level… 28 arXiv — Machine Learning research 12d ago Paired Exact-Reset Evaluation of a Prediction-Derived Medium-to-Full World-Model Cascade arXiv:2608.14650v1 Announce Type: new Abstract: Existing adaptive-inference and world-action-model systems use cheap-stage outputs or predicted futures to allocate additional computation. We study a narrower question: under paired exact-reset physical outcomes, can a… 23 arXiv — Machine Learning research 12d ago PureTD: Reinforcement Learning for Backgammon Money Games with No Evaluation-time Search arXiv:2608.15146v1 Announce Type: new Abstract: We revisit Tesauro's TD-Gammon for backgammon money games in the setting of no evaluation-time search. Both checker play and cube action (use of the doubling cube) are learned from scratch via self-play reinforcement learning (RL),… 21 arXiv — Machine Learning research 12d ago Structuring Semantic Embeddings for Principle Evaluation: A Prototype-Guided Contrastive Learning Approach arXiv:2608.15224v1 Announce Type: new Abstract: Reliable post-hoc evaluation asks whether already generated text satisfies a target criterion after generation. In this paper we study a focused frozen-embedding setting using principle-evaluation proxy tasks: toxicity detection,… 34 arXiv — Machine Learning research 12d ago QSMP: finding representative time series subsequences through Quick Shift+Matrix Profile arXiv:2608.15492v1 Announce Type: new Abstract: Finding representative waveforms in long time series has scientific and practical value in many domains, as it enables summarization and visualization of large time series datasets, and downstream tasks like classification and… 4 arXiv — NLP / Computation & Language research 12d ago HarmProfile: Characterizing Harmful Distributions in Frontier LLMs arXiv:2608.14577v1 Announce Type: new Abstract: Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome rather than as an object of analysis. Consequently, little is known about the harmful outputs produced during model… 10 arXiv — NLP / Computation & Language research 12d ago What to Forget in Unlearning? Forget Set Curation for Language Models arXiv:2608.14855v1 Announce Type: new Abstract: Machine unlearning aims to remove targeted data or behaviors from a trained model without retraining from scratch. Yet most evaluations assume that the examples to forget are already known. In realistic language-model deployments,… 7 arXiv — NLP / Computation & Language research 12d ago Interpretable Cross-Lingual Alignment in Small Language Models: Probing Cultural and Pragmatic Reasoning in Japanese-English Bilingual LLMs arXiv:2608.14896v1 Announce Type: new Abstract: Large language models work well on English and behave in poorly understood ways on languages typologically far from it. Japanese is a clean example, where evaluation still leans on translation quality and JGLUE-style benchmarks,… 21 arXiv — NLP / Computation & Language research 12d ago How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks arXiv:2608.14905v1 Announce Type: new Abstract: AI has long assisted scientific research, but the rapid advance of LLMs and agentic scaffolds is reshaping the landscape; a single system can now carry whole-stage research from an initial hypothesis all the way to final published… 26 arXiv — NLP / Computation & Language research 12d ago Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents arXiv:2608.15008v1 Announce Type: new Abstract: Memory is becoming core infrastructure for long-horizon LLM agents, yet existing evaluations offer limited guidance on which memory substrate, namely the underlying medium in which memory is represented and stored, should be used… 14 Page 3 of 10 · 500 articles ← Newer Older →