News / #funding Tag Funding 500 articles archived under #funding · RSS Sign in to follow arXiv — Machine Learning research 4d ago StateTune: Transforming LLM-Assisted EDA Flow Tuning into a Stateful, Closed-Loop Process arXiv:2608.23601v1 Announce Type: cross Abstract: EDA flow parameter tuning is critical for quality-of-results~(QoR), yet the parameter space is large, tightly coupled, and full evaluations are prohibitively expensive. Prior LLM-assisted tuners mainly use the LLM as an external… 10 TechCrunch — AI news-outlet 4d ago India’s Ringg gets backing from Peak XV as it pushes voice AI past the phone call Ringg has raised $10 million from Peak XV as a part of its Series A extension. 25 TechCrunch — AI news-outlet 4d ago Robotics startup Generalist reaches $3B valuation, sources say The $200 million extension comes just months after the physical AI startup reached a $2 billion valuation. 27 r/MachineLearning community 4d ago What would a fair benchmark for agent architecture look like? [D] I am working on an evaluation design and would appreciate criticism before running it. Most coding-agent benchmarks collapse the model and its harness into one score. If a run fails, it is difficult to tell whether the cause was model capability, context assembly, task… 20 TechCrunch — AI news-outlet 4d ago Accel-backed Keenable is indexing the web for AI agents Now exiting stealth mode with a $26 million seed round, Keenable has been building a vast web search index for AI agents. 8 Hugging Face Daily Papers research 5d ago One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows Abstract Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation, app APIs, and function calling. Yet completing consequential work beyond code requires more than producing a plausible response or valid tool call:… 33 Hugging Face Daily Papers research 5d ago Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection Abstract Task-CoEvolve improves LLM harness optimization by adaptively selecting validation tasks and estimating full-set performance from partial evaluations, cutting evaluation costs by 80%. Generated by thinkingmachines/Inkling-Small We present a novel approach to efficient… 30 arXiv — NLP / Computation & Language research 5d ago Wazobia Eval: A Benchmark for Nigerian Pidgin Emotion Understanding, Sarcasm Detection, and Cultural Reasoning arXiv:2608.21369v1 Announce Type: new Abstract: Nigerian Pidgin is one of Africa's most widely spoken languages, yet remains severely underrepresented in language model evaluation. Existing benchmarks primarily focus on translation, transcription, or generic sentiment analysis,… 21 arXiv — NLP / Computation & Language research 5d ago Forgotten in Weights, Recovered by Tools: Agentic Tool Unlearning for LLM Agents arXiv:2608.21544v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as tool-augmented agents, where responses can depend on tool calls and external observations rather than model parameters alone. This creates an evaluation mismatch for LLM… 26 arXiv — NLP / Computation & Language research 5d ago Automating Multi-Hop RAG Evaluation via TRIAD: From Context Extraction to Validated Dataset Generation arXiv:2608.21558v1 Announce Type: new Abstract: Recent advances in LLMs and the adoption of RAG systems in industry have created a need for domain-specific question-answer datasets that can assess RAG performance on proprietary data. Existing datasets, such as HotpotQA,… 31 arXiv — NLP / Computation & Language research 5d ago Can LLMs Truly Forget? Revealing Unlearning Gaps Through Adversarial Evaluation arXiv:2608.21606v1 Announce Type: new Abstract: Machine unlearning aims to remove the influence of targeted training data from a model while preserving its remaining capabilities, but evaluating whether such information has truly become inaccessible remains challenging. Existing… 19 arXiv — NLP / Computation & Language research 5d ago Evaluation Awareness in Language Models: Representation, Verbalization, and Control arXiv:2608.21766v1 Announce Type: new Abstract: Both capability and safety benchmarks rest upon the assumption that the behavior of language models undergoing a test is informative about their behavior in deployment. This assumption can fail, should models infer that they are… 16 arXiv — NLP / Computation & Language research 5d ago Lexical Coupling in GUI Element Grounding: Sentence Embeddings Track Labels across Mobile and Web arXiv:2608.21794v1 Announce Type: new Abstract: GUI grounding evaluations that expose UI elements as text metadata often treat high instruction-element embedding similarity as evidence of semantic grounding. Across three mobile and web benchmarks, we show that this… 28 arXiv — NLP / Computation & Language research 5d ago LLM assisted writing deserves empirical evaluation arXiv:2608.22124v1 Announce Type: new Abstract: LLM-assisted writing is often treated as a detection problem, as it raises questions about clarity, integrity, equity, and evaluation. An analysis of 69,209 Health Informatics papers links it to more focused presentation, broader… 6 arXiv — NLP / Computation & Language research 5d ago LLM Evaluation on Unseen Questions: Contextual Multidimensional IRT Model arXiv:2608.22295v1 Announce Type: new Abstract: Evaluation of large language models (LLMs) increasingly requires predicting how a model will perform on new questions or tasks before collecting large amounts of new annotations. This problem is challenging because question… 11 arXiv — NLP / Computation & Language research 5d ago Register Shifts Break LLM Safety: A Bengali Benchmark with Culturally Grounded Harms arXiv:2608.22335v1 Announce Type: new Abstract: Bengali is the seventh-most-spoken language globally, yet LLM safety evaluation remains overwhelmingly English-centric. We introduce BanglaSafe, a benchmark of 879 Bengali prompts combining 309 natively authored prompts with 570… 12 arXiv — NLP / Computation & Language research 5d ago SchemaGUI: A Schema-Driven Benchmark for Controllable GUI Generation Evaluation arXiv:2608.22390v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong potential in graphical user interface (GUI) generation, but reliable evaluation remains challenging due to uncontrolled data distributions, noisy annotations, and limited layout… 35 arXiv — NLP / Computation & Language research 5d ago Aligned Alone, Misaligned Together: Forecasting Adversarial Capture in LLM Agent Populations arXiv:2608.22444v1 Announce Type: new Abstract: The unit of AI safety evaluation is still the individual model, yet language-model agents are increasingly deployed in interacting populations that read and write one another's decisions. This raises a question no single-agent… 33 Hugging Face Daily Papers research 5d ago Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision Abstract A hierarchical taxonomy and dense supervision strategy improve diffusion-based image editing through fine-grained concepts, large-scale paired data, and granular evaluation. Generated by thinkingmachines/Inkling-Small Existing image editing frameworks predominantly… 17 Hugging Face Daily Papers research 5d ago Same Agent, Different Answers: A Repeat-Aware Audit of Corpus-Induced Answer Churn in Retrieval-Augmented QA Abstract Retrieval-augmented QA systems can exhibit hidden answer churn during index updates without noticeable accuracy changes, motivating compatibility audits alongside utility evaluations. Generated by thinkingmachines/Inkling-Small A retrieval-augmented QA system can return… 21 TechCrunch — AI news-outlet 5d ago Valor, Point72 back General Intuition at $6B valuation as AI startup pushes into robotics General Intuition, the startup building a foundation model that trains generalized AI agents how to move through space and time, is in talks to raise at a $6 billion pre-money valuation from new investors including Valor Ventures, Point72 Ventures, Seven Seven Six. 19 The Information — AI news-outlet 5d ago Shein Expects IPO Valuation of $26 Billion Fast fashion company Shein expects to go public at a valuation of around $26 billion, the company said in a filing with the Hong Kong Stock Exchange. That would be a sharp decline from the $80 billion to $90 billion valuation it was hoping for three years ago. That decline… 10 arXiv — Machine Learning research 6d ago Machine Learning and ARIMA Model Averaging for Adaptive Public Health Forecasting: Comparative Evaluation and an Ontario COVID-19 Case Study arXiv:2608.20406v1 Announce Type: new Abstract: Public health forecasts must respond to abrupt changes in surveillance data without over-extrapolating noise, reporting artifacts, or temporary trends. We evaluated autoregressive integrated moving average (ARIMA), random forest,… 32 arXiv — Machine Learning research 6d ago Nothing Changed but the Model: CellFill -- Bounded In-Cell Learning for Bit-Identical, Revocable Updates to Quantized LLMs arXiv:2608.20873v1 Announce Type: new Abstract: Every way of teaching a deployed language model something new -- full fine-tuning, adapter merging, model editing -- replaces the released checkpoint, and with it every evaluation and cache that referred to those exact bits. We… 13 arXiv — Machine Learning research 6d ago Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment arXiv:2608.21057v1 Announce Type: new Abstract: Agentic large language model (LLM) systems are reshaping scientific workflows in chemistry and drug discovery, but evaluating their open-ended, tool-augmented outputs remains a fundamental bottleneck. Reference-based metrics such… 12 arXiv — Machine Learning research 6d ago Rigorous Evaluation of Large Language Models for Malaria Drug Discovery: Trade-offs in Performance, Scale, and Resource Utility arXiv:2608.20418v1 Announce Type: cross Abstract: We introduce Malaria-Instruct, a curated instruction-following dataset derived from the ChEMBL Legacy Malaria corpus for Malaria virtual screening, and conduct a systematic evaluation of five open-source LLMs; Gemma-2 2B/9B,… 36 arXiv — Machine Learning research 6d ago aiXamine: Unified Black-Box Evaluation of Cross-Dimensional Trade-offs in LLM Safety, Security, and Privacy arXiv:2608.20554v1 Announce Type: cross Abstract: The critical failure modes in deployed large language models (LLMs) are cross-dimensional: a model can score 99.3 in safety alignment while refusing one in three benign queries, or improve across every capability metric while… 29 arXiv — NLP / Computation & Language research 6d ago Who Do Language Models Think Is Competent? A Mechanistic Analysis of Occupational Bias arXiv:2608.20347v1 Announce Type: new Abstract: Language models (LMs) often pass behavioral bias evaluations, but it remains unclear whether they no longer represent the underlying associations that give rise to biases, or have merely learned not to express them. In this study,… 36 arXiv — NLP / Computation & Language research 6d ago Evaluation-as-Search: Adaptive Discovery of Grounding Failures in Meeting Assistants arXiv:2608.20392v1 Announce Type: new Abstract: LLM-powered meeting assistants are deployed at scale, yet systematic evaluation of their grounding fidelity remains limited to static benchmarks that miss failure modes tied to specific discourse structures or reasoning demands. We… 14 arXiv — NLP / Computation & Language research 6d ago Source-Free MT Evaluation Is Not MT Evaluation arXiv:2608.20925v1 Announce Type: new Abstract: Reference-based metrics remain the standard choice in machine translation evaluation, partly because quality estimation methods often correlate less well with human judgments. As a result, source-free, reference-based evaluation… 7 arXiv — NLP / Computation & Language research 6d ago Extractive Summarization for Arabic Documents Using SAraBERT with a Semantic Siamese Similarity Evaluation Metric arXiv:2608.20964v1 Announce Type: new Abstract: In this research, we introduce SAraBERT, an enhanced version of AraBERT which proposes inter-sentence transformer layers for extractive summarization tasks. To ensure that the summaries generated by SAraBERT achieve a high coverage… 21 arXiv — NLP / Computation & Language research 6d ago Free-Text Evaluation of LLMs for 5G Domain Knowledge and Fault Analysis using LLM-as-Judge arXiv:2608.21021v1 Announce Type: new Abstract: Real-world fault analysis in 5G and emerging 6G networks demands domain expertise to analyze free-text diagnostics, including root-cause explanations and recommended actions. LLMs have emerged as a promising approach to automating… 16 arXiv — NLP / Computation & Language research 6d ago Evidence-Consistent Generative Detection under Scenario-Level Distribution Shift arXiv:2608.21043v1 Announce Type: new Abstract: Conventional in-distribution evaluation can overestimate robustness when training and test data share recurring task-specific patterns or surface cues. This risk is especially relevant in social-engineering fraud detection, where… 6 arXiv — NLP / Computation & Language research 6d ago No PUN Intended: Plausible Unknown Names for Person-Centred LLM Evaluation arXiv:2608.21206v1 Announce Type: new Abstract: Person names are widely used as prompt variables in LLM evaluations of factuality, privacy leakage, bias and abstention, but when a name's evidential status is uncontrolled, measurements may conflate memorisation, retrieval, name… 26 arXiv — NLP / Computation & Language research 6d ago ProofJudge: Tool-Grounded LLM Evaluation of Formal Proof Quality in Mathlib arXiv:2608.20432v1 Announce Type: cross Abstract: Formal proofs in Lean 4 that pass the kernel's type checker can nonetheless vary widely in quality. We introduce ProofJudge, an agentic LLM-as-judge system that scores formal proof quality along five dimensions beyond… 7 arXiv — NLP / Computation & Language research 6d ago Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems arXiv:2608.21095v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) grounds Large Language Model (LLM) outputs in external knowledge, but RAG systems usually trust whatever they retrieve, creating a Security-Reliability Gap: high semantic relevance does not… 12 arXiv — NLP / Computation & Language research 6d ago Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation arXiv:2505.16222v2 Announce Type: replace Abstract: With the growing use of large language models(LLMs) as evaluators, their application has expanded to code evaluation tasks, where they assess the correctness of generated code without relying on reference implementations. While… 12 The Information — AI news-outlet 6d ago Nvidia Discusses Perplexity Investment at $30 Billion-Plus Valuation Nvidia is discussing investing in Perplexity as part of an equity round that would value the AI startup at more than $30 billion, according to people with knowledge of the discussion. The round would be worth billions of dollars and boost the startup’s valuation more than 50%… 14 r/MachineLearning community 7d ago The evaluation resolution has been shown to have a significant impact on the identification of the "learning rule" that exhibits the most brain-like characteristics at V1. [R] The preprint can be accessed via the following link: https://arxiv.org/abs/2608.12408 (q-bio.NC / cs.LG). And for the code: https://github.com/nilsleut/evaluation-resolution-rsa The following assertion is frequently made in model-brain comparisons: untrained convolutional neural… 9 The Information — AI news-outlet 8d ago Dragoneer Founder Stad to Buy Timberwolves Controlling Stake Marc Stad, the founder of Dragoneer Investment Group, is buying a controlling stake in the Minnesota Timberwolves and Minnesota Lynx professional basketball teams, according to The New York Times’ Athletic publication. Stad is buying the interest at a $4.5 billion valuation from… 15 Hugging Face Daily Papers research 8d ago QuoteBench: How Matched Scores Can Hide Command-Path Failures Abstract QuoteBench reveals that execution-boundary parsing errors significantly reduce LLM coding agent success, and disclosing the boundary helps recover performance, showing that evaluation must account for deployment configuration rather than treating matched scores as… 18 arXiv — Machine Learning research 9d ago Quantifying Event Impacts on Time Series via Multiscale Contrastive Learning arXiv:2608.19447v1 Announce Type: new Abstract: Shocks that spread through the web, such as cybersecurity breach disclosures, can abruptly disrupt financial time series and cause substantial abnormal losses. While these events are disclosed as discrete records through news… 31 arXiv — Machine Learning research 9d ago Systematic Evaluation of TabPFN-TS for Zero-Shot Probabilistic Heat Load Forecasting in District Heating Networks arXiv:2608.20024v1 Announce Type: new Abstract: District heating energy hubs require reliable heat load forecasts for efficient operational scheduling. Conventional forecasting workflows train system-specific models on historical data, which can become burdensome when networks… 11 arXiv — Machine Learning research 9d ago A Standardized Framework for Machine Learning in Power System Protection arXiv:2608.20181v1 Announce Type: new Abstract: Studies of machine-learning-based power-system protection increasingly report near-perfect scores, yet the meaning of those scores depends strongly on the evaluation setting. Protection task, physical scope, measurements, timing,… 27 arXiv — Machine Learning research 9d ago TorchDCM: A Unified PyTorch-Native Package for Discrete Choice Modeling arXiv:2608.19231v1 Announce Type: cross Abstract: Estimating large and simulation-intensive discrete choice models (DCMs) requires repeated evaluation of utilities, probabilities, derivatives, and simulated likelihoods over many observations, alternatives, and draws. Existing… 12 arXiv — Machine Learning research 9d ago SCAPE: Scenario-Conditioned Simulation-Augmented Policy Evaluation arXiv:2608.19425v1 Announce Type: cross Abstract: Reliable performance evaluation is a central bottleneck for deploying robot-learning policies in real-world conditions. Real-world testing is faithful but costly and difficult to scale, whereas simulation-based testing scales… 6 arXiv — NLP / Computation & Language research 9d ago A Speech Corpus for Mizo Automatic Speech Recognition: Whisper and SraVaani 1.0 Fine-Tuning with Morphology-Aware Evaluation arXiv:2608.19361v1 Announce Type: new Abstract: This study reports the development of an Automatic Speech Recognition (ASR) system in Mizo, a low-resource language. The development included collecting 17.62 hours of speech data, curating it, and fine-tuning the Mizo ASR system… 4 arXiv — NLP / Computation & Language research 9d ago Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM Agents arXiv:2608.19564v1 Announce Type: new Abstract: Persistent memory can personalize an LLM agent, but an incorrect durable update can silently distort future behavior. We study the memory-clarification boundary: whether interaction-derived information should be persisted, used… 17 arXiv — NLP / Computation & Language research 9d ago One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows arXiv:2608.19741v1 Announce Type: new Abstract: Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation, app APIs, and function calling. Yet completing consequential work beyond code requires more than producing a… 23 arXiv — NLP / Computation & Language research 9d ago Stopping and Routing LLM Judge Panels arXiv:2608.19802v1 Announce Type: new Abstract: LLM evaluation pipelines often have many candidate judges: general LLM-as-a-judge prompts, reward models, safety classifiers, confidence variants, and task-specific verifiers. The deployment question is not only which judge is… 20 Page 2 of 10 · 500 articles ← Newer Older →