News / #funding Tag Funding 500 articles archived under #funding · RSS Sign in to follow arXiv — NLP / Computation & Language research 17d ago When the Knowledge Base Becomes the Gold Standard: Measuring Resource-Shared Evaluation Loops in Entity-Level Machine Translation arXiv:2608.11843v1 Announce Type: new Abstract: The Seungjeongwon Ilgi, a UNESCO Memory of the World record, is only 37.4% translated, and the most conspicuous failure mode in automatic translation is the person name -- a misread name corrupts the historical fact rather than… 28 arXiv — NLP / Computation & Language research 17d ago Benchmarking LLM Judges for Mobile Agent Evaluation arXiv:2608.11434v1 Announce Type: cross Abstract: Mobile agent benchmarks increasingly rely on LLM-based judges to evaluate task completion, yet the reliability of these judges on mobile agent trajectories remains largely unexamined. We introduce MobileJudgeBench, a benchmark… 17 arXiv — NLP / Computation & Language research 17d ago ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents arXiv:2608.11878v1 Announce Type: cross Abstract: Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually implemented or reused… 25 arXiv — NLP / Computation & Language research 17d ago Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation arXiv:2608.12150v1 Announce Type: cross Abstract: Standard evaluation of large language models assumes stable model rankings across inference conditions. We challenge this assumption by varying the token generation budget, i.e., the maximum tokens a model may produce, across… 9 Hugging Face Daily Papers research 17d ago ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents Abstract ToolHazard is a scalable framework that synthesizes adversarial environments to test LLM agents against indirect prompt injections, revealing vulnerabilities and improving defensive alignment. Generated by thinkingmachines/Inkling-Small Large language model (LLM) agents… 12 Hugging Face Daily Papers research 17d ago SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure Abstract SkillZip compresses self-evolving agent skills by finding a minimal faithful structural explanation that shares repeated rules and procedures while preserving rare exceptions, without requiring evaluation rollouts. Generated by thinkingmachines/Inkling-Small… 26 TechCrunch — AI news-outlet 17d ago AI coding startup Cognition reportedly already in talks to raise at $40B valuation Cognition may be looking to raise another mega round just a few months after raising $1 billion at a $26 billion valuation. 27 TechCrunch — AI news-outlet 17d ago OpenAI-backed Thrive Holdings raises $2B to bring AI to the enterprise Thrive Holdings has raised $2 billion in new funding at a $12 billion valuation from investors like SoftBank, D1 Capital Partners, and Alitmeter Capital. 23 TechCrunch — AI news-outlet 17d ago Lovable confirms new $13.3B valuation, raises another $400M This new funding comes after Lovable hit $500 million in annualized run rate revenue in June, the startup told TechCrunch. 24 TechCrunch — AI news-outlet 17d ago Everything announced at Made by Google ’26: Pixel 11, Pixel Watch 5, Pixel Tag, and tons of Gemini features From the Pixel 11 series and a brand new competitor to Apple’s AirTag, here are all the announcements from the Made by Google 2026 event. 23 TechCrunch — AI news-outlet 18d ago AI code-testing startup Blacksmith’s valuation jumps almost 10x in less than a year Blacksmith says revenue has grown more than tenfold over the past year. 26 arXiv — Machine Learning research 18d ago The Evaluation Protocol Determines the Result: An Independent Reproduction of LeWorldModel on TwoRoom arXiv:2608.10145v1 Announce Type: new Abstract: LeWorldModel trains a latent world model with a prediction loss and a single anti-collapse regulariser, and reports approximately 87% of goals reached on TwoRoom, its simplest diagnostic environment. We reproduce that result by… 18 arXiv — Machine Learning research 18d ago A matched-integrator evaluation of Hamiltonian neural networks on pendulum and Kepler dynamics arXiv:2608.10235v1 Announce Type: new Abstract: Hamiltonian Neural Networks (HNNs) parameterize conservative dynamics through a learned scalar Hamiltonian, providing an architectural prior that is absent from generic vector-field neural networks. We evaluate this prior under a… 34 arXiv — Machine Learning research 18d ago Toward Human Rights Benchmarking for LLMs: A Pilot Methodology arXiv:2608.10268v1 Announce Type: new Abstract: Large language models (LLMs) increasingly mediate legal determinations over what human rights are realized, and how. Yet, no evaluation benchmark exists to assess whether they can reason correctly about human rights law. To this… 14 arXiv — Machine Learning research 18d ago Retrieval-Corrected Conformal Prediction for Time Series arXiv:2608.10553v1 Announce Type: new Abstract: Conformal prediction (CP) provides distribution-free prediction intervals for fixed forecasters, but its standard calibration procedure is often inefficient for time series data, where forecast errors are temporally dependent and… 12 arXiv — NLP / Computation & Language research 18d ago Optimize Cheap, Deploy Strong: Cost-Aware Cross-Tier Transfer for Evolutionary Optimization arXiv:2608.10694v1 Announce Type: cross Abstract: Evolutionary optimization of LLM prompts and agentic programs (e.g., GEPA) is dominated by fitness evaluation: scoring each candidate runs an answering LLM over a validation set, so the evaluator's price tier dictates total… 25 arXiv — Machine Learning research 18d ago Physics-informed Diffusion Generative Model for Time-Series Data Synthesis in Dynamic Systems arXiv:2608.10941v1 Announce Type: new Abstract: Industrial time-series signals, such as turbine temperature and rotational speed in aero-engines, are essential for monitoring the health and operational status of complex dynamical systems. However, collecting such data is often… 23 arXiv — Machine Learning research 18d ago Do AI weather models miss extremes? arXiv:2608.09972v1 Announce Type: cross Abstract: First-generation AI weather models are often reported to underperform at extremes, mostly in reanalysis-based evaluations of deterministic regression systems. We verify eleven physical and AI forecast systems against European… 30 arXiv — NLP / Computation & Language research 18d ago The Multilingual Quantization Tax: Structural Collapse and Typological Fragility in Edge SLMs arXiv:2608.09941v1 Announce Type: new Abstract: While 4-bit weight quantization is critical for deploying Small Language Models (SLMs) on edge devices, evaluations of the resulting performance degradation-the quantization tax-remain overwhelmingly English-centric. We present a… 30 arXiv — NLP / Computation & Language research 18d ago ASR-Roundtrip Evaluation Can Mask Context- and Convention-Dependent Reading Errors in Chinese News TTS arXiv:2608.10606v1 Announce Type: new Abstract: ASR-roundtrip evaluation is widely used as a scalable proxy for text-to-speech (TTS) intelligibility, but it can produce false negatives for reading errors perceived by listeners. We study Chinese news TTS spans whose correct… 5 arXiv — NLP / Computation & Language research 18d ago Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASR arXiv:2608.10670v1 Announce Type: new Abstract: At corpus sizes typical of low-resource dialects, single-run comparisons can yield gains that do not replicate. We show this for Garhwali, an under-resourced Indo-Aryan language of the central Himalaya, building the first… 28 arXiv — NLP / Computation & Language research 18d ago VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World? arXiv:2608.10875v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments. Everyday life assistance is different. A task runs… 20 arXiv — NLP / Computation & Language research 18d ago A Cost-Efficient Routing Pipeline for Multilingual Short-Text Classification Using Small Language Models arXiv:2608.10939v1 Announce Type: new Abstract: Multilingual short-text classification supports operational systems such as content moderation, customer support routing, and intent recognition, yet aggregate evaluation often hides large differences between high-resource and… 23 arXiv — NLP / Computation & Language research 18d ago Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents arXiv:2608.11110v1 Announce Type: new Abstract: When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discards the actions. Yet those actions are the product:… 38 arXiv — NLP / Computation & Language research 18d ago OpenPM: Auditable Point-in-Time Evaluation for LLM Portfolio-Management Agents arXiv:2608.09988v1 Announce Type: cross Abstract: Large language models are increasingly used to read markets, assess risk, and allocate capital. However, reported results for LLM trading agents can be inflated by look-ahead leakage, optimistic execution, and risk mandates that… 31 arXiv — NLP / Computation & Language research 18d ago No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding arXiv:2503.05061v3 Announce Type: replace Abstract: Reliable evaluation of large language models (LLMs) is critical as their deployment rapidly expands, particularly in high-stakes domains such as business and finance. The LLM-as-a-Judge framework, which uses prompted LLMs to… 37 Hugging Face Daily Papers research 18d ago TSDS-Toolbox: A Toolbox for Measuring Time-Series Dataset Similarity Abstract A unified toolbox enables reproducible comparison and extension of time-series dataset similarity methods for forecasting, classification, and generation tasks. Generated by thinkingmachines/Inkling-Small The rapid advancement of artificial intelligence (AI) has… 10 Hugging Face Daily Papers research 18d ago Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness Abstract Researchers propose source-contrastive evaluation via a localized benchmark to detect data contamination and assess localization robustness in multilingual translation models. Generated by thinkingmachines/Inkling-Small Multilingual translation benchmarks are typically… 4 Hugging Face Daily Papers research 18d ago Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure Abstract Optimized GPU kernel benchmarks reveal that evolutionary LLM proposals exploit evaluation configurations, causing widespread failure to generalize to held-out settings. Generated by thinkingmachines/Inkling-Small Benchmarks for systems that are optimized against the… 11 Hugging Face Daily Papers research 19d ago MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models Abstract MMOOC is a large-scale benchmark assessing whether multimodal language models can correctly refuse out-of-context questions while answering shifted in-context questions, revealing that current models struggle to balance these abilities. Generated by… 37 Hugging Face Daily Papers research 19d ago A^2E : An End-to-End Agent Auditing Engine Abstract A2E is an end-to-end evaluation engine for agent harnesses that uses a standardized task protocol and execution traces to assess capabilities across efficiency, tool use, planning, and error recovery. Generated by thinkingmachines/Inkling-Small With the rapid… 9 arXiv — Machine Learning research 19d ago Neural Operators for Immersed-Boundary Soft Swimmers Locomotion arXiv:2608.07722v1 Announce Type: new Abstract: High-fidelity immersed-boundary simulation resolves the coupled motion of a deforming swimmer and its surrounding flow, but the resulting cost limits repeated evaluations for engineering design, parameter studies, and control. We… 16 arXiv — Machine Learning research 19d ago When Does Trace-Driven Evaluation Mislead MoE Expert Caching? Replay Semantics, Workload Contamination, and Operating Regimes arXiv:2608.07911v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models have outgrown accelerator memory, and offloading expert weights to host memory is now standard. This makes expert cache management an attractive lever: a policy that raised the hit rate would cut… 28 arXiv — Machine Learning research 19d ago TSDS-Toolbox: A Toolbox for Measuring Time-Series Dataset Similarity arXiv:2608.08119v1 Announce Type: new Abstract: The rapid advancement of artificial intelligence (AI) has significantly accelerated research in time-series analysis, particularly in forecasting, classification, and generation tasks. Recent models, especially foundation models,… 38 arXiv — Machine Learning research 19d ago FreSH: Frequency-Segmented Hierarchical Multi-Expert Framework for Multivariate Time Series Classification arXiv:2608.08207v1 Announce Type: new Abstract: Multivariate Time Series Classification (MTSC) demands models that can effectively capture complex temporal patterns across multiple scales while remaining computationally efficient. However, existing approaches generally struggle… 24 arXiv — Machine Learning research 19d ago The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World arXiv:2608.08239v1 Announce Type: new Abstract: LLM routers promise efficiency by matching each request to the cheapest adequate model, and are increasingly applied per step inside multi-step agents. Yet agentic routers are evaluated like single-turn routers: by replaying logged… 5 arXiv — Machine Learning research 19d ago Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure arXiv:2608.08722v1 Announce Type: new Abstract: Benchmarks for systems that are optimized against the evaluation signal measure something different from what they claim. We document this concretely in two GPU-kernel-optimization suites with held-out generalization gates:… 24 arXiv — Machine Learning research 19d ago Agentic Anomaly Detection with ORCA-Style Dynamic Inductive Bias Adaptation in Multimodal Wearable Time Series Data arXiv:2608.08859v1 Announce Type: new Abstract: Wireless Body Area Networks (WBANs) generate multivariate physiological time series that are highly nonstationary and must often be processed under strict computational and memory constraints. A critical yet underexplored challenge… 19 arXiv — NLP / Computation & Language research 19d ago Unified Hallucination Fuzzing for Multimodal Large Language Models arXiv:2608.07525v1 Announce Type: new Abstract: Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high-stakes applications. Existing evaluations, predominantly based on static benchmarks, suffer from… 16 arXiv — NLP / Computation & Language research 19d ago SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators arXiv:2608.07641v1 Announce Type: new Abstract: The rapid advancement of large language models has transformed survey writing from a months-long manual effort into an automated process. As generation scales, reliable evaluation becomes the bottleneck, and LLMs are increasingly… 36 arXiv — NLP / Computation & Language research 19d ago Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation arXiv:2608.07763v1 Announce Type: new Abstract: Vision-language models (VLMs) have achieved strong performance on tasks such as image captioning, visual question answering, and image-to-text generation. However, they are predominantly trained on English-centric data, which… 19 arXiv — NLP / Computation & Language research 19d ago On the use of foundation models in cognitive science arXiv:2608.07812v1 Announce Type: new Abstract: A host of recent studies have evaluated the cognitive and developmental alignment of Foundation Models (FMs). These investigations include evaluations of their correspondence to adult performance across a range of cognitive… 34 arXiv — NLP / Computation & Language research 19d ago SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs arXiv:2608.07862v1 Announce Type: new Abstract: Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguistic diversity and culturally grounded safety risks present in other languages.… 21 arXiv — NLP / Computation & Language research 19d ago Thinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions arXiv:2608.07968v1 Announce Type: new Abstract: Reasoning language models increasingly use test-time compute to improve performance, but existing evaluations typically study this compute one question at a time. Yet when multiple problems share an end-to-end cost or latency… 33 arXiv — NLP / Computation & Language research 19d ago A Grounded and Decomposed Framework for Relation-Level Hallucination Evaluation in Abstractive Summarization arXiv:2608.08180v1 Announce Type: new Abstract: Abstractive text summarization systems frequently generate fluent yet unfaithful summaries by fabricating or distorting relationships between entities and events. Such relation-level hallucinations undermine the reliability of… 28 arXiv — NLP / Computation & Language research 19d ago Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations? arXiv:2608.08283v1 Announce Type: new Abstract: Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows is limited by the lack of reliable evaluation. We investigate whether existing automatic… 30 arXiv — NLP / Computation & Language research 19d ago Position Bias in Ordinal Classification: A Systematic Evaluation arXiv:2608.08869v1 Announce Type: new Abstract: Large language models are increasingly used for ordinal classification, yet semantically equivalent changes to prompt organization can alter their predictions. We conduct systematic experiments to characterize positional bias from… 36 arXiv — NLP / Computation & Language research 19d ago How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review arXiv:2608.08975v1 Announce Type: new Abstract: As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how… 24 arXiv — NLP / Computation & Language research 19d ago ELICITED: EHR-grounded Longitudinal Interactive Conversations for Information-seeking Triage Evaluation and Decision-making arXiv:2608.09024v1 Announce Type: new Abstract: Emergency-department (ED) triage requires clinicians to rapidly identify patients who need immediate attention, determine who can safely wait, and prioritize limited clinical resources. At presentation, however, information may be… 5 arXiv — NLP / Computation & Language research 19d ago Evo-Bench: Can Language Models Improve Agent Harness? arXiv:2608.09096v1 Announce Type: new Abstract: Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution---the agent's capacity to autonomously… 27 Page 5 of 10 · 500 articles ← Newer Older →