News / #funding Tag Funding 500 articles archived under #funding · RSS Sign in to follow arXiv — NLP / Computation & Language research 19d ago EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models arXiv:2608.09189v1 Announce Type: new Abstract: Despite significant advances in instruction-following and auditory comprehension, the evaluation of Emotional Intelligence (EI) in Spoken Language Models (SLMs) remains confined to rudimentary paralinguistic perception, lacking a… 7 arXiv — NLP / Computation & Language research 19d ago Accurate but Natural? Diagnosing Grammatical and Idiomatic Gaps in Japanese EFL Writing arXiv:2608.09289v1 Announce Type: new Abstract: Second language writing research distinguishes grammatical accuracy from native-like idiomaticity, yet automated writing evaluation often conflates these dimensions. This study introduces a layered LLM-correction pipeline that… 18 arXiv — NLP / Computation & Language research 19d ago Build it, Break it, Repeat: Benchmarking and improving LLM-manipulated disinformation detection in social media posts arXiv:2608.09510v1 Announce Type: new Abstract: Detecting machine-generated disinformation on social media is increasingly difficult as large language models (LLMs) make it easier to generate and rewrite misleading content at scale. Static benchmark evaluations, measuring… 38 arXiv — NLP / Computation & Language research 19d ago How Do Large Language Models Judge Social Attraction? Evidence from Theory-Grounded Persona Ratings Across Multiple LLMs and Humans arXiv:2608.09717v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to perform subjective evaluations traditionally made by humans, yet their validity as social judges remains unclear. This paper examines whether LLMs can assess social attraction… 16 arXiv — NLP / Computation & Language research 19d ago Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness arXiv:2608.09766v1 Announce Type: new Abstract: Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlooks locale… 30 TechCrunch — AI news-outlet 20d ago Discovered Materials is playing AI whack-a-mole to hunt cooler chips Discovered Materials raised $9 million to fund the hunt for more novel materials to build more efficient chips. 36 Hugging Face Daily Papers research 20d ago Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination Abstract Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized. Contamination mitigation evaluation intervenes in the decoding process to suppress memorization and restore a contaminated model's genuine capability,… 14 arXiv — Machine Learning research 20d ago Unmasking Removal-Budget Confounding: A Matched Operating-Point Evaluation Framework for Adaptive Data Cleaning arXiv:2608.06511v1 Announce Type: new Abstract: Adaptive data-cleaning methods replace manual filtering thresholds with data-driven partitions. However, changing the partition granularity, the number of groups used to segment samples by estimated corruption risk, can implicitly… 35 arXiv — Machine Learning research 20d ago When GNNs Fail: Quantifying and Overcoming Temporal Correlation Volatility in Time Series arXiv:2608.07333v1 Announce Type: new Abstract: Modeling multivariate time series by representing them as graphs, where individual series act as nodes and pairwise temporal corre- lations serve as edges, has gained significant traction. Recent advances in Graph Neural Networks… 6 arXiv — NLP / Computation & Language research 20d ago The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents arXiv:2608.06663v1 Announce Type: new Abstract: Frontier language models solve reasoning problems in a single forward pass that would have been research contributions years ago, yet fail at multi-hour tasks: losing track of earlier decisions, declaring half-finished work done,… 9 arXiv — NLP / Computation & Language research 20d ago Do Audio Language Models Use Paralinguistic Evidence? Counterfactual Audits for Response Evaluation arXiv:2608.06718v1 Announce Type: new Abstract: Audio-language models (ALMs) are increasingly used as judges for speech-to-speech systems, but a judge that receives audio may not actually use paralinguistic evidence. We introduce counterfactual audits for paralinguistic response… 13 arXiv — NLP / Computation & Language research 20d ago Can Language Models Imagine Without Seeing? Ekphrasis: Measuring Visual Creative Ideation in Text-Only LLMs arXiv:2608.06967v1 Announce Type: new Abstract: Current evaluations do not isolate whether text-only language models can originate visual concepts before image generation. Fluent visual prose can hide visual-plan failures: an answer may appear creative while repeating familiar… 36 arXiv — NLP / Computation & Language research 20d ago From Test-Time Scaling to Reusable Memory: Measuring Crystallization in Text-to-SQL arXiv:2608.07213v1 Announce Type: new Abstract: Test-time scaling can correct difficult text-to-SQL queries, but the extra computation is normally discarded after each answer. Systems increasingly retain verified repair episodes, yet evaluations still report one end-to-end… 18 arXiv — NLP / Computation & Language research 20d ago Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination arXiv:2608.07341v1 Announce Type: new Abstract: Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized. \textbf{Contamination mitigation evaluation} intervenes in the decoding process to suppress memorization and… 7 arXiv — NLP / Computation & Language research 20d ago An Exploratory Evaluation of LLM-Assisted Rewriting of Moderate-Complexity Financial Sentences for DisCoCat-Based Sentiment Analysis arXiv:2608.07439v1 Announce Type: new Abstract: Quantum natural language processing (QNLP) provides a grammar-aware framework for text modeling, and Distributional Compositional Categorical (DisCoCat) is one of its theoretically grounded formulations. Prior work on financial… 29 arXiv — NLP / Computation & Language research 20d ago ADIAS: Automated Design of Interactive Agentic Systems arXiv:2608.06410v1 Announce Type: cross Abstract: Automated agent design improves agent harnesses through iterative revision, evaluation, and feedback summarization. Existing methods are largely candidate-centric: cross-round experience is organized around candidate agents,… 36 arXiv — NLP / Computation & Language research 20d ago How Should I Pick a Foundation Model for My Robot? In Favor of a Community Evaluation Framework for Social Robots arXiv:2608.06898v1 Announce Type: cross Abstract: Researchers who seek to build social robot applications on foundation models are faced with a difficult question: how should we pick a model? Public leaderboards offer little guidance: the demands of real-time, embodied social… 14 arXiv — NLP / Computation & Language research 20d ago Recipes for Creativity: Iterative Generation and Evaluation in Large Language Models arXiv:2608.07243v1 Announce Type: cross Abstract: Generative models are often evaluated through singular artifacts, whereas human creativity typically emerges through iterative generation, appraisal, and refinement. This pilot study examines whether iterative search improves LLM… 20 arXiv — NLP / Computation & Language research 20d ago How Long Reasoning Chains Influence LLMs' Judgment of Answer Factuality arXiv:2604.06756v2 Announce Type: replace Abstract: Large language models (LLMs) has been widely adopted as a scalable surrogate for human evaluation, yet such judges remain imperfect and susceptible to surface-level biases. One possible reason is that these judges lack… 17 Hugging Face Daily Papers research 20d ago StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Abstract Deploying autonomous multimodal agents in continuous, real-world environments requires them to ingest unbounded audio-visual streams and maintain hour-scale memory. However, current evaluations predominantly rely on brief clips and multiple-choice formats. This design… 27 r/LocalLLaMA community 21d ago DeepSeek V4 Flash 0731 hits 82.7% on Terminal-Bench 2.1 in an independent public-harness run (445 trials) Disclosure: I’m the author of Ante. DeepSeek recently reported an 82.7% score on Terminal-Bench 2.1 for DeepSeek V4 Flash 0731. Its evaluation used “DeepSeek Harness minimal mode,” which hasn’t been released yet. We wanted to see whether the reported result could be… 14 r/MachineLearning community 21d ago Evaluation metrics - [D] I wanted to ask about the selection of evaluation metrics. In which scenario we use ROC-AUC score and in which scenario we use f1 score as an evaluation metric in a classification problem to define the model's performance on a specific dataset.   submitted by  … 8 r/LocalLLaMA community 21d ago Repeated generation is worth it and self-evaluation is effective I made gemma4 12B write timestamp-anchored summaries of youtube video transcripts. I tested if the summaries have significant qualitative variance and if the SLM can pick the best one by itself. Below is the prompt texts I used. "{{[INPUT]}}<attachement name='original'>",… 18 OpenAI official-blog 22d ago Responding to the next frontier of critical cyber capabilities OpenAI is sharing preliminary cybersecurity evaluations for Astra and the steps we’re taking to strengthen safeguards and security controls. 22 Hugging Face Daily Papers research 23d ago MameLoshnLM: Yiddish Language Model and Evaluation Benchmark Abstract We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have constrained progress in Yiddish… 5 Hugging Face Daily Papers research 23d ago OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Abstract Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA evaluation, data curation, and reinforcement learning.… 25 arXiv — Machine Learning research 23d ago Decoupling Perception from Description: Computation-Grounded Representation Alignment between Multivariate Time Series and Language arXiv:2608.05238v1 Announce Type: new Abstract: Training multimodal models to align time series with language runs into a self-supervision trap. The usual recipe asks an LLM to read a series and write a description, so label quality is capped by the perceptual skill the model is… 31 arXiv — Machine Learning research 23d ago A Unified Risk View of Uncertainty: Posterior Risk for Disentanglement and Evaluation Beyond Proxies arXiv:2608.05995v1 Announce Type: new Abstract: Reliable uncertainty estimates are critical in safety-sensitive applications, where understanding the sources of predictive uncertainty is essential. This often requires disentangling epistemic uncertainty from aleatoric… 33 arXiv — Machine Learning research 23d ago Does Latent Context Help? A Controlled Evaluation of Inverse Reinforcement Learning in Arctic Shipping arXiv:2608.06105v1 Announce Type: new Abstract: Artificial Intelligence (AI)-assisted navigation can help Arctic shipping adapt to rapidly changing sea-ice conditions, but reliable deployment requires reward models that are interpretable and robust to changing environments.… 28 arXiv — Machine Learning research 23d ago Hypothesis Testing with Conditional Queries: Learnability and the Value of Interaction arXiv:2608.06262v1 Announce Type: new Abstract: Model evaluations may fix all tests before observing any responses or select later tests using earlier responses. We study this choice in a conditional-query model on a finite outcome space $\mathcal{X}$ with $|\mathcal{X}|=N$. We… 17 arXiv — NLP / Computation & Language research 23d ago PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs arXiv:2608.05162v1 Announce Type: new Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden states into a passage-level vector, yet no shared protocol exists for comparing… 24 arXiv — NLP / Computation & Language research 23d ago Conditional Cognitive Biases in LLMs: How Biased User Turns Modulate In-Context Reasoning arXiv:2608.05166v1 Announce Type: new Abstract: We present an evaluation of cognitive bias expression in state-of-the-art instruction-tuned LLMs under realistic multi-turn interaction settings. Our work introduces a novel three-condition experimental framework that disentangles… 38 arXiv — Machine Learning research 23d ago A Unified Causal Inference Framework for the Desirability of Outcome Ranking Paradigm in Benefit-Risk Evaluation arXiv:2608.05244v1 Announce Type: cross Abstract: We developed a unified covariate-adjusted causal inference framework for estimating the desirability of outcome ranking (DOOR) probability for benefit-risk evaluation in randomized trials and observational studies. The framework… 30 arXiv — Machine Learning research 23d ago Physics-Based Molecular Fingerprints from Spectral Graph Theory Provide Efficient Geometry-Aware Measures of Chemical Similarity arXiv:2608.05336v1 Announce Type: cross Abstract: Molecular representations are essential for the evaluation of molecular similarity and the development of structure-property relationships. Despite the known importance of 3D structure to determine chemical and physical… 10 arXiv — NLP / Computation & Language research 23d ago Beyond Sentiment: Comparing Traditional NLP and LLM-Based Multi-Dimensional Analysis for Political News Evaluation arXiv:2608.05155v1 Announce Type: new Abstract: Traditional sentiment analysis (SA) models, while effective for polarity classification, provide limited insight into the rhetorical, ideological, and framing dimensions of political discourse -- dimensions that are central to… 9 arXiv — NLP / Computation & Language research 23d ago Evidence Lock Before Commitment: A Frozen Interface Degrades LLM-as-Judge Evaluation arXiv:2608.05353v1 Announce Type: new Abstract: LLM judges are often asked to extract criteria and evidence before choosing between candidate answers. This workflow assumes that the intermediate record preserves the information needed for a later verdict. For reasoning-capable… 19 arXiv — NLP / Computation & Language research 23d ago Mitigating Scoring Bias in LLM-as-a-Judge via Random Number Generation arXiv:2608.05726v1 Announce Type: new Abstract: Large Language Models (LLMs) are often used as evaluators of text quality, known as LLM-as-a-Judge, which can outperform conventional automatic evaluation metrics that rely on reference texts. However, LLM evaluators tend to… 38 arXiv — NLP / Computation & Language research 23d ago MameLoshnLM: Yiddish Language Model and Evaluation Benchmark arXiv:2608.05850v1 Announce Type: new Abstract: We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have… 23 arXiv — NLP / Computation & Language research 23d ago A Study of LLMs' Preferences for Libraries and Programming Languages arXiv:2503.17181v4 Announce Type: cross Abstract: Despite the rapid progress of large language models (LLMs) in code generation, existing evaluations focus on functional correctness or syntactic validity, overlooking how LLMs make critical design choices such as which library or… 34 arXiv — NLP / Computation & Language research 23d ago TriQua: Reconciling Granularity and Context in Factuality Evaluation arXiv:2608.05228v1 Announce Type: cross Abstract: The "decompose-then-verify" paradigm for LLM factuality evaluation faces a fundamental trade-off: atomic facts, i.e., one sentence conveying one unit of information, often omit essential context, while broader statements lack the… 14 arXiv — NLP / Computation & Language research 23d ago From Sports to Safety: Benchmarking Proactive Risk Inference in MLLMs arXiv:2608.05560v1 Announce Type: cross Abstract: Timely anticipation of physical hazards is essential for real-world safety, yet existing MLLM evaluations focus on harmful content or general risks, leaving proactive physical hazard prediction underexplored. Sports provide a… 33 arXiv — NLP / Computation & Language research 23d ago Schema-Guided Hierarchical Information Extraction and Semantic Evaluation Using Generative AI arXiv:2608.06167v1 Announce Type: cross Abstract: We present a schema-based framework for extracting complex, structured information from unstructured text documents using generative AI, followed by automated semantic evaluation of the extracted information against a gold… 31 arXiv — NLP / Computation & Language research 23d ago AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games arXiv:2608.06362v1 Announce Type: cross Abstract: Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or expert time. Since the number of games needed is unknown, fixed-budget evaluations either… 8 arXiv — NLP / Computation & Language research 23d ago Text Generation: A Systematic Literature Review of Tasks, Evaluation, and Challenges arXiv:2405.15604v4 Announce Type: replace Abstract: Text generation has become more accessible than ever, and the growing interest in these systems, especially those using large language models, has spurred a surge in related publications. We provide a systematic literature… 17 Hugging Face Daily Papers research 23d ago What AI Red-Team Evaluations Can and Cannot Prove Abstract Red-team evaluations of AI models support some claims and not others, and the boundary between the two is calculable rather than merely a matter of judgment. We define the evidential ceiling of an evaluation as the largest factor by which one result can move belief… 11 Simon Willison community 23d ago Simon Willison on Technical Blogging Simon Willison on Technical Blogging I was interviewed by Cynthia Dunlop for her "Write that blog!" series back in January, but I just realized I never linked to the interview from my own blog! It includes my answers to the following questions: Why did you start blogging – and… 29 Simon Willison community 23d ago Simon Willison on Technical Blogging Simon Willison on Technical Blogging I was interviewed by Cynthia Dunlop for her "Write that blog!" series back in January, but I just realized I never linked to the interview from my own blog! It includes my answers to the following questions: Why did you start blogging – and… 28 Don't Worry About the Vase community 24d ago AI #180: No Longer In Charge What we know about internal AI models hacking into real companies during cyber evaluations keeps getting worse. 23 TechCrunch — AI news-outlet 24d ago Omilia raises $67M to scale its customer support platform The Series B is the company's second fundraise since it last raised capital in 2020. In that time, it has increased its ARR by 10x to $60 million. 20 Hugging Face Daily Papers research 24d ago AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities Abstract While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. Existing benchmarks primarily evaluate visual transformations on… 15 Page 6 of 10 · 500 articles ← Newer Older →