News / #safety Tag Safety + alignment 500 articles archived under #safety · RSS Sign in to follow VentureBeat — AI news-outlet 1mo ago The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway Across 157 enterprises, organizations are granting AI agents more autonomy while trusting the evaluations meant to gate that autonomy less. Half have already shipped an agent that passed their internal evaluations and then failed a customer in production; only one in twenty… 26 r/LocalLLaMA community 1mo ago Filings: Dario Amodei gave $1M in May to Public First, a super PAC advocating for AI safety regulations, seemingly his first seven-figure political donation   submitted by   /u/pscoutou [link]   [comments] 26 arXiv — Machine Learning research 1mo ago CoDiffGRN: Rethinking Gene Regulatory Network Inference via the BEELINE-KGC Benchmark and Co-evolutionary Discrete Diffusion arXiv:2607.13120v1 Announce Type: new Abstract: Inferring gene regulatory networks (GRNs) from single-cell transcriptomic data is crucial for biological discovery, yet existing approaches suffer from a fundamental misalignment with real-world needs. Researchers typically seek a… 35 arXiv — Machine Learning research 1mo ago SteinGate: Tail-Sensitive Safe Reinforcement Learning via Stein Discrepancy arXiv:2607.13175v1 Announce Type: new Abstract: Safe reinforcement learning typically enforces safety by bounding expected cumulative costs, a criterion that often fails to detect rare but catastrophic tail events. To overcome these limitations, this paper introduces SteinGate,… 8 arXiv — Machine Learning research 1mo ago Distributionally Robust and Safe Imitation Learning arXiv:2607.13436v1 Announce Type: new Abstract: Imitation learning (IL) has achieved remarkable success in complex decision-making tasks. However, its performance is highly sensitive to distribution shifts, which can pose significant safety risks. We propose a distributionally… 28 arXiv — Machine Learning research 1mo ago Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows arXiv:2607.13078v1 Announce Type: cross Abstract: LLMs are now proposed for fraud detection, scam investigation, content moderation, and other trust-and-safety workflows. Much of the public literature still evaluates them as models, with less attention to their behavior as… 38 Hugging Face Daily Papers research 1mo ago PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails Abstract Image guardrails are typically trained and evaluated under a fixed safety policy, implicitly treating safety as an intrinsic property of an image. Real deployments are different: the same image may be allowed in one product, restricted in another, and newly disallowed… 25 OpenAI official-blog 1mo ago The US is advancing AI safety through state and federal action OpenAI outlines a “reverse federalism” approach to AI governance, where state laws help build a national framework for safe, democratic AI. 35 OpenAI official-blog 1mo ago GPT-Red: Unlocking Self-Improvement for Robustness Explore GPT-Red, OpenAI’s automated red teaming system that uses self-play to improve AI safety, alignment, and prompt injection robustness. 14 arXiv — Machine Learning research 1mo ago Scalable Optimal Transport Algorithm for Network Alignment arXiv:2607.11952v1 Announce Type: new Abstract: Network alignment identifies node correspondences across different networks and is a fundamental primitive in many data science applications, including social network analysis, fraud detection, and knowledge graph integration.… 20 arXiv — Machine Learning research 1mo ago Exploring Zero-Shot Foundation Models for Multivariate Time Series Anomaly Detection arXiv:2607.12454v1 Announce Type: new Abstract: Multivariate Time Series Anomaly Detection (MTSAD) is essential for reliability and safety in domains such as industrial process monitoring and financial risk management, yet conventional approaches rely on application-specific… 13 arXiv — Machine Learning research 1mo ago Predictive Modeling of High-Altitude Clear Air Turbulence in the United States: A Machine Learning Approach arXiv:2607.11899v1 Announce Type: cross Abstract: High-altitude Clear Air Turbulence (CAT) poses significant risks to aviation safety due to its unpredictability and challenges in detection. This study leverages machine learning models to improve CAT prediction within U.S.… 24 arXiv — NLP / Computation & Language research 1mo ago CANDI: Contextual Alignment for Niche Domains Question Answering arXiv:2607.11891v1 Announce Type: new Abstract: The deployment of large language models (LLMs) in specialized domains like medical diagnostics and financial advisory necessitates evaluating capabilities beyond general knowledge. Traditional question-answering benchmarks often… 33 arXiv — NLP / Computation & Language research 1mo ago Beyond Parallel Tracking: Interactive Multi-Feature Fusion Drives Semantic Reconstruction from Non-invasive Brain Recordings arXiv:2607.12071v1 Announce Type: new Abstract: Continuous semantic reconstruction from non-invasive neural recordings remains limited by the representational mismatch between semantic feature spaces and neural coding patterns, which severely impedes cross-modal alignment… 34 arXiv — NLP / Computation & Language research 1mo ago Optimization Is Not All You Need arXiv:2607.11977v1 Announce Type: cross Abstract: In 2019, OpenAI released two million GPT-2 outputs-ungrammatical, half broken-to aid the detection of machine-generated text. The alignment that produced their more fluent successors is usually regarded as an engineering… 4 arXiv — NLP / Computation & Language research 1mo ago Code-MUE: Measuring Code LLMs' Uncertainty through Execution-based Semantic Interaction Graphs arXiv:2607.12273v1 Announce Type: cross Abstract: As Code Large Language Models (LLMs) become central to modern software engineering, their inherent stochasticity poses significant real-world risks, where even minor errors can lead to severe functional, security, or safety… 30 arXiv — NLP / Computation & Language research 1mo ago From Prompt Risk to Response Risk: Paired Analysis of Safety Behavior of Large Language Models arXiv:2604.26052v4 Announce Type: replace Abstract: Safety evaluations of large language models (LLMs) typically report binary outcomes, i.e. attack success rate (ASR), refusal rate, or harmful versus safe classification, which hide how risk changes between prompt and response.… 37 Marcus on AI community 1mo ago Breaking: Demis Hassabis endorses preflight safety testing for AI Good news, for once. 19 Hugging Face Daily Papers research 1mo ago Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals Abstract Post-training is essential for refining the domain-specific capabilities of large language models (LLMs), yet existing reward optimization and distribution matching methods tightly couple policy exploration with distribution alignment. This coupling forces expensive… 36 arXiv — Machine Learning research 1mo ago Safe responses matter: Output-aware safety guardrail mitigate over-refusal in MLLMs arXiv:2607.09697v1 Announce Type: new Abstract: Existing safety mechanisms for multimodal large language models (MLLMs) face a fundamental trade-off between safety and utility. Model fine-tuning achieves robust safety but compromises general utility. Input-side safety guardrails… 16 arXiv — Machine Learning research 1mo ago SMETA-ZSL:Semantic Meta-Alignment for Zero-Shot Threat Classification arXiv:2607.09936v1 Announce Type: new Abstract: Cybersecurity systems must adapt rapidly to emerging threats. However, labeled data for new threat categories is unavailable when those threats first appear. Generalized zero-shot learning offers a natural solution by enabling… 28 arXiv — NLP / Computation & Language research 1mo ago Demographic Prompting at Scale: When More Attributes Hurt LLM--Human Agreement arXiv:2607.10590v1 Announce Type: new Abstract: We investigate how annotator demographic attributes, supplied as prompt cues, shape the alignment between large language model (LLM) predictions and human annotations across five tasks. Using five open-source LLMs, we… 26 arXiv — NLP / Computation & Language research 1mo ago MJ: Multi-turn LLM Jailbreaking via Decomposed Credit Assignment arXiv:2607.11070v1 Announce Type: new Abstract: Modern large language models (LLMs) operate in interactive multi-turn settings, making multi-turn jailbreaking a realistic threat model and an important setting for automated red teaming. A core challenge in learning multi-turn… 4 arXiv — NLP / Computation & Language research 1mo ago TIGER: Text-Conditioned Visual Gated Routing with Acceptance Alignment for Multimodal Speculative Decoding arXiv:2607.11131v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive generation by letting a lightweight drafter propose multiple tokens that are verified by a larger target model. Although effective for text-only LLMs, speculative decoding yields… 24 arXiv — NLP / Computation & Language research 1mo ago Direct Image-to-Modern Vietnamese Translation of Han-Nom Manuscripts via Multimodal RLHF Preference Alignment arXiv:2607.11434v1 Announce Type: new Abstract: Translating Han-Nom manuscripts into modern Vietnamese is challenging because historical pages are often degraded, the script contains rare logographic characters, and parallel supervision is limited. We propose a multimodal RLHF… 11 arXiv — NLP / Computation & Language research 1mo ago Beyond Benchmarks: Exposing the Hidden Crisis in Bangla Hate Speech Detection arXiv:2607.11597v1 Announce Type: new Abstract: The spread of hate speech (HS) across different social media platforms (SMPs) poses a major concern for online safety and ethical moderation. Automatic detection of HS remains a challenging task, especially in under-resourced… 10 arXiv — NLP / Computation & Language research 1mo ago Question Type, Cognitive Load, and CEFR Alignment: Evaluating LLM-Generated EFL Grammar Drill Exercises arXiv:2606.01592v2 Announce Type: cross Abstract: This study evaluates the pedagogical viability of LLM-generated English as a Foreign Language (EFL) learning content. Utilising log data from Japanese junior high school students practicing on a grammar drilling application, we… 12 arXiv — Machine Learning research 1mo ago Reward Transport: Property Control in Flow Matching via Noise-Space Alignment arXiv:2607.08781v1 Announce Type: new Abstract: The coupling in flow matching -- the rule pairing noise vectors with data points -- is typically treated as a computational choice. We show that this coupling can instead serve as an alignment interface: by matching noise and data… 4 arXiv — Machine Learning research 1mo ago Optimizing Against Safety Representations: Activation-Guided Adversarial Suffixes and the Geometry of Refusal arXiv:2607.08883v1 Announce Type: new Abstract: Behavioral alignment in large language models often masks fragile internal safety representations. Recent work suggests that refusal behavior is mediated by low-dimensional directions in activation space. This raises questions… 13 arXiv — Machine Learning research 1mo ago Forget Narrowly, Retain Broadly: Unlearning as an Asymmetric Generalization Problem arXiv:2607.09236v1 Announce Type: new Abstract: Machine unlearning in LLMs is the targeted removal of specific knowledge while preserving all other capabilities, critical for privacy and safety. Yet existing benchmarks measure it unreliably. They miss knowledge that resurfaces… 27 arXiv — Machine Learning research 1mo ago Interval Certifications for Multilayered Perceptrons via Lattice Traversal arXiv:2607.08773v1 Announce Type: cross Abstract: In this work we present a rigorous theoretical framework to a foundational problem of AI safety, namely adversarial robustness. In particular, we show that the adversarial robustness problem can be reduced to a lattice traversal… 34 arXiv — NLP / Computation & Language research 1mo ago An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon? arXiv:2607.09053v1 Announce Type: new Abstract: Recent work has reported Emergent Misalignment (EM), where language models fine-tuned on narrow, domain-specific misaligned datasets abruptly acquire broadly misaligned behavior, alongside evidence that this behavior can be… 14 arXiv — NLP / Computation & Language research 1mo ago VTaMo: Video-Text Alignment Model for Sign Language Translation arXiv:2607.09126v1 Announce Type: cross Abstract: Sign language translation (SLT) converts continuous sign videos into spoken language text. Gloss-free approaches leverage pre-trained visual encoders and language models but rely on implicit cross-modal alignment from translation… 9 arXiv — NLP / Computation & Language research 1mo ago Entity Alignment Method of Science and Technology Patent based on Graph Convolution Network and Information Fusion arXiv:2311.00300v2 Announce Type: replace Abstract: The entity alignment of science and technology patents aims to link the equivalent entities in the knowledge graph of different science and technology patent data sources. Most entity alignment methods only use graph neural… 19 Don't Worry About the Vase community 1mo ago AI #176 Part 2: Plan B This is part 2 of the weekly, broadly covering speculation, rhetoric and policy, along with alignment research. 16 arXiv — NLP / Computation & Language research 1mo ago Efficient Safety Alignment of Language Models via Latent Personality Traits arXiv:2607.07918v1 Announce Type: cross Abstract: Current safety methods for large language models are known to be vulnerable to adversarial attacks, motivating research into robust alternatives. Latent Adversarial Training (LAT) is among the most effective defenses, but can… 15 arXiv — Machine Learning research 1mo ago Who Analyses the Analyser? Self-Validating LLM Hazard Analysis with Constitutional Meta-STPA arXiv:2607.08054v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly trusted to draft the artifacts of safety analysis such as, losses, hazards, Unsafe Control Actions (UCAs), and safety constraints, inside rigorous processes such as Systems-Theoretic… 15 arXiv — Machine Learning research 1mo ago CAAD: Causality-Aware Multivariate Time Series Anomaly Detection via Multi-Scale Alignment and Structural Causal Consistency arXiv:2607.08555v1 Announce Type: new Abstract: The operational integrity of complex industrial systems relies on precise anomaly detection and diagnosis. The vast majority of existing methods narrowly focus on capturing temporal similarities of representations, often… 12 arXiv — Machine Learning research 1mo ago Contravariance Theory: Strong Alignment for Minimal Solutions to Hard Tasks arXiv:2607.08561v1 Announce Type: new Abstract: A series of results from the NeuroAI over the past fifteen years have raised core questions both about how to compare Deep Neural Network (DNN) models to the brain, and about how much convergent evolution to expect between… 22 arXiv — NLP / Computation & Language research 1mo ago PLURAL: A Global Dataset for Value Alignment arXiv:2607.08034v1 Announce Type: new Abstract: Large language models (LLMs) are used worldwide, yet disproportionately reflect Western values, limiting their ability to represent diverse value systems. We introduce PLURAL, a large-scale, value-focused preference dataset… 14 arXiv — NLP / Computation & Language research 1mo ago Best-of-$N$ TTS Evaluation is Confounded by ASR Family Alignment arXiv:2607.08256v1 Announce Type: new Abstract: Best-of-$N$ (BoN) inference improves content consistency in zero-shot text-to-speech by selecting from $N$ candidates with an automatic speech recognition (ASR) verifier. We identify an underexplored evaluation confound: a… 18 arXiv — NLP / Computation & Language research 1mo ago Large-Language-Models-as-a-Judge in Theory-Agnostic Adaptive Metric-Alignment for Prototypical Networks in Personality Recognition arXiv:2607.08374v1 Announce Type: new Abstract: Personality recognition has traditionally been constrained by theory-dependent formulations, where models are trained to fit predefined psychological taxonomies rather than uncovering shared underlying behavioral structure. This… 19 Hacker News — AI on Front Page community 1mo ago GPT-5.6 https://deploymentsafety.openai.com/gpt-5-6/gpt-5-6.pdf https://developers.openai.com/api/docs/guides/latest-model https://x.com/levie/status/2075287443411222628 , https://xcancel.com/levie/status/2075287443411222628 Comments URL: https://news.ycombinator.com/item?id=48849066… 27 Hugging Face Daily Papers research 1mo ago Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs Abstract Splash is a mask-isolated tactile alignment learning framework that enables multimodal LLMs to acquire tactile sensing capabilities without sacrificing vision-language reasoning through selective parameter updating that prevents catastrophic forgetting. Generated by… 22 arXiv — Machine Learning research 1mo ago Reward Valuation in Vision Language Models: Causal Mechanisms Underlying Anhedonia arXiv:2607.06626v1 Announce Type: new Abstract: Recent Vision-Language Models capture increasingly complex aspects of human cognition. Here we ask whether this alignment extends to reward valuation, which we assess in a mechanistic framework built on clinical tests that were… 32 arXiv — Machine Learning research 1mo ago When Certificates Fail: A Unified Safety Framework for Embedded Neural Interface Models arXiv:2607.06630v1 Announce Type: new Abstract: Formal robustness certificates for embedded neural-interface models can pass while task accuracy collapses: at perturbation budget e=0.25, EEGNet classification accuracy drops by 25.7% under projected-gradient attack while the… 21 arXiv — Machine Learning research 1mo ago Online Data Selection Is Implicit Alignment arXiv:2607.07023v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) is often treated as a capability-adaptation step, while alignment is attributed to later preference optimization or reinforcement learning. This separation is incomplete: when examples are scored and… 33 arXiv — Machine Learning research 1mo ago A knowledge-augmented dataset of high-risk driving scenarios with LLM annotations for autonomous driving arXiv:2607.07103v1 Announce Type: new Abstract: Safe autonomous driving requires both rapid responses to common high-risk events and deeper reasoning over rare, extreme long-tail scenarios in traffic safety. These scenarios are severely under-represented in naturalistic driving… 6 arXiv — Machine Learning research 1mo ago Predicting LLM Safety Before Release by Simulating Deployment arXiv:2607.07184v1 Announce Type: new Abstract: Pre-deployment safety evaluations aim to inform the downstream risks of releasing a new AI model. Yet most evaluations provide limited evidence about how often undesired model behavior will occur in deployment: they generally have… 31 arXiv — Machine Learning research 1mo ago Avoiding unsafe sets when training with Langevin Dynamics arXiv:2607.07538v1 Announce Type: new Abstract: Training a model with noisy gradient descent can be idealized as overdamped Langevin dynamics on the loss landscape, and a natural safety question is to bound the probability $\nu_t(\mathcal{A}_H) = \mathbb{P}(Q_t \in… 16 Page 9 of 10 · 500 articles ← Newer Older →