News / #safety Tag Safety + alignment 500 articles archived under #safety · RSS Sign in to follow arXiv — NLP / Computation & Language research 1mo ago Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs arXiv:2607.06831v1 Announce Type: new Abstract: Speech-to-text alignment means finding the temporal boundaries of each word in the audio. Some models provide such an alignment directly and others do not. Connectionist temporal classification (CTC) and transducer models have an… 14 arXiv — NLP / Computation & Language research 1mo ago Riemannian Geometry for Pre-trained Language Model Embeddings arXiv:2607.07047v1 Announce Type: new Abstract: Understanding the geometric structure of pre-trained language model embeddings matters for interpretability and safety. We ask whether sentence-level classification signal lives in the Riemannian geometry of contextual token… 21 arXiv — NLP / Computation & Language research 1mo ago R^3: Advertisement Compliance Rectification via Group-Relative Experience Extractor and Curriculum Reinforcement arXiv:2607.07318v1 Announce Type: new Abstract: Rigorous content moderation is crucial for online advertising but leads to millions of daily rejections. This scale renders manual rectification infeasible, particularly for video advertisements. However, existing safety-driven… 28 arXiv — NLP / Computation & Language research 1mo ago Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents arXiv:2607.07474v1 Announce Type: cross Abstract: Agentic red-teaming benchmarks report whether an injected agent was compromised as a single bit: the attack succeeded, or it did not. We argue that this binary attack-success rate discards the information a defender most needs,… 20 arXiv — NLP / Computation & Language research 1mo ago $C$-$\Delta\Theta$: Circuit-Restricted Weight Arithmetic for Selective Refusal arXiv:2602.04521v2 Announce Type: replace Abstract: Modern deployments require LLMs to enforce safety policies at scale, yet many controls rely on inference-time interventions that add recurring compute cost and serving complexity. Activation steering is widely used, but it… 38 arXiv — NLP / Computation & Language research 1mo ago Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection arXiv:2604.07831v2 Announce Type: replace-cross Abstract: Existing red-teaming studies on GUI agents face two fundamental limitations: adversarial perturbations require white-box access unavailable in commercial deployments, while prompt injection is increasingly neutralized by… 37 r/MachineLearning community 1mo ago Agentic safety triggers aren't textual safety triggers — MCP attacks that beat SOTA guardrails more than half the time (code + dataset) [R] Most safety alignment work treats "detect the attack" as a text classification problem — does the prompt contain language the model's safety guardrails should catch. That assumption breaks down for LLM agents with real tool access. Here's a concrete case: take a known, public… 14 OpenAI official-blog 1mo ago Our approach to government and national security partnerships Learn how OpenAI approaches government and national security partnerships, with principles for responsible AI use, democratic accountability, and public safety. 8 arXiv — Machine Learning research 1mo ago TILDE: TILt-based Distributional Erasure for Concept Unlearning arXiv:2607.06432v1 Announce Type: new Abstract: Concept unlearning in text-to-image diffusion models is critical for safe and practical deployment: with rising privacy concerns, copyright disputes, trademark constraints, and safety regulations, deployed systems must be able to… 22 arXiv — NLP / Computation & Language research 1mo ago Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability arXiv:2607.06196v1 Announce Type: new Abstract: Current AI safety evaluation and benchmarking frameworks predominantly rely on Western-centric culture-agnostic defaults that mask critical regional laws, socio-linguistic nuances, and cultural taboos, leaving Vision-Language… 7 arXiv — NLP / Computation & Language research 1mo ago PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails arXiv:2607.05910v1 Announce Type: cross Abstract: Image guardrails are typically trained and evaluated under a fixed safety policy, implicitly treating safety as an intrinsic property of an image. Real deployments are different: the same image may be allowed in one product,… 27 arXiv — NLP / Computation & Language research 1mo ago Decoding the Multimodal Mind: Generalizable Brain-to-Text Translation via Multimodal Alignment and Adaptive Routing arXiv:2505.10356v3 Announce Type: replace Abstract: Decoding language from the human brain remains a grand challenge for Brain-Computer Interfaces (BCIs). Current approaches typically rely on unimodal brain representations, neglecting the brain's inherently multimodal… 31 arXiv — NLP / Computation & Language research 1mo ago Quantifying Retriever-Generator Alignment in RAG with Local Explanations arXiv:2601.21803v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) systems combine dense retrievers and language models to ground their outputs in external documents. However, the interaction between these components remains opaque, creating challenges for… 8 arXiv — NLP / Computation & Language research 1mo ago Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis arXiv:2602.00846v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) struggle with alignment due to the limitations of existing reward models (RMs), which are predominantly vision-centric, dependent on costly human labels, and provide opaque scalar scores… 13 arXiv — NLP / Computation & Language research 1mo ago Geometric Stability: The Missing Axis of Representations arXiv:2601.09173v5 Announce Type: replace-cross Abstract: Representational similarity analysis and related methods compare the internal geometries of neural networks, but they measure only alignment between spaces, leaving a blind spot -- whether a representation's structure is… 15 Hugging Face Daily Papers research 1mo ago Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment Abstract A supervised contrastive alignment framework maps WavLM embeddings from English and Mandarin into a shared clinical space for depression detection, addressing cross-lingual generalization challenges and revealing performance artifacts caused by speaker identity leakage.… 38 Hugging Face Daily Papers research 1mo ago Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory Abstract Light-Omni is a multimodal agent framework that enables efficient video understanding through dual contextual states, achieving faster and more accurate video processing by eliminating iterative reasoning while maintaining semantic alignment. Generated by… 14 r/MachineLearning community 1mo ago Mid research got me thinking what about reversed alignment, would trained "bad" model exhibit"good" behavior later and/or secretly [D] late night thoughts as I was working on my paper that is about specific behavior that arises from RHLF, it got me thinking what if train a model in an environment where bad behavior is rewarded: deception, selfishness, harmful behavior etc. and then find it occasionally and/or… 23 Hugging Face Daily Papers research 1mo ago PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space Abstract PixWorld presents a unified pixel-space diffusion approach for 3D reconstruction and generation that overcomes limitations of latent-space methods through direct image-level supervision and geometry-aware feature alignment. Generated by Qwen/Qwen2.5-Coder-32B-Instruct… 19 arXiv — Machine Learning research 1mo ago Federated Learning for Object Detection: Enabling Collaborative Drone Learning Without Centralizing Data arXiv:2607.02636v1 Announce Type: new Abstract: Object detection is a fundamental capability for AI-driven perception in safety-critical drone and edge-vision systems, including disaster response, operational security environments, infrastructure monitoring and defense… 31 arXiv — Machine Learning research 1mo ago Safe Inference-Time Alignment via Lagrangian Reward Augmentation arXiv:2607.02781v1 Announce Type: new Abstract: Inference-time alignment steers a frozen language model during decoding using auxiliary reward signals, avoiding the cost of repeated weight updates. However, existing inference-time alignment methods typically optimize a single… 28 arXiv — Machine Learning research 1mo ago Bootstrap Flow-Map Tree Sampling Enables Online Feedback Driven Search arXiv:2607.02915v1 Announce Type: new Abstract: In many scientific and engineering domains, maximizing discovery within a limited sampling budget demands strategic, observation-guided exploration. While generative models have enabled training-free reward alignment, current… 27 arXiv — Machine Learning research 1mo ago Robustness Meets Uncertainty: Evidential Adversarial Training for Robust Selective Classification arXiv:2607.03075v1 Announce Type: new Abstract: Safety-critical applications require classifiers that are both robust and reliable. Adversarial training is a widely adopted defense for improving robustness in deep neural networks; however, its effect on the reliability of… 34 arXiv — Machine Learning research 1mo ago Integrating Physics-Informed Neural Networks for Safe Reinforcement Learning in a 1-DoF Helicopter System arXiv:2607.03125v1 Announce Type: new Abstract: Deep reinforcement learning (DRL) offers powerful control for industrial cyber-physical systems (ICPSs), but its "black-box" exploration risks violating strict hardware safety limits. Typically, these constraints are managed… 35 arXiv — Machine Learning research 1mo ago Unbiased Alignment for Large Language Models with Noisy Preferences arXiv:2607.03248v1 Announce Type: new Abstract: The alignment of large language models with human preferences is commonly achieved through Reinforcement Learning from Human Feedback or Direct Preference Optimization. However, these methods are vulnerable to the significant noise… 37 arXiv — Machine Learning research 1mo ago Best-of-Better-$N$: Generating Pre-Aligned Responses with In-Context Learning arXiv:2607.03453v1 Announce Type: new Abstract: Inference-time alignment methods, such as Best-of-$N$, offer a flexible alternative to training-based alignment by using reward models to select high-quality responses generated by a reference LLM. However, the efficacy of these… 15 arXiv — NLP / Computation & Language research 1mo ago Improving LLMs via Validator-to-Generator Alignment arXiv:2607.02668v1 Announce Type: new Abstract: Large language models are inconsistent: varying prompts or including unrelated information can lead to unexpected changes in model outputs. The generator-validator (G-V) gap is one manifestation of this phenomenon, where LLMs… 26 arXiv — NLP / Computation & Language research 1mo ago Alignment-Guided Largest Table Overlap Size Estimation arXiv:2607.03049v1 Announce Type: new Abstract: Fast estimation of the size of the largest overlap between tables enables blocking and query-by-table retrieval in large table repositories. The first and the state-of-the-art estimator Armadillo improves efficiency by embedding… 31 arXiv — NLP / Computation & Language research 1mo ago KARMA: Knowledge graph-based Automated Reasoning Materialization and Alignment arXiv:2607.03166v1 Announce Type: new Abstract: Template-based contrastive synthesis is scalable, but its candidates often differ only in a few entity-slots while sequence-level optimization spreads supervision over mostly shared templates. We formalize this as the Resolution… 10 arXiv — NLP / Computation & Language research 1mo ago Optimizing Large Language Models for Causality Assessment in Pharmacovigilance: Developing a Performance Metric as Objective for Bayesian Hyperparameter Optimization arXiv:2607.03704v1 Announce Type: new Abstract: Background: Growing individual case safety report (ICSR) volumes have intensified demand for scalable automated causality assessment. Large Language Models (LLMs) show promise, yet performance on clinically demanding tasks remains… 29 arXiv — NLP / Computation & Language research 1mo ago Uncertainty-Aware Abstention in Large Language Models with Provable Alignment Guarantees arXiv:2607.04430v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in question answering (QA) systems, yet they may generate hallucinated or misaligned responses without reliable confidence estimates. Uncertainty quantification (UQ) offers a… 34 arXiv — NLP / Computation & Language research 1mo ago Transplanting, inverting, and preventing a misalignment persona: method-conditional emergent misalignment in Qwen2.5 arXiv:2607.04510v1 Announce Type: new Abstract: Emergent misalignment (EM) -- the broad misbehaviour a language model acquires after fine-tuning on narrow harmful data -- is mediated in Qwen2.5 models by a latent persona direction, and that direction is causal in open weights.… 7 arXiv — NLP / Computation & Language research 1mo ago Retroactive Chain-of-Thought (RetroCoT): Forensic Reconstruction Prompts as a Safety Diagnostic Across Model Generations arXiv:2607.04645v1 Announce Type: new Abstract: Safety alignment in large language models is typically evaluated against direct, imperative harmful requests. We show that this alignment is highly conditioned on pragmatic register: models that refuse a direct request frequently… 35 arXiv — NLP / Computation & Language research 1mo ago FormalRx: Rectify and eXamine Semantic Failures in Autoformalization arXiv:2607.04655v1 Announce Type: new Abstract: The veracious semantic alignment in autoformalization is significant for formal mathematical reasoning. However, existing evaluations provide only opaque binary verdicts or scalar scores, offering no interpretable insight into… 34 arXiv — NLP / Computation & Language research 1mo ago Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment arXiv:2607.04728v1 Announce Type: new Abstract: Reinforcement learning (RL) post-training for large language models (LLMs) follows a efficient paradigm of "rollout then update", which inevitably results in off-policy training data. To resolve this, Importance sampling (IS) is… 36 Hugging Face Daily Papers research 1mo ago Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification Abstract Automated safety testing framework Vera uses a three-stage pipeline to identify and test safety risks in LLM agents through structured risk taxonomies, combinatorial case generation, and adaptive sandbox execution with evidence-based verification. Generated by… 7 r/LocalLLaMA community 1mo ago ThinkingCap-Qwen3.6-27B: same accuracy as base Qwen3.6 with ~50% fewer thinking We rigorously evaluate the resulting checkpoint across general reasoning, non-reasoning multiple-choice question answering, everyday multi-turn conversations, system prompt adherence, safety, math, code and agentic use cases. Due to the high variability of reasoning quality at… 12 Hugging Face Daily Papers research 1mo ago Interpretation-Oriented Cloud Removal via Observation-Anchored Residual Flow with Geo-Contextual Alignment Abstract Geo-Anchored Cloud Removal framework addresses semantic drift in cloud removal by combining physically grounded residual inversion with semantic manifold constraints from vision foundation models. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Cloud removal (CR) is… 29 Hugging Face Daily Papers research 1mo ago Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teaming Abstract AI-Infra-Guard is an open-source framework that addresses AI infrastructure security through layered detection paradigms spanning infrastructure, protocol, agent behavior, and model layers. Generated by Qwen/Qwen2.5-Coder-32B-Instruct The fast growth of open-source AI… 4 r/MachineLearning community 1mo ago Best models for generating red-team attacks? Also looking for public datasets [R] Hi everyone, I'm currently working on a framework to evaluate the security of LLM applications and AI agents, and I've been stuck on one part for a while. Most red-teaming frameworks rely on an LLM to generate adversarial prompts. My question is more about which model to use .… 29 r/MachineLearning community 1mo ago What does "Safe AI" look like? [D] ​ For open-weight LLMs, how practical is it to study defenses against post-release fine-tuning that weakens refusal or safety behavior? I've been seeing “uncensored” or “heretic” variants of new models appear very quickly after release, which raises a question I’m curious… 28 Hugging Face Daily Papers research 1mo ago Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR Abstract Transfer-Aware Curriculum (TAC) improves multi-domain reinforcement learning by prioritizing domains that provide broad benefits to other domains, using gradient-geometry alignment to estimate cross-domain transferability. Generated by Qwen/Qwen2.5-Coder-32B-Instruct… 25 arXiv — Machine Learning research 1mo ago IonSense-QKG: A Quantum-Readiness Metadata Framework for Lithium-Ion Battery Dataset Discovery arXiv:2607.01286v1 Announce Type: new Abstract: Public lithium-ion battery datasets are increasingly used for state-of-health estimation, remaining-useful-life prediction, anomaly detection, electrochemical diagnostics, second-life analytics, and battery safety research.… 36 arXiv — Machine Learning research 1mo ago Multi-modal Rail Crossing Safety Analysis arXiv:2607.01365v1 Announce Type: new Abstract: Given one or more images of a railway crossing, can we leverage visual cues that allow us to robustly estimate how safe it is? Can we improve our ability to do so by introducing structured data (such as official accident reports)… 9 arXiv — Machine Learning research 1mo ago CALM: Interpretable Cross-Modal Alignment for Biomarker Discovery from Unpaired Data arXiv:2607.01656v1 Announce Type: new Abstract: The interaction between brain structure and genetic influences is key to understanding neuropsychiatric disorders. However, most large-scale datasets are unimodal, providing either neuroimaging or genetics data. We propose CALM, a… 15 arXiv — NLP / Computation & Language research 1mo ago Safeguarding LLM Agents from Misalignment through Provenance Analysis arXiv:2607.01236v1 Announce Type: new Abstract: As LLM agents gain increasing access to powerful tools, ensuring that their actions are aligned with the user's intent becomes critical. When an agent's proposed tool invocation deviates from the user's intent -- a phenomenon… 17 arXiv — NLP / Computation & Language research 1mo ago Breaking Safety at the Token Boundary: How BPE Tokenization Creates Exploitable Gaps in LLM Alignment arXiv:2607.01239v1 Announce Type: new Abstract: Character-level perturbations bypass safety alignment in modern LLMs despite leaving prompts human-readable. We identify and test a central structural mechanism: BPE tokenization fragments safety-critical words into sub-word… 9 arXiv — NLP / Computation & Language research 1mo ago Multi-Objective Exploration and Preference Optimization via Mutual Information arXiv:2607.01392v1 Announce Type: new Abstract: Aligning large language models with diverse and heterogeneous human values requires multi-objective alignment methods to effectively trade off conflicting preference dimensions. Current methods achieve this trade-off by training… 33 arXiv — NLP / Computation & Language research 1mo ago MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering arXiv:2607.01420v1 Announce Type: new Abstract: As grounded QA systems are increasingly deployed in AI assistants, accurately attributing generated answers to evidence is critical for user trust and model safety. While unimodal attributions have been explored in depth, the… 10 arXiv — NLP / Computation & Language research 1mo ago HaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI Safety arXiv:2607.02079v1 Announce Type: new Abstract: We present HaloGuard 1.0, an open-weights implementation of the constitutional-classifier paradigm for input safety. It achieves state-of-the-art performance on English and multilingual prompt-safety benchmarks at roughly one-tenth… 25 Page 10 of 10 · 500 articles ← Newer