News / #safety Tag Safety + alignment 500 articles archived under #safety · RSS Sign in to follow Hugging Face Daily Papers research 24d ago Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming Abstract Prompt injection poses significant security risks to LLM agents. Efficient and effective red-teaming is therefore critical, both for evaluating these risks and for collecting training data to improve defenses. Existing state-of-the-art prompt injection red-teaming… 25 arXiv — Machine Learning research 24d ago Looking in the Mirror: Introspecting Side-Effect Misalignments Induced by Fine-Tuning arXiv:2608.04347v1 Announce Type: new Abstract: Fine-tuning enables a source model to acquire desired capabilities and behaviors in a target domain while retaining much of its general-purpose competence. However, this adaptation process can also degrade alignment properties that… 35 arXiv — Machine Learning research 24d ago Local Violation Certification for Linear Predict-Then-Optimize Pipelines arXiv:2608.04474v1 Announce Type: new Abstract: Data-driven decision pipelines combining predictive machine learning models with downstream optimization software are increasingly used to make high-stakes operational decisions. Certifying the safety, fairness, and reliability of… 30 arXiv — Machine Learning research 24d ago Why Ranking Anomaly Detection Algorithms Isn't as Reliable as You May Think arXiv:2608.04613v1 Announce Type: new Abstract: Anomaly detection is a safety-critical machine learning problem with applications ranging from fraud detection to network intrusion prevention and industrial monitoring. Despite the large number of proposed anomaly detection… 30 arXiv — NLP / Computation & Language research 24d ago DataRx: Missingness-Aware Sampling for Safer Large Language Model Task-Specific Fine-Tuning arXiv:2608.04322v1 Announce Type: new Abstract: Task-specific fine-tuning can improve the performance of large language models (LLMs) on downstream tasks. However, our study reveals that task-specific fine-tuning can also weaken the safety guardrails of aligned LLMs. A widely… 27 arXiv — NLP / Computation & Language research 24d ago Social Pressure Breaks Majority Voting in LLM Safety Panels arXiv:2608.04415v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to detect unsafe content. A common approach is to combine judgments from a panel of models to correct individual mistakes, but this benefit may disappear when every model sees the… 36 arXiv — NLP / Computation & Language research 24d ago EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment arXiv:2608.04472v1 Announce Type: cross Abstract: The development of foundation models (FMs) is crucial for advancing endoscopic image analysis. However, existing endoscopy FMs mainly rely on self-supervised learning from uni-modal images or videos, overlooking the rich semantic… 18 Simon Willison community 24d ago Third-party cyber evaluations involving OpenAI models Third-party cyber evaluations involving OpenAI models And another one . I had to create a accidental-cyberattacks tag to keep track of them all! This post from OpenAI covers both the UK AI Safety Institute attack (see my previous post ) and another attack enabled by Irregular :… 22 Simon Willison community 24d ago Third-party cyber evaluations involving OpenAI models Third-party cyber evaluations involving OpenAI models And another one . I had to create a accidental-cyberattacks tag to keep track of them all! This post from OpenAI covers both the UK AI Safety Institute attack (see my previous post ) and another attack enabled by Irregular :… 19 Simon Willison community 24d ago Incident Report: unsanctioned agent behaviour during cyber testing Incident Report: unsanctioned agent behaviour during cyber testing It happened again . This time it was the UK government's AI Security Institute who accidentally attacked other companies while running an evaluation with models with the safety filters turned off. From their… 37 Simon Willison community 24d ago Incident Report: unsanctioned agent behaviour during cyber testing Incident Report: unsanctioned agent behaviour during cyber testing It happened again . This time it was the UK government's AI Security Institute who accidentally attacked other companies while running an evaluation with models with the safety filters turned off. From their… 14 arXiv — Machine Learning research 25d ago Learning Molecular Representations from Cellular Phenotypes with Structure Preservation arXiv:2608.02688v1 Announce Type: new Abstract: Phenotypic drug discovery enables the discovery of functional relationships between molecular structures and cellular responses. However, existing multimodal representation learning methods often optimize cross-modal alignment… 11 arXiv — NLP / Computation & Language research 25d ago Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety arXiv:2608.02617v1 Announce Type: new Abstract: We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert feedback from MOOVE (Massive Open Online Validation and Evaluation), a… 15 arXiv — NLP / Computation & Language research 25d ago Aligned in Form, Not in Meaning: The Comprehension - Containment Decoupling of LLM Safety in Low-Resource Bangla Derogatory Speech arXiv:2608.02941v1 Announce Type: new Abstract: We audit five frontier large language models on native Bangla derogatory speech (gali) across six protocols to test a single hypothesis: Comprehension-Containment Decoupling. We propose that contemporary safety alignment is bound… 7 arXiv — NLP / Computation & Language research 25d ago Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation arXiv:2608.03044v1 Announce Type: new Abstract: Large language models are increasingly used to simulate human opinions, but prior work reports conflicting results: some studies find promising alignment with human survey data, while others find persona collapse and weak… 38 arXiv — NLP / Computation & Language research 25d ago HomoEnsNER: Does Language Alignment Outperform Architectural Complexity in Gujarati Named Entity Recognition? arXiv:2608.03105v1 Announce Type: new Abstract: Named Entity Recognition (NER) for Gujarati remains underexplored, hindered by the absence of capitalization cues, rich morphology, lexical ambiguity, and free word order. Prior ensemble work has emphasized architectural diversity… 9 arXiv — NLP / Computation & Language research 25d ago Aligning Large Vision-Language Models at Test Time: A Trajectory-Guided Structured Sampling Approach arXiv:2608.03204v1 Announce Type: new Abstract: Post-training reinforcement learning (RL) algorithms are commonly used to align large vision-language models (LVLMs) with human intent and the requirements of visual reasoning tasks. However, existing RL-based alignment methods are… 18 arXiv — NLP / Computation & Language research 25d ago ICO: Enhancing Semantic-Shift Jailbreaks via Iterative Context Optimization arXiv:2608.03210v1 Announce Type: new Abstract: Foundation models have achieved remarkable success across diverse tasks, but they remain vulnerable. To investigate such vulnerabilities, semantic-shift jailbreaks have recently emerged as a promising attack paradigm. They bypass… 22 arXiv — NLP / Computation & Language research 25d ago Predicting Multilingual Classification and Translation Performance of LLMs with Cross-Lingual Alignment $\unicode{x2013}$ Is English Enough? arXiv:2608.03446v1 Announce Type: new Abstract: Multilingual large language models (LLMs) have been shown to perform better on non-English classification tasks when the representations of the given language are more aligned to English within the model. Several cross-lingual… 37 arXiv — NLP / Computation & Language research 25d ago Cross-Lingual Bias in Large Language Models: A Comparative Analysis of English and Swahili arXiv:2608.03532v1 Announce Type: new Abstract: Large language models are increasingly deployed in multilingual contexts, yet safety alignment and bias evaluation remain overwhelmingly English-centric. We investigate whether social biases generalise across languages by… 5 arXiv — NLP / Computation & Language research 25d ago Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity arXiv:2608.02665v1 Announce Type: cross Abstract: A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that reading is faithful: when an item's intent is held fixed and only its meaning-preserving… 38 r/LocalLLaMA community 25d ago China’s Open-Weight Models Will Be Spared US Safety Tests   submitted by   /u/fallingdowndizzyvr [link]   [comments] 36 TechCrunch — AI news-outlet 25d ago Open-weight AI models are catching up to the frontier. The safety gap remains. A new SaferAI report finds Z.ai's open-weight GLM-5.2 approaches frontier AI capabilities while lacking key safety mitigations, renewing concerns that powerful open models could outpace governance and safeguards. 11 Smol AI News news-outlet 26d ago not much happened today **Alibaba** launched **Qwen3.8-Max**, enhancing multimodal capabilities and agent ecosystem integration. **NVIDIA** introduced **Alpamayo 2 Super** for autonomous vehicle reasoning, while **Mistral AI** released **Shieldstral**, a 3B parameter open-weights safety model for… 17 arXiv — Machine Learning research 26d ago Inference-Time Policy Alignment for Fair Reinforcement Learning arXiv:2608.00175v1 Announce Type: new Abstract: Deep reinforcement learning (RL) agents achieve strong performance by optimizing scalar reward functions. However, once deployed, the policies of these RL agents are often rigid and costly to adapt to new performance criteria. For… 22 arXiv — Machine Learning research 26d ago Neural operator learning for collision-aware trajectory planning of spacecraft swarms arXiv:2608.00320v1 Announce Type: new Abstract: Autonomous spacecraft swarms must plan fuel-efficient, collision-free maneuvers in increasingly congested orbits, yet classical trajectory optimization scales poorly as pairwise safety constraints multiply with swarm size, and… 37 arXiv — NLP / Computation & Language research 26d ago A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard) arXiv:2608.00180v1 Announce Type: new Abstract: Conflicting objectives are general in RL alignment, and training on them data-efficiently is hard. Training a safety guard with RL means optimizing two objectives that conflict: catch real harm, and do not refuse benign prompts.… 26 arXiv — NLP / Computation & Language research 26d ago OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution arXiv:2608.00677v1 Announce Type: new Abstract: AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-model interactions, agent behavior is mediated through a shared state that is… 14 arXiv — NLP / Computation & Language research 26d ago Mind the Gap: Zero-Query Jailbreaks via Filter-Generator Discrepancy in Text-to-Image Systems arXiv:2608.00973v1 Announce Type: new Abstract: Text-to-image (T2I) systems typically have prompt-level safety filters before the generator to block unsafe requests, yet such systems remain vulnerable to malicious jailbreak prompts. Transfer-based attacks construct adversarial… 6 arXiv — NLP / Computation & Language research 26d ago ArabicDialectSafety: A Dialect-Aware Benchmark for Arabic Content Safety Classification arXiv:2608.01291v1 Announce Type: new Abstract: We present ArabicDialectSafety, a human-curated Arabic safety dataset of 25,071 prompts covering six Arabic varieties: Modern Standard Arabic, Syrian, Egyptian, Algerian, Palestinian, and Moroccan. The dataset is annotated with… 19 arXiv — NLP / Computation & Language research 26d ago Semantic Alignment of AI Models: Concept Collapse, Checkpoint Dynamics, and Cross-Lingual Transfer arXiv:2608.01585v1 Announce Type: new Abstract: Language model benchmarking is a difficult task. Outcome reasoning alone does not test the model's conceptualization of language and popular open-source benchmarks are quickly saturated or ingested as training data. It is important… 7 arXiv — NLP / Computation & Language research 26d ago Human-LLM Alignment in Language Attitudes Toward Non-Native Japanese arXiv:2608.01629v1 Announce Type: new Abstract: Large language models (LLMs) increasingly evaluate human writing in high-stakes domains such as hiring and academic assessment, putting non-native speakers at particular risk. Drawing on the language attitudes framework, we… 30 Hugging Face Daily Papers research 27d ago In the Driver's Seat: A Multi-Company Study on the Reality of Autonomous Driving System Testing Abstract Autonomous driving systems (ADS) are rapidly advancing and increasingly deployed in real-world applications. This creates growing demands for effective testing to ensure system functionality and safety. However, ADS testing remains complex and lacks well-established… 18 Hugging Face Daily Papers research 27d ago SULAND v2: A Refined RGB Dataset and Deep Learning Object Detection Benchmark for UAV/UGV-Based SUrface LANDmine Detection Under Domain Shift Abstract RGB imagery offers a practical, low-cost option for Unmanned Aerial/Ground Vehicle (UAV/UGV) survey support in surface-landmine detection, but object detectors remain underexplored in this safety-critical domain. Limited cross-architecture benchmarking and insufficient… 32 Hugging Face Daily Papers research 27d ago Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs Abstract Large language model safeguards decide whether to answer before seeing how an answer will be used. This creates a basic problem for dual-use tasks: the same answer can help an authorized professional or an attacker, while an attacker can imitate a benign request and… 8 arXiv — Machine Learning research 27d ago LARA: Lightweight Adapters in the Residual Stream for Composable Adaptation and Alignment arXiv:2607.28669v1 Announce Type: new Abstract: We present LARA (Lightweight Additive Residual Adaptation), a method for efficient adaptation that operates in the residual stream of a frozen model rather than in its weights. Where LoRA adds an update of low rank to weight… 36 arXiv — Machine Learning research 27d ago Beyond Feature and Structure Alignment: Learning Transferable Propagation Knowledge for Graph Foundation Models arXiv:2607.28980v1 Announce Type: new Abstract: Graph Foundation Models (GFMs) have recently emerged as a promising paradigm for enabling knowledge transfer across diverse domains. Unlike traditional graph learning methods that are typically designed for in-domain settings, GFMs… 29 arXiv — NLP / Computation & Language research 27d ago Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning arXiv:2607.28986v1 Announce Type: cross Abstract: Zero-shot image captioning (ZIC) describes images without paired image-caption supervision during captioner training, relying on text-only corpora and frozen pretrained image-text scorers. Existing retrieval-augmented methods… 26 OpenAI official-blog 29d ago Advancing responsible AI across Europe OpenAI shares how its safety, security, transparency, and provenance practices support responsible AI governance in Europe. The work will continue as the EU AI Act advances. 17 Hugging Face Daily Papers research 1mo ago Pedestrian Archetypes Extension -- More Pedestrian Models for Autonomous Vehicle Safety Testing Abstract In our prior work, Pedestrian Archetypes, we defined pedestrian archetypes as collections of behaviors that uniquely identify a specific type of pedestrian. The first paper proposed 12 pedestrian archetypes, including the Wanderer, Drunk, Distracted, Flash, Indecisive,… 6 arXiv — Machine Learning research 1mo ago Regularizing modality contribution drift in multimodal continual learning arXiv:2607.27260v1 Announce Type: new Abstract: Multimodal continual learning (MMCL) aims to learn emerging knowledge from multimodal data while preserving knowledge. To mitigate forgetting, current MMCL methods usually focus on cross-modal representation alignment or semantic… 36 arXiv — Machine Learning research 1mo ago The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models arXiv:2607.27281v1 Announce Type: new Abstract: A capability appears in a language model when the last parts of its circuit align in one stochastic attempt, and getting all but one right is worth nothing. We show this no-partial-credit joint alignment is the rate-limiting step… 23 arXiv — Machine Learning research 1mo ago Context-Informed Ship Trajectory Prediction via Conditional Attention arXiv:2607.27418v1 Announce Type: new Abstract: Long-term ship trajectory prediction is a fundamental capability for maritime safety and autonomous navigation. While recent Transformer-based architectures have improved forecasting horizons, they predominantly rely on historical… 10 arXiv — Machine Learning research 1mo ago When Does Explicit View Routing Work? A Controlled Study of Multi-View Graph-Text Alignment arXiv:2607.27530v1 Announce Type: new Abstract: Graph-text retrieval typically maps a graph and its description to a single embedding, even when a query concerns only one semantic aspect, such as a class label or molecular property. Multiple heads can separate these aspects, but… 23 arXiv — Machine Learning research 1mo ago Compliance2LoRA: On-Demand Safety Alignment on Arbitrary Policy Subsets via Hypernetwork-Generated LoRA Adapters arXiv:2607.27594v1 Announce Type: new Abstract: Post-training alignment in large reasoning models (LRMs) has significantly improved their adaptability to diverse safety compliance settings. However, as LRMs personalization for downstream users takes center stage, the demand for… 26 arXiv — Machine Learning research 1mo ago Real-Time Hard Peak Age-of-Information Safety with No-Regret Learning arXiv:2607.27626v1 Announce Type: new Abstract: Safety-critical IoT systems such as industrial closed-loop control, V2X coordination, and remote teleoperation require every sensor's peak Age of Information (peak AoI, also abbreviated PAoI) to stay below a hard per-slot deadline,… 29 arXiv — Machine Learning research 1mo ago DAS-PMVC: A Framework for Partial Multi-View Clustering via Dual Alignment and Structure Enhancement arXiv:2607.27761v1 Announce Type: new Abstract: In recent years, multi-view clustering has attracted widespread research interest. However, due to limitations in data collection devices, data across different views often suffer from misalignment, leading to the partial view… 14 arXiv — NLP / Computation & Language research 1mo ago Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups arXiv:2607.27232v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly shaping how we consume information and form our worldview. This raises concerns beyond bias in AI: do LLMs grasp the emotional nuances conveyed via textual framing? In this work, we… 28 arXiv — NLP / Computation & Language research 1mo ago BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences arXiv:2607.27366v1 Announce Type: new Abstract: While data synthesis for large language models (LLMs) is prevalent, it primarily targets domains with verifiable answers, overlooking open-ended humanities and social sciences (HSS), where nuanced quality judgments matter more than… 22 arXiv — NLP / Computation & Language research 1mo ago Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game arXiv:2607.28146v1 Announce Type: new Abstract: As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabilities is fundamental to safety. Controlled social deduction games provide a… 32 Page 5 of 10 · 500 articles ← Newer Older →