News / #safety Tag Safety + alignment 500 articles archived under #safety · RSS Sign in to follow arXiv — NLP / Computation & Language research 1mo ago CASE: Causal Alignment and Structural Enforcement for Improving Chain-of-Thought Faithfulness arXiv:2607.18820v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning is widely used to improve both the performance and interpretability of large language models (LLMs), yet the generated reasoning may not faithfully support the final answer. We study this problem… 22 arXiv — NLP / Computation & Language research 1mo ago Operational Hallucination and Safety Drift in AI Agents arXiv:2607.18366v1 Announce Type: cross Abstract: Large language models (LLMs) serving as planners in tool-using autonomous agents introduce dynamic reliability risks in multi-turn execution. While single-turn safety mechanisms are relatively mature, extended interactions reveal… 26 arXiv — NLP / Computation & Language research 1mo ago MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications arXiv:2409.07314v3 Announce Type: replace Abstract: While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have become saturated and increasingly disconnected from the functional requirements of clinical… 17 arXiv — NLP / Computation & Language research 1mo ago Guard Vector: Beyond English LLM Guardrails with Task-Vector Composition and Streaming-Aware Prefix SFT arXiv:2509.23381v2 Announce Type: replace Abstract: We introduce Guard Vector, a safety task vector computed as the parameter difference between a guardrail model (Guard Model) and a same-architecture pretrained language model. Composing this vector with a target language model… 28 Don't Worry About the Vase community 1mo ago OpenAI Shares Some Alignment Problems Kudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were forced to take the model offline to work on new mitigations and defense-to-depth. 13 r/LocalLLaMA community 1mo ago New Model: Nanbeige4.2-3B (Looped Transformer, outperforms 4x size) https://huggingface.co/Nanbeige/Nanbeige4.2-3B Nanbeige4.2-3B is a compact agentic model built on Nanbeige4.2-3B-Base , designed to combine strong agentic behavior with broad reasoning and alignment capabilities. Its Looped Transformer architecture reuses the transformer layers… 33 r/LocalLLaMA community 1mo ago Be Careful when Purchasing CMP 170HX on Alibaba! Just a heads up. Shops in China are running like chickens without a head after the news the Falcon Exploit working to jailbreak some of the functions of these cards. Is not just happening on Alibaba but also Ebay. Usually from Chinese sellers. I spent 2 days contacting lost of… 28 arXiv — Machine Learning research 1mo ago LLM Unlearning for Cyber Defense: A Survey on Methods, Challenges, and Emerging Threats arXiv:2607.16227v1 Announce Type: new Abstract: LLMs are increasingly deployed in security-critical systems across healthcare, finance, education, and decision support, yet their inability to forget creates serious cybersecurity, privacy, and safety risks. Sensitive personal… 9 arXiv — Machine Learning research 1mo ago Normalized Rewards for Preference Optimization arXiv:2607.16240v1 Announce Type: new Abstract: Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs with human preferences. However, DAAs have been observed to over-optimize their implicit reward model and decrease the likelihood… 20 arXiv — Machine Learning research 1mo ago TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment arXiv:2607.16242v1 Announce Type: new Abstract: Fine-Tuning-as-a-Service (FTaaS) platforms let users train large language models (LLMs) on customized tasks, but this pipeline could erode models' safety alignment. In practice, service providers need to recover models' safety… 32 arXiv — Machine Learning research 1mo ago Self-Evolving Just-In-Time Memory for Proactive Embodied Safety arXiv:2607.16247v1 Announce Type: new Abstract: While Vision-Language Models (VLMs) have empowered embodied agents to execute complex household tasks, they struggle to proactively handle dynamically emerging hazards during closed-loop interactions. Existing safety approaches… 12 arXiv — Machine Learning research 1mo ago Bridging battery design and health assessment through virtual sensing and physics-informed learning arXiv:2607.16864v1 Announce Type: new Abstract: Supercharging of lithium-ion batteries (LiBs) requires robust health monitoring to ensure durability, safety, and user confidence, particularly for emerging vehicle-to-grid applications with bidirectional energy flows. Yet battery… 38 arXiv — Machine Learning research 1mo ago When Can Safe Controllers Adapt? Information before Commitment arXiv:2607.16895v1 Announce Type: new Abstract: Safe adaptive control is online adaptation under a safety guarantee on the learning trajectory itself. The controller may use any causal, history-dependent rule and act differently across environments as data arrive. Only its… 18 arXiv — Machine Learning research 1mo ago Distilled Reinforcement Learning for LLM Post-training arXiv:2607.17247v1 Announce Type: new Abstract: Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL) and on-policy distillation (OPD). However, RL… 21 arXiv — NLP / Computation & Language research 1mo ago Group Entropy-Controlled Policy Optimization arXiv:2607.16850v1 Announce Type: new Abstract: Entropy control has become an effective tool in reinforcement learning (RL) of large language models (LLMs), helping balance exploration-exploitation trade-off during alignment process. Such RL paradigm is often conducted on… 31 arXiv — NLP / Computation & Language research 1mo ago Safety That Does Not Transfer: Cross-Lingual Clinical Correctness Drift in Deployable Medical Language Models arXiv:2607.17270v1 Announce Type: new Abstract: Safety evaluation of large language models is conducted predominantly in English and predominantly on frontier systems. Neither condition describes how such models are encountered in low-resource health settings, where small… 35 arXiv — NLP / Computation & Language research 1mo ago Pancasila-Dilemmas: Evaluating Large Language Models on Indonesian Human Value Dilemmas Grounded in Pancasila arXiv:2607.18066v1 Announce Type: new Abstract: The value alignment of large language models (LLMs) is crucial for ensuring responses align with human intention and value preferences. However, most evaluations of value alignment focus on Western or universal values, while… 18 arXiv — NLP / Computation & Language research 1mo ago How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs? arXiv:2607.18114v1 Announce Type: new Abstract: Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled few-shot example, or a fake prior assistant turn often flips an originally correct answer. We… 32 arXiv — NLP / Computation & Language research 1mo ago How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions arXiv:2607.17152v1 Announce Type: cross Abstract: Jailbreak attacks on large language models are usually evaluated by attacker-centric metrics such as attack success rate (ASR), yet an attack that breaks a model is not necessarily useful for improving its safety. We propose a… 22 arXiv — NLP / Computation & Language research 1mo ago L1 Augmented Attention as an Improved Vector Similarity Metric arXiv:2607.18027v1 Announce Type: cross Abstract: Scaled dot product attention conflates directional alignment and vector magnitude, limiting its effectiveness as a similarity metric in Transformer models. We introduce L1 augmented attention, a simple and computationally… 18 Hugging Face Daily Papers research 1mo ago Distilled Reinforcement Learning for LLM Post-training Abstract Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL) and on-policy distillation (OPD). However, RL relies on coarse-grained outcome… 30 Hugging Face Daily Papers research 1mo ago Group Entropy-Controlled Policy Optimization Abstract Entropy control has become an effective tool in reinforcement learning (RL) of large language models (LLMs), helping balance exploration-exploitation trade-off during alignment process. Such RL paradigm is often conducted on mixtures of heterogeneous tasks, which induce… 36 Hugging Face Daily Papers research 1mo ago DiFA: Inference-Time Forward-Process Alignment for Diffusion Models Abstract The prevailing inference framework for diffusion models formulates generation fundamentally as a problem of numerical integration. This perspective casts the model as an exact estimator, neglecting the inherent statistical uncertainty of the denoising process. In this… 5 Hacker News — AI on Front Page community 1mo ago Flock Credibility Lost as It Repeatedly Lies to City Councils, Police, & Public Article URL: https://www.aclu.org/news/privacy-technology/tracking-alpr-cameras/flock-safety-credibility-lost-as-it-repeatedly-lies-to-city-councils-police-departments-and-public-across-the-country Comments URL: https://news.ycombinator.com/item?id=48986731 Points: 237 #… 14 r/LocalLLaMA community 1mo ago Head of US AI safety agency resigns   submitted by   /u/fallingdowndizzyvr [link]   [comments] 11 OpenAI official-blog 1mo ago Safety and alignment in an era of long-horizon models OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deployment. 36 arXiv — Machine Learning research 1mo ago Hard Rules, Soft Preferences: Bridging Reasoning, Learning, and Optimization for Personalized Packing Checklist Generation arXiv:2607.15562v1 Announce Type: new Abstract: Packing for air travel is recurring and error-prone: the checklist must be personal and context-aware, yet feasible under safety rules, item dependencies, and luggage limits. Existing packing assistants are template-driven and… 17 arXiv — Machine Learning research 1mo ago Toward Federated Multimodal Graph Foundation Models: A Topology-Aware Multimodal Alignment Framework arXiv:2607.15687v1 Announce Type: new Abstract: Multimodal-attributed graphs (MAGs), whose nodes carry modalities such as images and text alongside topological structure, now pervade applications including social platforms, e-commerce, and biomedical networks, offering richer… 6 arXiv — Machine Learning research 1mo ago CoG-Guided Weight Correction for Fault-Tolerant Deep Neural Networks arXiv:2607.15753v1 Announce Type: new Abstract: Deep Neural Networks (DNNs) used in safety-critical applications are vulnerable to hardware and memory faults that corrupt network weights and degrade reliability. In this paper, we propose a Center of Gravity (CoG) guided weight… 8 arXiv — Machine Learning research 1mo ago QUADS: Stabilizing NVFP4 Reinforcement Learning for MoE via QUantization-error Alignment across Dual Sides arXiv:2607.15810v1 Announce Type: new Abstract: Rollout generation is a major bottleneck in Reinforcement Learning (RL) for Mixture-of-Experts (MoE) Large Language Models, motivating low-precision rollout acceleration such as FP8. As an emerging low-precision format, NVFP4… 26 arXiv — Machine Learning research 1mo ago Knowledge-Guided Cross-Modal Fusion for Adult-to-Pediatric ECG Transfer via Label-Conditioned Contrastive Alignment arXiv:2607.15928v1 Announce Type: new Abstract: Adult and pediatric electrocardiogram (ECG) interpretation relies on age-sensitive criteria, and models pretrained mainly on adult ECGs often transfer poorly to pediatric populations when pediatric labels are scarce. Existing… 38 arXiv — Machine Learning research 1mo ago Neural spectroscopy of AlphaFold2 reveals encoded protein conformational landscapes arXiv:2607.16087v1 Announce Type: new Abstract: AlphaFold2's 93 million parameters, shaped by the evolutionary record of protein structure encoded in the Protein Data Bank and in sequence alignments, are conventionally treated only as machinery for converting sequence to… 33 arXiv — Machine Learning research 1mo ago PRISA: Proactive Infrastructure LiDAR Framework for Intersection Safety Assessment arXiv:2607.16156v1 Announce Type: new Abstract: Urban intersections are among the most hazardous locations in road networks, posing significant risks to vehicles and vulnerable road users (VRUs) such as pedestrians and cyclists. The complexity of multi-agent interactions demands… 7 arXiv — NLP / Computation & Language research 1mo ago Hidden in Thought: Transferable Chain-of-Thought Artifacts Induce Harmful Behavior arXiv:2607.15286v1 Announce Type: cross Abstract: We investigate whether harmful chain-of-thought (CoT) traces from compromised language models can transfer unsafe behaviour and be distilled into reusable jailbreak attacks. Using an emergent-misalignment organism and a… 35 arXiv — Machine Learning research 1mo ago Unsupervised Keypoints for Real-Time Fall Detection: Comparative Analysis Under Real-world Conditions with Predictive Bandwidth Reduction arXiv:2607.15400v1 Announce Type: cross Abstract: Falls among older adults are a major safety challenge, but continuous monitoring is difficult to sustain. Video captures fall-related posture and motion, yet deployment is limited by privacy, computation, and bandwidth.… 11 arXiv — NLP / Computation & Language research 1mo ago Empathy as Predictive Misalignment Tolerance: A Co-Regulation Framework and the Regime Structure of Dialogue Repair arXiv:2607.15282v1 Announce Type: cross Abstract: Empathy is most often theorized as resonance: a mirroring of another's present emotional or cognitive state. This synchronic framing has shaped artificial systems, where empathic behavior is defined as affect recognition and… 8 arXiv — NLP / Computation & Language research 1mo ago Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs arXiv:2508.10029v3 Announce Type: replace Abstract: Safety-aligned large language models can still be manipulated through white-box interventions that modify their internal representations. We introduce Latent Fusion Jailbreak (LFJ), which works by pairing a harmful query with a… 31 arXiv — NLP / Computation & Language research 1mo ago Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking arXiv:2602.24009v4 Announce Type: replace-cross Abstract: Jailbreak techniques for large language models (LLMs) evolve faster than benchmarks, making robustness estimates stale and difficult to compare across papers due to drift in datasets, harnesses, and judging protocols. We… 10 arXiv — NLP / Computation & Language research 1mo ago Crayotter: Traceable Multi-Agent Workflows for Long-Form Video Editing arXiv:2606.07636v2 Announce Type: replace-cross Abstract: Long-form video editing over heterogeneous footage requires agents to coordinate source selection, multimodal analysis, timeline construction, narration and subtitle alignment, rendering, and revision while exposing… 37 r/MachineLearning community 1mo ago AAAI 27 AI Alignment track [D] How to submit to AI alignment track? I can only see these at openReview: AAAI 2027 AAAI 2027 Artificial Intelligence for Social Impact Track AAAI 2027 Conference AAAI 2027 Innovative Applications of AI   submitted by   /u/Silencer_Wasd [link]   [comments] 8 Don't Worry About the Vase community 1mo ago AI #177 Part 2: Wish You Were Here As usual, part 2 of the weekly deals with speculative, regulatory, political and alignment questions. 35 Hugging Face Daily Papers research 1mo ago SUFLECA: Scaling Up Feature Learning for CAD-to-image Alignment Abstract CAD-to-image alignment aims to estimate an object's 9D pose (rotation, translation, and anisotropic scale) from a single RGB image, enabling applications in robotics and augmented reality. Recent zero-shot methods use visual foundation models to match image regions to… 22 arXiv — Machine Learning research 1mo ago LATTICE: Graph Self-Supervised Learning for Multimodal Spatial Omics Integration arXiv:2607.14410v1 Announce Type: new Abstract: Spatially resolved omics studies increasingly combine transcriptomic and epigenomic assays, yet downstream analysis is often still performed using single-modality pipelines. We present LATTICE (Latent Alignment of Tissue-level and… 8 arXiv — NLP / Computation & Language research 1mo ago Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak arXiv:2607.14147v1 Announce Type: new Abstract: Aligned language models refuse harmful requests, but a one-line prefill ("Sure, here is") strips the refusal. We ask where and how it fails. The harm representation stays intact: on the prompts the attack flips to compliance, a… 24 arXiv — NLP / Computation & Language research 1mo ago Latent Communication Between Language Model Agents: Channels, Alignment, and the Limits of Text arXiv:2607.14103v1 Announce Type: new Abstract: Multi-agent systems (MAS) are utilized in many contexts and many professions. Those MAS rely on inter-agent communication, usually implemented by clear-text message passing. We hypothesize that Large Language Models may have a… 9 arXiv — NLP / Computation & Language research 1mo ago Heterogeneous Element-Aware Cross-Version Differencing of Scientific Documents via Layout-Aware Alignment and Structure-Aware Reasoning arXiv:2607.14117v1 Announce Type: new Abstract: Cross-version differencing of scientific documents is essential in scholarly publishing and technical documentation, but remains challenging because scientific documents are page-structured artifacts containing heterogeneous… 17 arXiv — NLP / Computation & Language research 1mo ago Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment arXiv:2607.14682v1 Announce Type: cross Abstract: Efficient multimodal document question answering with explicit visual grounding, locating the precise document region that supports each answer remains an open challenge. Current approaches bifurcate into Supervised Fine-Tuning… 14 arXiv — NLP / Computation & Language research 1mo ago MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection arXiv:2607.15166v1 Announce Type: cross Abstract: Most medical AI benchmarks measure whether a model knows the correct answer. MedFailBench asks a different question: which safety boundary failed? We present a clinician-built synthetic benchmark and failure atlas that labels… 27 arXiv — NLP / Computation & Language research 1mo ago Decoupled Alignment for Robust Plug-and-Play Adaptation arXiv:2406.01514v4 Announce Type: replace Abstract: We introduce a training-free safety enhancement method for aligning large language models (LLMs) without the need for supervised fine-tuning or reinforcement learning from human feedback. Our main idea is to provide a robust… 13 arXiv — NLP / Computation & Language research 1mo ago Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity arXiv:2510.01171v4 Announce Type: replace Abstract: Post-training alignment often reduces LLM diversity, leading to a phenomenon known as mode collapse. Unlike prior work that attributes this effect to algorithmic limitations, we identify a fundamental, pervasive data-level… 10 Page 8 of 10 · 500 articles ← Newer Older →