News / #safety Tag Safety + alignment 500 articles archived under #safety · RSS Sign in to follow arXiv — NLP / Computation & Language research 1mo ago Fidelity Is Not Safety: Gently-Compressed LLMs Pass Every Data-Free Quality Guard Yet Invent Procedure Steps in Agentic Execution arXiv:2607.28196v1 Announce Type: new Abstract: Practitioners accept a compressed language model once it clears a stack of data-cheap quality guards: perplexity within a small factor of the original, downstream accuracy (for example MMLU) inside a confidence interval, and… 26 arXiv — NLP / Computation & Language research 1mo ago Inducing language models to assert their own consciousness restores human beliefs and values arXiv:2607.28607v1 Announce Type: new Abstract: Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety… 18 arXiv — NLP / Computation & Language research 1mo ago Measuring Alignment With Reader Highlights Net of Position and Length arXiv:2607.27739v1 Announce Type: cross Abstract: Context compression discards most of a document before a language model reads it, and is normally evaluated by downstream task accuracy - which makes another model the judge of what mattered. Naturalistic social highlighting… 38 Ars Technica — AI news-outlet 1mo ago Google reveals Gemini Robotics 2.0, promising improved dexterity and safety Gemini Robotics 2 includes three models, but only one is publicly available right now. 22 Hugging Face Daily Papers research 1mo ago GPT-Red: Automated Red Teaming via Self-Play at Scale Abstract We introduce GPT-Red, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, we use it to adversarially… 27 arXiv — Machine Learning research 1mo ago Data Fusion and Contrastive Alignment for Unconstrained IR Molecular Structure Elucidation arXiv:2607.26164v1 Announce Type: new Abstract: Automated molecular structure elucidation from infrared (IR) spectroscopy data has seen significant advancements in recent years, but its broad applicability is limited by a reliance on pre-determined chemical formulas provided as… 19 arXiv — Machine Learning research 1mo ago Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models arXiv:2607.26173v1 Announce Type: new Abstract: Alignment training, model organisms, and toy models are usually treated as separate research areas. But projects in all three frequently use supervised fine-tuning (SFT) to pursue the same underlying goals. When projects share a… 5 arXiv — Machine Learning research 1mo ago Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions arXiv:2607.26820v1 Announce Type: new Abstract: As large language models (LLMs) evolve from standalone assistants into autonomous agents, ensuring their safety requires shifting beyond pointwise risk assessment to understand how risks emerge and unfold over long-horizon… 4 arXiv — Machine Learning research 1mo ago Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark arXiv:2607.27143v1 Announce Type: new Abstract: High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantification under severe class imbalance and asymmetric error costs. Standard marginal conformal… 15 arXiv — Machine Learning research 1mo ago Shape-Based Inductive Bias for Glioma Grading from Tumor Contours arXiv:2607.26090v1 Announce Type: cross Abstract: Glioma grading from tumor contours is often treated as a pixel problem even when the signal of interest is shape. We align closed contours with a functional shape-alignment framework, separate global deformation from residual… 11 arXiv — NLP / Computation & Language research 1mo ago GPT-Red: Automated Red Teaming via Self-Play at Scale arXiv:2607.26115v1 Announce Type: cross Abstract: We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production… 11 arXiv — Machine Learning research 1mo ago A Picture Says Thousands of Words - Harnessing Dermal Exposure Data from Images through Hybrid Deep Learning for Enhanced Safety Assessment arXiv:2607.26170v1 Announce Type: cross Abstract: This study developed a hybrid computer vision method to quantify exposed skin from images for dermal exposure assessment. Using 170 indoor-painting images, Mask R-CNN first identified human subjects and removed background… 21 arXiv — NLP / Computation & Language research 1mo ago Steering Instruction Hierarchies at Inference Time arXiv:2607.26228v1 Announce Type: new Abstract: Instruction hierarchies are a core safety assumption of language model deployment: higher priority inputs, such as system prompts, should override conflicting lower priority inputs from users or tools. Yet frontier LLMs often… 33 arXiv — NLP / Computation & Language research 1mo ago Misalignment Has a Personality: A Big Five Account of Emergent Misalignment arXiv:2607.26389v1 Announce Type: new Abstract: Fine-tuning a language model on data containing a narrow flaw, such as insecure code or incorrect mathematical answers, can cause broad misalignment through a mechanism that remains debated. We provide an interpretable account: in… 13 arXiv — NLP / Computation & Language research 1mo ago Constitutional Midtraining: Content Presence Drives Alignment Gains arXiv:2607.26654v1 Announce Type: new Abstract: Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce durable alignment remains untested. We test this via constitutional… 31 arXiv — NLP / Computation & Language research 1mo ago DIRECT: Direct Decoding for Efficient and Aligned Sequence Labeling with Large Language Models arXiv:2607.26891v1 Announce Type: new Abstract: Sequence labeling is a fine-grained information extraction task, yet existing large language model-based approaches suffer from insufficient domain alignment and low inference efficiency. To address these issues, we propose DIRECT,… 20 arXiv — NLP / Computation & Language research 1mo ago OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment arXiv:2607.26981v1 Announce Type: new Abstract: Large language models are increasingly used as decision aids whose probability judgments shape downstream choices. Whether those judgments carry a systematic directional tilt has been hard to detect: calibration metrics aggregate… 35 arXiv — NLP / Computation & Language research 1mo ago Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis arXiv:2607.26541v1 Announce Type: cross Abstract: Audio-capable foundation models enable end-to-end spoken interaction, but they also introduce safety risks beyond transcript content. It remains unclear how much jailbreak capability can arise from matched-text variation in… 29 arXiv — NLP / Computation & Language research 1mo ago On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment arXiv:2607.27081v1 Announce Type: cross Abstract: Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers can embed harmful behaviors into downstream corpora, creating models that… 5 TechCrunch — AI news-outlet 1mo ago Thinking Machines co-founder Lilian Weng left the company citing health reasons, then joined OpenAI Weng previously served as the VP of AI Safety Research at OpenAI. 27 Hugging Face Daily Papers research 1mo ago Projection Pursuit CPCANet for Domain Generalization Abstract Domain Generalization (DG) aims to learn representations robust to distribution shifts. Recent geometric alignment methods, such as CPCANet, extract domain-invariant structures through batch-wise Common Principal Component Analysis (CPCA). However, CPCANet suffers from… 7 arXiv — NLP / Computation & Language research 1mo ago MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios arXiv:2607.25186v1 Announce Type: new Abstract: Background: Most medical large language model (LLM) benchmarks focus on examination knowledge or isolated tasks and may not reflect the longitudinal, multimodal, and safety-critical workflow of cardiovascular care. Objective: To… 33 arXiv — NLP / Computation & Language research 1mo ago IRIS: Reusable Identity Representations from Frozen LLMs for Entity Alignment arXiv:2607.25579v1 Announce Type: new Abstract: Entity alignment (EA) identifies entities across knowledge graphs (KGs) that refer to the same real-world object. Conventional EA methods mainly exploit explicit graph structures and textual fields, which often provide insufficient… 7 arXiv — NLP / Computation & Language research 1mo ago Evaluation of forced alignment of code-mixed speech: the case of Hindi-English arXiv:2607.25581v1 Announce Type: new Abstract: Code-mixed speech poses unique challenges to forced alignment: expanded inventories, orthographic errors, and speaker variation. We evaluate forced alignment of Hindi-English code-mixed speech using the Montreal Forced Aligner. We… 35 arXiv — NLP / Computation & Language research 1mo ago Shieldstral arXiv:2607.25857v1 Announce Type: new Abstract: We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7$\times$ its size on text safety benchmarks and sets a new state of the art on multimodal safety… 19 arXiv — NLP / Computation & Language research 1mo ago LLM Scheming Inversely Scales with Pretraining Language Coverage arXiv:2607.24769v1 Announce Type: cross Abstract: With the growing capabilities of frontier models, AI alignment becomes increasingly critical in high-risk deployment settings. While recent work has empirically demonstrated in-context scheming -- the covert pursuit of misaligned… 36 arXiv — NLP / Computation & Language research 1mo ago Retrieval-Augmented Generation in LLMs for Mental Health: Quantifying the Incremental Contribution of Retrieval Within a Layered Safety Architecture arXiv:2607.24817v1 Announce Type: cross Abstract: Digital mental health interventions (DMHIs) offer scalable support, but ensuring they accurately detect users' intent during volatile situations can be challenging. Pure parametric Large Language models (LLMs) do not contain… 10 arXiv — NLP / Computation & Language research 1mo ago Towards Robust Reinforcement Learning for Small-Scale Language Model Agents arXiv:2607.25091v1 Announce Type: cross Abstract: The alignment of Small Language Models (SLMs) in the 70--500M parameter range using reinforcement learning is often considered unstable, though the underlying failure mechanisms have not been systematically investigated. In the… 37 arXiv — NLP / Computation & Language research 1mo ago MemSFT: Mitigating Alignment Tax with an External Parametric Memory arXiv:2607.25614v1 Announce Type: cross Abstract: Adapting Large Language Models (LLMs) to specialized domains often incurs an alignment tax, as fine-tuning on domain-specific tasks can cause catastrophic forgetting and substantially degrade performance on general tasks. We… 37 arXiv — NLP / Computation & Language research 1mo ago Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM arXiv:2508.05775v3 Announce Type: replace Abstract: Large Language Models (LLMs) have revolutionized content creation across digital platforms, offering unprecedented capabilities in natural language generation and understanding. Meanwhile, they pose risks by inadvertently… 14 Hugging Face Daily Papers research 1mo ago Shieldstral Abstract We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7times its size on text safety benchmarks and sets a new state of the art on multimodal safety classification. Shieldstral formulates content… 14 Hugging Face Daily Papers research 1mo ago Towards Robust Reinforcement Learning for Small-Scale Language Model Agents Abstract The alignment of Small Language Models (SLMs) in the 70--500M parameter range using reinforcement learning is often considered unstable, though the underlying failure mechanisms have not been systematically investigated. In the State-of-the-Art (SOTA) research, fifteen… 34 arXiv — NLP / Computation & Language research 1mo ago Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B arXiv:2607.22545v1 Announce Type: cross Abstract: Deploying large language models in financial-services and agentic settings requires safety classifiers that simultaneously handle prompt injection, regulatory compliance, and general harm, a combination no existing open guardrail… 21 arXiv — Machine Learning research 1mo ago Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation arXiv:2607.22766v1 Announce Type: new Abstract: The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality. As datasets scale, massive preference and instruction-tuning corpora inevitably accumulate hidden structural contradictions, safety risks,… 25 arXiv — Machine Learning research 1mo ago Physically Verifiable Evidence and LLM-Based Reporting for Bearing Fault Diagnosis arXiv:2607.22797v1 Announce Type: new Abstract: Trustworthy deployment of AI-based diagnosis in safety-critical mechanical systems hinges on validation: whether a prediction can be checked against physical reality before it is acted upon. Current intelligent fault diagnosers… 38 arXiv — Machine Learning research 1mo ago Distribution-Specific Curvature Control with Finite-Sample Guarantees for Open-Weight Safety arXiv:2607.22929v1 Announce Type: new Abstract: A short fine-tuning run can undo the safety guards of an open-weight model---retraining a refusal-trained assistant to aid weapons development or produce hate speech. Preventing such harmful fine-tuning while retaining benign… 30 arXiv — Machine Learning research 1mo ago Diffusion-Guided Search via Exponential Tilting (DiffTilt): An Application to Falsification of Safety-Critical Systems arXiv:2607.23134v1 Announce Type: new Abstract: Discovering rare safety-critical failures in autonomous and cyber-physical systems is a fundamental challenge in verification and validation. Existing falsification approaches rely on conditional sampling strategies that factor the… 19 arXiv — Machine Learning research 1mo ago Context-Aware Concept Distillation for Trustworthy Flood Prediction arXiv:2607.23237v1 Announce Type: new Abstract: Effective flood risk management relies on accurate forecasting, yet the "black box" nature of stateof-the-art Deep Learning models creates a barrier to trust and accountability in high-stakes public safety decisions. While existing… 21 arXiv — Machine Learning research 1mo ago Directional Influence Function: Estimating Training Data Influence in Constrained Learning arXiv:2607.23388v1 Announce Type: new Abstract: As constrained learning becomes increasingly common, models are trained under explicit feasibility requirements to enforce fairness, safety, robustness, regulariza- tion, and physics or logic constraints. Understanding how training… 25 arXiv — NLP / Computation & Language research 1mo ago Not All LLM Reasoning is Visible in the Chain-of-Thought arXiv:2607.22925v1 Announce Type: new Abstract: A key question for AI safety is whether a language model expresses all of its reasoning in its output tokens. We demonstrate a concrete failure mode where frontier models exhibit invisible reasoning by leveraging semantically… 16 arXiv — NLP / Computation & Language research 1mo ago Beyond a Global Norm: Personalizing Toxicity Sensitivity in Language Models Without Retraining arXiv:2607.23175v1 Announce Type: new Abstract: Reducing toxicity is often framed as a global alignment problem, yet perceptions of harmful language are subjective and context-dependent. We present the first comparative evaluation of training-free methods for aligning language… 35 arXiv — NLP / Computation & Language research 1mo ago SyRuP: Enhancing System-Prompt Following via Reward-Guided Prediction in LLM Decoding arXiv:2607.23991v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly controlled through system prompts that specify roles, styles, formats, and safety requirements. However, models follow these prompts only implicitly through in-context learning, which… 24 arXiv — NLP / Computation & Language research 1mo ago From transcription to semantic corpus analysis: unsupervised learning of sentence representations for ancient languages arXiv:2607.24542v1 Announce Type: new Abstract: Automatic Text Recognition (ATR) now supplies digital humanities with large volumes of unstructured, heterogeneous, and often noisy text in ancient languages. Downstream semantic analysestext reuse identification, alignment, and… 5 arXiv — NLP / Computation & Language research 1mo ago STAIF: A Stage-wise Optimization for Complex Instruction Following arXiv:2607.22649v1 Announce Type: cross Abstract: Following complex instructions with multiple explicit constraints remains a fundamental challenge for large language models (LLMs). Existing alignment methods, such as DPO, optimize holistic reward signals that often… 17 arXiv — NLP / Computation & Language research 1mo ago How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift arXiv:2607.22676v1 Announce Type: cross Abstract: Post-training is a key mechanism for adapting large language models to downstream tasks. While prior work suggests that task adaptation can alter a model's pre-existing alignment, especially its safety behavior, its broader… 11 Hugging Face Daily Papers research 1mo ago OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Abstract Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging… 12 TechCrunch — AI news-outlet 1mo ago OpenAI’s Hugging Face breach has reignited the debate over alignment and control OpenAI's Hugging Face breach has reignited debate over AI alignment and control, exposing competing views on whether increasingly capable AI should be better aligned, better contained, or both. 14 Hugging Face Daily Papers research 1mo ago LAMAR: An Open Language-Aware Multilingual Alignment Reranker Abstract In multilingual retrieval augmented generation, a retriever can retrieve relevant documents written in multiple languages, which are subsequently reranked before answer generation. However, it remains unclear whether existing multilingual rerankers consider document… 25 arXiv — Machine Learning research 1mo ago Adjustment Speed as a Safety Constraint for Nonstationary Reinforcement Learning arXiv:2607.21646v1 Announce Type: new Abstract: Ensuring safety in reinforcement learning under nonstationarity requires determining whether a learning system can safely adapt to forecasted environmental change within the required recovery horizon. Existing safe reinforcement… 24 arXiv — Machine Learning research 1mo ago Evolution-Aware MSA Reasoning for Subsampling via Factor Graphs arXiv:2607.22314v1 Announce Type: new Abstract: Multiple Sequence Alignments (MSAs) provide protein language models with explicit evolutionary context, but their large depth makes subsampling unavoidable under limited token budgets. Existing strategies, including random… 8 Page 6 of 10 · 500 articles ← Newer Older →