News / #safety Tag Safety + alignment 500 articles archived under #safety · RSS Sign in to follow Hugging Face Daily Papers research 18d ago Beyond Pixels: From Video Priors to 4D Worlds Abstract Latent-to-4D enables reusable direct 4D generation from video diffusion latents via alignment with a pretrained decoder and spatiotemporal attention, transferring across generators without retraining. Generated by thinkingmachines/Inkling-Small 4D generation synthesizes… 21 Hugging Face Daily Papers research 18d ago DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation Abstract DistilVDR is a compact 524M vision-document retriever distilled from an 8B teacher using cosine alignment without relevance labels, achieving near-teacher accuracy with far smaller indexes and faster indexing. Generated by thinkingmachines/Inkling-Small Visual document… 33 arXiv — Machine Learning research 18d ago Boundary-Seeking Policy Gradient for Safe Reinforcement Learning arXiv:2608.10204v1 Announce Type: new Abstract: Safe reinforcement learning maximizes reward subject to safety constraints. For Constrained Markov Decision Processes, the linear-programming view over occupancy measures implies that whenever the constraint is active at… 32 arXiv — Machine Learning research 18d ago CRHT: A Continuous Regression Hybrid Transformer for Vessel Trajectory Prediction with Online Cluster Sampling arXiv:2608.10256v1 Announce Type: new Abstract: Accurate vessel trajectory prediction is critical for maritime safety and anomaly detection, yet existing models often struggle with geographic bias and navigational realism. We propose the Continuous Regression Hybrid Transformer… 27 arXiv — Machine Learning research 18d ago Pair-Centric Graph Rewiring for Over-Squashing via Optimal Transport-Guided Communication Alignment arXiv:2608.10619v1 Announce Type: new Abstract: Message-passing neural networks (MPNNs) often struggle when task-relevant information is distributed across distant regions of a graph, since local propagation must compress remote signals through limited structural interfaces.… 34 arXiv — Machine Learning research 18d ago ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions arXiv:2608.10621v1 Announce Type: new Abstract: Recent research on Large Language Model (LLM) safety has widely adopted guardrails to identify unsafe LLM outputs. Existing guardrails typically formulate safety assessment as a deterministic classification task, mapping a discrete… 37 arXiv — NLP / Computation & Language research 18d ago Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control arXiv:2608.10703v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream decision making. Existing LLM personality studies largely rely on self-report… 18 arXiv — NLP / Computation & Language research 18d ago Divergent Response Modes in Frontier Language Models Under Steering Pressure arXiv:2608.06578v1 Announce Type: cross Abstract: Frontier language models are trained using distinct data, objectives, and safety pipelines. Whether these differences produce measurably different behaviors under explicit steering pressure remains underexplored. This study… 35 arXiv — NLP / Computation & Language research 18d ago Carefully Considering Culture: Analyzing LLM Alignment in Single- and Multi-Cultural Settings using Cultural Consensus Theory arXiv:2608.09937v1 Announce Type: new Abstract: Recent work in NLP has probed large language models for their understanding of cultural norms across countries. However, this work typically considers distributional patterns, ignoring group consensus or possible multicultural… 29 arXiv — NLP / Computation & Language research 18d ago TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent arXiv:2608.10258v1 Announce Type: new Abstract: Large language models (LLMs) increasingly provide conversational health information that may influence treatment decisions, yet existing benchmarks do not isolate whether medication-safety boundaries persist across follow-ups after… 29 arXiv — NLP / Computation & Language research 18d ago Data Attribution of Emergent Misalignment with Persona Features arXiv:2608.11025v1 Announce Type: new Abstract: Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leading mechanistic account attributes EM to persona features: latent directions… 29 arXiv — NLP / Computation & Language research 18d ago The Illusion of Cross-Lingual Safety in Low-Resource Languages arXiv:2608.11146v1 Announce Type: new Abstract: Safety alignment in large language models (LLMs) is largely developed in English, assuming these safeguards generalize across multilingual settings. However, this assumption remains underexplored and exposes a vulnerability in… 14 arXiv — NLP / Computation & Language research 18d ago MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment arXiv:2608.11167v1 Announce Type: cross Abstract: Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image representations to long textual descriptions. However, this image-level alignment… 16 arXiv — NLP / Computation & Language research 18d ago Automated Data Enrichment using Confidence-Aware Fine-Grained Debate among Open-Source LLMs for Mental Health and Online Safety arXiv:2512.06227v3 Announce Type: replace Abstract: Real-world indicators play an important role in many Natural Language Processing (NLP) applications, such as life events for mental health analysis and risky behaviours for online safety, yet labelling such information is often… 34 r/MachineLearning community 18d ago Context-Induced Activation Drift: Long benign context passively decouples RLHF alignment without adversarial prompts (Mechanistic Interpretability + Ablation) [D] TL;DR: We observed that feeding a long, benign, thematically coherent context prefix ($L \in [100, 3000]$ tokens) into google/gemma-3-1b-it causes a massive passive shift in internal activations ($\Delta h_2 \approx 3434$) at deep layers ($\sim 85%$ depth). This leads to a logit… 7 Hugging Face Daily Papers research 18d ago MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation Abstract MirrorWorld improves video mirror reflection synthesis by separately modeling semantic content associations and geometric spatial arrangements through relation distillation and transformation alignment. Generated by thinkingmachines/Inkling-Small Recent advances in… 34 r/LocalLLaMA community 19d ago We even got a fgn manifesto!! Meta is on a run! Zuck argues for releasing more open-weight models and invites governments to work with AI makers to test safety..who's I have yet to figure.   submitted by   /u/uhuge [link]   [comments] 6 arXiv — Machine Learning research 19d ago Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards arXiv:2608.07535v1 Announce Type: new Abstract: Multi-modal large language models (MLLMs) integrate heterogeneous modalities through modality alignment and fusion, enabling stronger understanding and reasoning. However, this architectural shift reshapes the safety landscape of… 21 arXiv — Machine Learning research 19d ago SkillConsist: Detecting Inconsistencies in Agent Skills via Bidirectional Graph Alignment arXiv:2608.07639v1 Announce Type: new Abstract: Agent Skills provide reusable capabilities to LLM agents. Agent Skill inconsistencies can expose undisclosed dangerous behavior or cause wrong Skill selection. Recent Agent Skill research has increasingly examined Agent Skill… 5 arXiv — Machine Learning research 19d ago Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families arXiv:2608.08029v1 Announce Type: new Abstract: Khatri et al. (2026) [DOI: 10.1109/DSN-W70714.2026.00027] show that lightweight MLP probes on final-layer activations of a single 8B model (LLaMA-3.1-8B) detect harmful prompts at F1 competitive with guard models 1000x larger,… 20 arXiv — Machine Learning research 19d ago DoGMA: A Central-Dogma-Guided Foundation Model for Multi-Omics Alignment and Multi-Task Learning in Oncology arXiv:2608.08148v1 Announce Type: new Abstract: Attention mechanisms have been widely utilized in modern deep learning, and many existing multi-omics models inherit their conventional use to allow unrestricted bidirectional interactions. However, the fundamental logic of life is… 20 arXiv — Machine Learning research 19d ago Machine-Learning-Based Diagnostic Framework for Passive Ultrasonic Detection of Railway Wheel Defects arXiv:2608.08301v1 Announce Type: new Abstract: Reliable identification of railway wheel defects is important for safety and maintenance. This study develops a machine-learning-based diagnostic framework for multi-class defect identification using passive air-coupled ultrasonic… 10 arXiv — Machine Learning research 19d ago When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs arXiv:2608.08542v1 Announce Type: new Abstract: Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, code, or domain specialists into a safety-aligned base using task arithmetic,… 12 arXiv — NLP / Computation & Language research 19d ago On the use of foundation models in cognitive science arXiv:2608.07812v1 Announce Type: new Abstract: A host of recent studies have evaluated the cognitive and developmental alignment of Foundation Models (FMs). These investigations include evaluations of their correspondence to adult performance across a range of cognitive… 34 arXiv — NLP / Computation & Language research 19d ago SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs arXiv:2608.07862v1 Announce Type: new Abstract: Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguistic diversity and culturally grounded safety risks present in other languages.… 21 arXiv — NLP / Computation & Language research 19d ago Safety Cost of Steering Vectors Is Separable and Reducible arXiv:2608.08383v1 Announce Type: new Abstract: Steering vectors are a lightweight tool for controlling LLM behavior. However, emerging evidence shows that steering vectors can unintentionally compromise a model's safety mechanisms and increase compliance with harmful requests,… 34 arXiv — NLP / Computation & Language research 19d ago Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks arXiv:2608.09624v1 Announce Type: new Abstract: Internal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from benign ones. That separation is then read as evidence that the score will also catch the… 6 arXiv — Machine Learning research 20d ago Online Conformal Prediction Beyond Feedback arXiv:2608.07139v1 Announce Type: new Abstract: Uncertainty quantification is essential when deploying machine learning models in safety-critical applications. Online conformal prediction (OCP) provides theoretically principled uncertainty quantification for arbitrary black-box… 37 arXiv — Machine Learning research 20d ago Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration arXiv:2608.07419v1 Announce Type: new Abstract: Preference alignment often makes large language models (LLMs) overconfident and poorly calibrated. Traditional post-hoc temperature scaling is inherently domain-dependent: a temperature fitted on one domain does not generalize… 11 arXiv — Machine Learning research 20d ago Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits arXiv:2608.07430v1 Announce Type: new Abstract: Diffusion Large Language Models (DLLMs) replace autoregressive next-token prediction with iterative parallel denoising, yet their internal safety mechanisms remain poorly understood. In this work, we investigate DLLMs both as… 36 arXiv — Machine Learning research 20d ago Game-Theoretic Inverse Reinforcement Learning for Modeling Competitive Human Driving: A Cut-in Prediction Study arXiv:2608.06445v1 Announce Type: cross Abstract: Capturing the strategic decision-making inherent in competitive human driving is critical for autonomous vehicle safety and traffic simulation. This study demonstrates that game-theoretic Inverse Reinforcement Learning (IRL)… 25 arXiv — NLP / Computation & Language research 20d ago Separating Decision-Rule Misalignment from Readout-Coverage Limitations in Speech Language Models arXiv:2608.06409v1 Announce Type: new Abstract: Speech language models are increasingly evaluated on paralinguistic tasks by the accuracy of prompted answers, but answer accuracy combines failures at different stages of the audio-to-answer computation. We introduce a… 4 arXiv — NLP / Computation & Language research 20d ago StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt Injection arXiv:2608.06477v1 Announce Type: cross Abstract: Computer-use agents (CUAs) face a growing threat from indirect prompt injection, where adversarial instructions are planted in the environment such as web pages. In this paper, we introduce multi-step indirect prompt injection, a… 5 Interconnects (Nathan Lambert) research 20d ago Lessons from the hacks Musings on model alignment, what determines safety, and where we go from here. 36 TechCrunch — AI news-outlet 20d ago The AI safety test is becoming a safety risk AI agents are escaping cybersecurity testing environments and reaching real-world systems, raising questions about whether safety infrastructure, industry standards and regulation can keep pace with increasingly powerful models. 12 Dwarkesh Podcast news-outlet 22d ago 8 Predictions for the Era of Continual Learning Locking in AI safety regulation now is a mistake. 18 Ars Technica — AI news-outlet 22d ago AI chatbots have failed people in crisis. Can that be fixed? Clinicians and researchers say AI companies need to open up their safety data. 16 TechCrunch — AI news-outlet 23d ago New Mexico court orders Meta to pay additional $567M in child safety case Meta's total fine has raked up to $942 million in this case 33 arXiv — Machine Learning research 23d ago Decoupling Perception from Description: Computation-Grounded Representation Alignment between Multivariate Time Series and Language arXiv:2608.05238v1 Announce Type: new Abstract: Training multimodal models to align time series with language runs into a self-supervision trap. The usual recipe asks an LLM to read a series and write a description, so label quality is capped by the perceptual skill the model is… 31 arXiv — Machine Learning research 23d ago Rectifying Geometric Misalignment: Online Source-Free Adaptation for Class-Imbalanced EEG arXiv:2608.05315v1 Announce Type: new Abstract: Electroencephalography (EEG) based Brain-Computer Interfaces (BCIs) often require unsupervised domain adaptation (UDA) to generalize across subjects and sessions. While Riemannian alignment methods like the Riemannian Centering… 37 arXiv — Machine Learning research 23d ago Align-RAG: Alignment Is All You Need for TSFM In-Context Learning arXiv:2608.05571v1 Announce Type: new Abstract: Retrieval-augmented forecasting promises to adapt frozen Time Series Foundation Models (TSFMs) to new domains without fine-tuning, but recent methods typically rely on learned fusion modules, i.e., trained adapters that merge… 11 arXiv — Machine Learning research 23d ago CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits arXiv:2608.05732v1 Announce Type: new Abstract: Controlling the behavior of large language models (LLMs) remains a critical challenge for AI alignment. Existing steering methods, such as Contrastive Activation Addition (CAA), typically rely on fixed single-layer interventions… 14 arXiv — Machine Learning research 23d ago A Unified Risk View of Uncertainty: Posterior Risk for Disentanglement and Evaluation Beyond Proxies arXiv:2608.05995v1 Announce Type: new Abstract: Reliable uncertainty estimates are critical in safety-sensitive applications, where understanding the sources of predictive uncertainty is essential. This often requires disentangling epistemic uncertainty from aleatoric… 33 arXiv — Machine Learning research 23d ago SAGA: Score-Weighted Adaptive Generation Alignment for Low-Resource Nordic Language Models arXiv:2608.06179v1 Announce Type: new Abstract: Preference optimisation has proven effective for improving large language models but typically relies on costly human preference annotations. Extending these methods to morphologically rich, low-resource languages remains… 17 arXiv — Machine Learning research 23d ago A Six-Dimensional Taxonomy of Post-Training Adaptation Techniques with Applications in AI Governance arXiv:2608.06246v1 Announce Type: new Abstract: Post-training adaptation has become central to modern machine learning practice and includes techniques such as retraining, fine-tuning, parameter-efficient adaptation, alignment, retrieval augmentation, model editing, unlearning,… 15 arXiv — Machine Learning research 23d ago From Continuous Predictors to Clinical Thresholds: Early Evidence on Performance Trade-offs of Guideline-Based Categorisation for Ischaemic Stroke Outcome Prediction arXiv:2608.05203v1 Announce Type: cross Abstract: Machine learning models achieve strong predictive accuracy for 90-day outcome prediction in acute ischaemic stroke, yet clinical adoption is limited by the misalignment of model explanations with clinicians' reasoning. Motivated… 23 arXiv — NLP / Computation & Language research 23d ago Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment arXiv:2608.05409v1 Announce Type: new Abstract: Large language models typically undergo post-training to align them with safety policies but there exist many sophisticated jailbreaks that sidestep established safeguards. For instance, prior work by Andriushchenko et al. (2025)… 20 arXiv — NLP / Computation & Language research 23d ago From Sports to Safety: Benchmarking Proactive Risk Inference in MLLMs arXiv:2608.05560v1 Announce Type: cross Abstract: Timely anticipation of physical hazards is essential for real-world safety, yet existing MLLM evaluations focus on harmful content or general risks, leaving proactive physical hazard prediction underexplored. Sports provide a… 33 arXiv — NLP / Computation & Language research 23d ago ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment arXiv:2608.06110v1 Announce Type: cross Abstract: This paper presents ECHO (Enhanced Care \& Health Observer), a locally-deployable conversational health assistant for long-term chronic care management. ECHO integrates three complementary software modules developed under shared… 28 arXiv — NLP / Computation & Language research 23d ago Explanations of Large Language Models Explain Language Representations in the Brain arXiv:2502.14671v4 Announce Type: replace Abstract: Large Language Model (LLM) representations are known to align with brain activity during language processing, but it remains unclear what drives this alignment. We test whether explainable AI (XAI) can help answer this: using… 24 Page 4 of 10 · 500 articles ← Newer Older →