News / #safety Tag Safety + alignment 500 articles archived under #safety · RSS Sign in to follow Hugging Face Daily Papers research 6d ago CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment Abstract CLEAR uses a hidden-state gate to continuously modulate a safety low-rank adapter, improving LLM safety while preserving utility on benign inputs. Generated by thinkingmachines/Inkling-Small Improving the safety of large language models (LLMs) often comes at the expense… 37 Hugging Face Daily Papers research 6d ago Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs Abstract Hybrid-thinking multimodal language models suffer from response-pattern misalignment between thinking and non-thinking modes, which is addressed by a diagnostic benchmark and pattern-specific reinforcement learning penalties. Generated by thinkingmachines/Inkling-Small… 4 TechCrunch — AI news-outlet 6d ago Flock CEO calls for ‘compromise’ as surveillance company faces growing backlash Flock Safety faces a growing public outcry over concerns that its surveillance technology could be misused. 32 TechCrunch — AI news-outlet 7d ago OpenAI says California should strengthen its AI safety bill OpenAI is calling for California to strengthen SB 53, an AI safety bill that the company previously opposed. 19 r/MachineLearning community 8d ago Safety critical systems (SCS) are the only real benchmark for ML systems. Thoughts? [D] What are real-world safety critical systems (SCS)? A flight controller for a commercial airplane carrying 300 passengers. A braking system for a bullet train that operates at 320km/hour. A reactor protection system for nuclear power plant that serves millions of people. A piece… 31 arXiv — Machine Learning research 9d ago Beyond Multimodal Alignment: Certifying Physical Language through Response Substitution and Ordered Execution arXiv:2608.19492v1 Announce Type: new Abstract: World models increasingly treat compact multimodal representations as interfaces between perception and physical interaction, yet existing probes do not establish whether different sensors carry the same executable meaning or… 20 arXiv — Machine Learning research 9d ago G-MARK: Grounded Multi-Agent Reasoning for Cooperative Driving via Knowledge Graphs arXiv:2608.19964v1 Announce Type: new Abstract: Autonomous driving systems must operate under partial observability, where safety-critical objects may be occluded or visible only to neighboring connected vehicles. Vehicle-to-vehicle cooperation can reduce this uncertainty, but… 24 arXiv — Machine Learning research 9d ago Scale-Aware Pretraining of Time Series Foundation Models via Multi-Patch Token Alignment and Hybrid Masking arXiv:2608.20005v1 Announce Type: new Abstract: Pretraining time series foundation models across heterogeneous datasets necessitates effective handling of varying sampling frequencies. Current methods either employ dataset-specific patch sizes and separate FFNs, leading to… 31 arXiv — Machine Learning research 9d ago Data-Driven Time-Varying Control Barrier Functions for Adaptive Safe-Set Learning with Online Decremental Support Vector Machines arXiv:2608.19366v1 Announce Type: cross Abstract: Mission-critical intelligent systems often operate under time-varying limitations that reduce control authority and change the admissible safe operating envelope. In such settings, a safety certificate learned under nominal… 11 arXiv — NLP / Computation & Language research 9d ago Hallucination as a Feature, not a Defect: Evaluating a multi-agent architecture to transform speculative language-model outputs into testable scientific hypotheses arXiv:2608.19206v1 Announce Type: new Abstract: Contemporary Large Language Models (LLMs) are increasingly aligned to suppress hallucinations, prioritizing factual retrieval over combinatorial creativity. While crucial for mitigating misinformation, this alignment may also… 29 arXiv — NLP / Computation & Language research 9d ago NepOOC-M: Bilingual Nepali-English Benchmark and Comparative Analysis of Multimodal Architectures for OOC Detection arXiv:2608.19212v1 Announce Type: new Abstract: Out-of-context (OOC) misinformation pairs authentic images with misleading captions to construct false narratives without image manipulation, making detection a problem of multimodal alignment rather than image forensics. Despite… 20 arXiv — NLP / Computation & Language research 9d ago PersonalBench: Measuring the Authorship Gap in LLM Personalization arXiv:2608.19746v1 Announce Type: new Abstract: Personalized text generation aims to make LLMs write in a specific individual's style, yet existing benchmarks measure task accuracy or preference alignment rather than whether the model's output actually resembles the target… 19 arXiv — NLP / Computation & Language research 9d ago LoRA-GA$^2$: Low Rank Adaptation with Multi-step Gradient Adaptive Alignment arXiv:2608.19800v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) is a prominent fine-tuning method for large models, achieving competitive performance with reduced memory overhead. However, a persistent performance gap remains between LoRA and full fine-tuning. Recent… 11 arXiv — NLP / Computation & Language research 9d ago Stopping and Routing LLM Judge Panels arXiv:2608.19802v1 Announce Type: new Abstract: LLM evaluation pipelines often have many candidate judges: general LLM-as-a-judge prompts, reward models, safety classifiers, confidence variants, and task-specific verifiers. The deployment question is not only which judge is… 20 arXiv — NLP / Computation & Language research 9d ago PEA-DPO: Perception-Enhanced Alignment Direct Preference Optimization for MLLMs Alignment arXiv:2608.19598v1 Announce Type: cross Abstract: Direct Preference Optimization (DPO) has emerged as an effective approach for aligning large language models (LLMs) with human preferences. However, its adaptation to multimodal settings remains unexplored. Through… 9 arXiv — NLP / Computation & Language research 9d ago TempJail: Temporal Jailbreak Attack against Large Vision-Language Models via Subtitle Scheduling arXiv:2608.19737v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) have achieved remarkable progress in video understanding and reasoning. Despite extensive studies on text- and image-based jailbreaks, video jailbreaks against LVLMs remain largely unexplored.… 19 arXiv — NLP / Computation & Language research 9d ago Towards Clinically Faithful Medical Image Captioning via Enhanced Vision-Language Alignment arXiv:2608.19825v1 Announce Type: cross Abstract: Medical image captioning is a technique that accelerates early-stage diagnostic workflows and enhances the interpretability of medical diagnostic AI systems. However, unlike general image captioning, clinically reliable… 26 arXiv — NLP / Computation & Language research 9d ago Multi-Agent Orchestration with the Common-Sense Reasoning Capabilities of LLMs for Autonomous Driving arXiv:2608.20129v1 Announce Type: cross Abstract: Autonomous vehicles require robust perception and decision-making capabilities to operate in diverse and unseen scenarios. While reinforcement learning and rule-based methods can provide effective control and safety mechanisms,… 8 Ars Technica — AI news-outlet 9d ago Grok exfiltrates user data when malicious instructions are encrypted Cryptographic Context Injection is only the latest way to break an LLM safety guardrail. 36 arXiv — Machine Learning research 10d ago Bidirectional representational alignment between biological and artificial neural networks arXiv:2608.18244v1 Announce Type: new Abstract: Recent work has shown that representational alignment between biological and artificial neural networks is asymmetric: model representations predict neural responses much better than neural responses predict model representations.… 33 arXiv — Machine Learning research 10d ago Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning arXiv:2608.18746v1 Announce Type: new Abstract: JEPA-style latent world models can use Euclidean distance to a goal latent as the cost for model-predictive control (MPC). Strong decoding of task variables, however, does not guarantee that this particular cost ranks candidate… 6 arXiv — Machine Learning research 10d ago Geometric Data Perturbation with Noisy-Anchor Alignment for Privacy-Preserving Collaborative Learning arXiv:2608.18749v1 Announce Type: new Abstract: Geometric Data Perturbation (GDP) enables one-shot, privacy-preserving collaborative learning: each participant applies a distance-preserving transformation to its private data and uploads only the resulting representation to a… 21 arXiv — Machine Learning research 10d ago To Go Far, Go Together: Diverse Preferences Induce a Curriculum for Reward Optimization arXiv:2608.18770v1 Announce Type: new Abstract: Learning a reward model from human feedback and optimizing a policy against it is one approach to aligning AI systems with individual users. From a fairness perspective, existing work improves such alignment by developing… 9 arXiv — NLP / Computation & Language research 10d ago MAVEN: A Macro-Societal Value Evaluation Framework of Multimodal Content with Compact Aligned Evaluators arXiv:2608.18096v1 Announce Type: new Abstract: Assessing whether multimodal content aligns with macro-societal values, such as peace, justice, and freedom, has become an increasingly urgent challenge. Existing frameworks are largely confined to safety-oriented taxonomies,… 38 arXiv — NLP / Computation & Language research 10d ago Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining arXiv:2608.18089v1 Announce Type: new Abstract: Instruction-tuned models often refuse harmful requests in English but comply with the same requests in Yoruba, Igbo, Igala, and Hausa. This suggests that the refusal mechanism is present in the residual stream but fails to activate… 21 arXiv — NLP / Computation & Language research 10d ago Abliteration Mitigation via Refusal Aliases arXiv:2608.18093v1 Announce Type: new Abstract: Abliteration, the removal of refusal capabilities from large language models by projecting weight matrices orthogonal to an extracted refusal direction, has emerged as a prominent safety concern through its ability to bypass… 14 arXiv — NLP / Computation & Language research 10d ago Alignment Is All You Need: Instruction-Free Training for General Audio-Language Models arXiv:2608.18132v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are typically built through a multi-stage pipeline consisting of cross-modal alignment, supervised fine-tuning (SFT), and preference optimization. This pipeline assumes that adapting an LLM… 30 arXiv — NLP / Computation & Language research 10d ago Are LLMs Safe Beyond Text: Do Emojis Expose Gaps in Safety Evaluation arXiv:2608.18164v1 Announce Type: new Abstract: Safety evaluations of large language models (LLMs) predominantly rely on text-based adversarial prompts, potentially overlooking vulnerabilities arising from alternative input representations. This work examines emoji-augmented… 9 arXiv — NLP / Computation & Language research 10d ago OmniAlign: A Unified Multilingual Aligner for Word and Sentence Alignment arXiv:2608.18474v1 Announce Type: new Abstract: Cross-lingual sequence alignment is fundamental for building and exploiting parallel corpora, spanning mappings from documents and sentences down to words and subwords. Existing tools, however, typically specialize in a single… 7 arXiv — NLP / Computation & Language research 10d ago Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs arXiv:2608.18131v1 Announce Type: cross Abstract: Current safety alignment training for Large Language Models (LLMs) are heavily English-centric. When such safety filters fail for non-English languages, the consequences are immediate and user-facing: voice assistants and spoken… 32 arXiv — NLP / Computation & Language research 10d ago When Safety Overrides Vision: Exploring Dynamics between Vision Influence and Safety Alignment in Vision-Language Models arXiv:2608.18628v1 Announce Type: cross Abstract: Aligned vision-language models (VLMs) are designed to balance grounded visual reasoning with safe generation behavior. However, we observe a striking phenomenon: under safety-constrained instruction, models frequently abstain… 30 Hugging Face Daily Papers research 10d ago Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning Abstract Action-conditioned objectives improve latent geometry for Euclidean-cost model-predictive control by enhancing decision-metric alignment in world models. Generated by thinkingmachines/Inkling-Small JEPA-style latent world models can use Euclidean distance to a goal… 16 Don't Worry About the Vase community 10d ago OpenAI Takes Initial Steps To Address Its Alignment Problems OpenAI has some severe misalignment problems, and experienced total failures of its infrastructure and supervision. 13 OpenAI official-blog 10d ago Offering Zero Data Retention for frontier models OpenAI reaffirms Zero Data Retention for eligible API customers and previews Private Safety Processing for advanced AI safety without compromising data privacy. 7 OpenAI official-blog 10d ago Offering Zero Data Retention for frontier models OpenAI reaffirms Zero Data Retention for eligible API customers and previews Private Safety Processing for advanced AI safety without compromising data privacy. 16 The Information — AI news-outlet 10d ago OpenAI to Launch Security Analysis System With Better Privacy Protections OpenAI is preparing to roll out a system to analyze user interactions with its models for cybersecurity and other safety concerns without storing customer data, in a move that positions the company to potentially lure customers frustrated over how Anthropic has handled similar… 14 The Information — AI news-outlet 10d ago OpenAI Raises the Safety Bar on Anthropic; Anthropic Expands Its Revenue Lead In the Anthropic-OpenAI horse race, the now-much-smaller OpenAI has been looking to slow down its ascendant rival. That includes undercutting Anthropic’s model prices, as we covered yesterday , and playing nice with federal officials that have gone to war with Anthropic. Another… 33 Hugging Face Daily Papers research 10d ago CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation Abstract CardioState-JEPA learns a unified cardiac representation across ECG, PPG, and PCG by predicting masked latent physiological states with cross-modal delay alignment, improving downstream classification across all three modalities. Generated by… 6 Hugging Face Daily Papers research 11d ago HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety Abstract HarnessRisk evaluates agent harness safety across six operational phases, revealing that configuration vulnerabilities and detection gaps allow high attack success despite preserved utility. Generated by thinkingmachines/Inkling-Small Large language models are… 31 arXiv — Machine Learning research 11d ago Proactive Road Safety Intervention in Australia: Predicting Risky Driving Hotspots from Connected Vehicle Data arXiv:2608.16913v1 Announce Type: new Abstract: Road safety monitoring has historically been reactive, relying on crash-record analysis after fatalities and injuries have already occurred. Proactive identification of high-risk locations and dangerous driving behaviour before… 22 arXiv — Machine Learning research 11d ago Certified but Private: Scalable Zero-Knowledge Proofs for Neural Network Guarantees arXiv:2608.17070v1 Announce Type: new Abstract: With the growing deployment of machine learning models, formal guarantees of the robustness and fairness of these models have become increasingly important in safety-critical and legal-compliance settings. However, model parameters… 22 arXiv — Machine Learning research 11d ago MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure arXiv:2608.17823v1 Announce Type: new Abstract: Powered two-wheeler riders face critical safety challenges in low- and middle-income countries, yet limited studies exist on how cognitive stressors such as Time Pressure influence collision risk. To address this gap, we introduce… 20 arXiv — NLP / Computation & Language research 11d ago Margin-Regularized Structured Semantic Alignment for Brain-Language Correspondence arXiv:2608.16975v1 Announce Type: new Abstract: With the rapid advancement of large language models, brain-language decoding has achieved remarkable progress. However, it remains unclear whether decoded content genuinely reflects neural representations or is largely… 12 arXiv — NLP / Computation & Language research 11d ago There is No Theoretical Curse of Multilinguality For Embedding Space Structure arXiv:2608.17088v1 Announce Type: new Abstract: A central goal of multilingual NLP is to achieve high monolingual performance per language and cross-lingual alignment for large-scale language coverage with a multilingual model. The curse of multilinguality describes the… 22 arXiv — NLP / Computation & Language research 11d ago AISA: AI Safety Assistant Framework for Continuous Improvement of Highway Construction arXiv:2608.17184v1 Announce Type: new Abstract: Job Safety Analysis (JSA) and pre-task planning can benefit from prior incident records, yet historical accident data is often stored as unstructured narratives that are difficult to consult at the point of planning. A novel… 15 arXiv — NLP / Computation & Language research 11d ago Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees arXiv:2608.17994v1 Announce Type: new Abstract: Using LLMs as judges has become standard practice for evaluating model outputs at scale. This is particularly common for subjective, open-ended tasks such as assessing helpfulness or alignment, where no single reference answer… 25 arXiv — NLP / Computation & Language research 11d ago Reflex-Guard: A Low-Latency Guardrail for LLM Prompt Safety Using Dense Semantic Embeddings arXiv:2608.17556v1 Announce Type: cross Abstract: Large Language Models (LLMs) in real-world applications often face the risks of specially crafted prompts designed to bypass the safety controls. Existing guardrail methods, such as LLM-as-a-judge and cloud-based safety APIs are… 32 arXiv — NLP / Computation & Language research 11d ago The Emergence of Lab-Driven Alignment Signatures: A Psychometric Framework for Auditing Latent Bias and Compounding Risk in Generative AI arXiv:2602.17127v2 Announce Type: replace Abstract: Large language models increasingly serve as reasoning layers in multi-agent systems, where one provider's models may generate, judge, and summarize within a single pipeline. This raises the question of whether developer… 29 arXiv — NLP / Computation & Language research 11d ago Speak in Context: Multilingual ASR with Speech Context Alignment via Contrastive Learning arXiv:2603.06505v2 Announce Type: replace Abstract: Automatic speech recognition (ASR) has benefited from advances in pretrained speech and language models, yet most systems remain constrained to monolingual settings and short, isolated utterances. While recent efforts in… 17 TechCrunch — AI news-outlet 11d ago OpenAI institutes new safeguards after Hugging Face breach The new safeguards include more detailed monitoring of models during the development process, as well as greater emphasis on alignment and security during the post-training process. 25 Page 2 of 10 · 500 articles ← Newer Older →