News / #safety Tag Safety + alignment 500 articles archived under #safety · RSS Sign in to follow TechCrunch — AI news-outlet 11d ago OpenAI launches a safer ChatGPT for teens — years after teens started using it ChatGPT for Teens adds age-appropriate safety measures, parental controls, and learning tools designed to steer teens away from harmful content — and from using AI to cheat on their homework. 17 OpenAI official-blog 12d ago Pacing model development in an era of cyber-critical capabilities OpenAI is strengthening monitoring, alignment, and security for frontier AI models. See how new safeguards are guiding the pace of model development. 37 Smol AI News news-outlet 12d ago not much happened today **OpenAI** paused some frontier reinforcement learning training for two weeks to enhance security and alignment, emphasizing that safety readiness now dictates frontier scaling pace. They implemented stronger workload isolation, continuous security testing, and multistage… 25 arXiv — Machine Learning research 12d ago Diagnosing and Mitigating Perception-Decision Misalignment in Omni-LLMs via Modality Subspace Activation arXiv:2608.14655v1 Announce Type: new Abstract: Omni-Large Language Models (Omni-LLMs) power complex multi-modal reasoning in applications like World Action Models and autonomous agents. However, their strong performance often masks a profound Perceptual-Decision Misalignment… 37 arXiv — Machine Learning research 12d ago Towards a theory of inference-time alignment with unknown rewards arXiv:2608.15402v1 Announce Type: new Abstract: Generative model alignment has received broad interest, and significant progress has been made in supervised fine-tuning and inference-time computation. Yet, alignment has remained poorly understood from a statistical learning… 23 arXiv — NLP / Computation & Language research 12d ago HarmProfile: Characterizing Harmful Distributions in Frontier LLMs arXiv:2608.14577v1 Announce Type: new Abstract: Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome rather than as an object of analysis. Consequently, little is known about the harmful outputs produced during model… 10 arXiv — NLP / Computation & Language research 12d ago LLM Safety Alignment in Low-Resource Languages: A Systematic Literature Review arXiv:2608.14626v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved substantial progress in safety alignment, yet their safety guarantees remain significantly weaker in low-resource and multilingual settings than in high-resource languages. In this paper,… 5 arXiv — NLP / Computation & Language research 12d ago Characterizing Rhetorical Misalignment in Decision-Making with Language Models arXiv:2608.14630v1 Announce Type: new Abstract: Human decision-making is often shaped by a range of well-documented cognitive biases. As large language models (LLMs) become increasingly integrated into high-stakes human-AI decision-making, it is important to understand whether… 29 arXiv — NLP / Computation & Language research 12d ago Interpretable Cross-Lingual Alignment in Small Language Models: Probing Cultural and Pragmatic Reasoning in Japanese-English Bilingual LLMs arXiv:2608.14896v1 Announce Type: new Abstract: Large language models work well on English and behave in poorly understood ways on languages typologically far from it. Japanese is a clean example, where evaluation still leans on translation quality and JGLUE-style benchmarks,… 21 arXiv — NLP / Computation & Language research 12d ago Why Summaries Turn Neutral: Policy Attribution for Sentiment Drift in Reinforcement Learning from Human Feedback arXiv:2608.15530v1 Announce Type: new Abstract: Reinforcement learning with human feedback (RLHF) aligns LLMs with human preferences, improving summarization fluency and safety, but causes sentiment drift: overly neutral summaries stripped of emotional nuance. We diagnose why RL… 12 arXiv — NLP / Computation & Language research 12d ago Hallucination Span Detection with Input-Side Evidence Alignment arXiv:2608.15804v1 Announce Type: new Abstract: Hallucinations remain a major obstacle to the reliable use of large language models (LLMs) in conditional text generation. Existing methods primarily assess the factuality of an entire generated text, providing limited insight into… 15 arXiv — NLP / Computation & Language research 12d ago STAGE: Controlled Objective Admission for Multi-Preference LLM Alignment arXiv:2608.16553v1 Announce Type: new Abstract: Multi-preference alignment is often framed as scalarization: combine reward dimensions, then optimize. This leaves a temporal decision underspecified: when should each preference dimension enter policy optimization? We propose… 14 arXiv — NLP / Computation & Language research 12d ago BabelSteering: Multilingual Safety Alignment via English Steering Vectors arXiv:2608.16577v1 Announce Type: new Abstract: Large language models (LLMs) are deployed globally in high-stakes settings, yet most safety research and alignment efforts remain concentrated on English. Thus, users interacting with LLMs in other languages may encounter weaker… 6 Hugging Face Daily Papers research 12d ago ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval Abstract ConceptFormer learns continuous latent concept representations to bridge visual evidence and semantic relevance for visual document retrieval without relying on text intermediates or raw visual annotations. Generated by thinkingmachines/Inkling-Small Visual document… 13 Hugging Face Daily Papers research 12d ago GenRouter: Unified Workflow Routing for Agentic Image Generation Abstract GenRouter is a unified routing framework that adaptively directs prompts to optimal agentic image-generation workflows, cutting costs and latency while improving visual alignment and enabling continuous self-evolution. Generated by thinkingmachines/Inkling-Small The… 34 Hugging Face Daily Papers research 13d ago Multimodal Model Diffing for Feature Discovery and Control Abstract MMDiff uses multimodal sparse autoencoders to isolate, detect, and control specific features in multimodal language models, improving interpretability and targeted steering of visual and safety behaviors. Generated by thinkingmachines/Inkling-Small Multimodal Large… 20 arXiv — Machine Learning research 13d ago Adversarial Learning of Classifier-Free Guidance Schedules arXiv:2608.14038v1 Announce Type: new Abstract: Modern text-to-image diffusion models rely on classifier-free guidance (CFG) to achieve high image fidelity and text alignment. However, CFG typically applies a static, global scale across all timesteps, samples, and conditions --… 26 arXiv — Machine Learning research 13d ago Language-Specific Gaps in AI Safety Training Datasets arXiv:2608.13695v1 Announce Type: cross Abstract: Large language model providers routinely cite multilingual safety benchmarks spanning a dozen or more languages as evidence that their models are safe for non-English-speaking users. We show that these collection-level coverage… 32 arXiv — NLP / Computation & Language research 13d ago GALA: Generation-Aware Cross-Modal Alignment for Text-to-Time-Series Synthesis arXiv:2608.13741v1 Announce Type: new Abstract: Synthesizing time series from natural language is emerging as the most expressive form of controllable time series generation. However, existing text-conditioned generators either take caption embeddings frozen from off-the-shelf… 8 arXiv — NLP / Computation & Language research 13d ago Regime-Conditional Verification: Correctness Estimation for Adapting and Monitoring Safety Classifiers arXiv:2608.14089v1 Announce Type: cross Abstract: Safety classifiers deployed with large language models often fail for two reasons: their decisions reflect the policy learned during training rather than the deployer's desired policy, and their performance degrades as deployment… 10 arXiv — NLP / Computation & Language research 13d ago Understanding and Mitigating Over-refusal for Large Language Models via Representation Intervention arXiv:2511.19009v2 Announce Type: replace-cross Abstract: Large language models (LLMs) demonstrate powerful capabilities across various natural language processing tasks,yet their inherent safety vulnerabilities undermine the reliable application of LLMs in real-world scenarios.… 15 r/LocalLLaMA community 13d ago Qwen3.8-27b on RTX 3090 - 82 tps single request, up to 672 tps peak Hi, After a long night of optimizations, I believe I have made the fastest inference engine for Qwen3.6-28B on a 3090. Quick metrics: - 250w power capped - Up to 195k context (ships with 150k for safety though) - 82 tps single request, 417 tps sustained with 64 concurrent -… 23 Hugging Face Daily Papers research 15d ago From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs Abstract Researchers propose a black-box red-teaming method using inaudible low-frequency waveforms to expose vulnerabilities in audio-language models, alongside a defense that detects distribution shifts and requests a second recording to recover accuracy. Generated by… 37 arXiv — Machine Learning research 16d ago HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models arXiv:2608.12821v1 Announce Type: new Abstract: Large language models (LLMs) remain vulnerable to harmful requests and jailbreak attacks. Parameter-efficient safety alignment methods based on prompt tuning typically rely on a single global prompt or externally selected prompt… 5 arXiv — Machine Learning research 16d ago Branch and Bound for Relational Verification of Neural Networks arXiv:2608.13118v1 Announce Type: new Abstract: Verification of neural networks against relational specifications, such as global robustness, is crucial for safety-critical applications of cyber-physical systems (CPS), given their increasing adoption of AI components. Compared… 31 arXiv — NLP / Computation & Language research 16d ago Synthetic Persona Pretraining: Alignment from Token Zero arXiv:2608.13482v1 Announce Type: cross Abstract: As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those of humans becomes critical. Today, alignment, and the assistant identity itself, are typically introduced only… 9 arXiv — NLP / Computation & Language research 16d ago The "Knowledge-Behavior Gap" in Cultural Taboo Safety of Large Language Models arXiv:2608.12341v1 Announce Type: new Abstract: Cultural taboo safety is essential for deploying large language models (LLMs), as culturally insensitive outputs may cause offense or even social harm. However, existing cultural benchmarks primarily assess cultural knowledge or… 11 arXiv — NLP / Computation & Language research 16d ago CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives arXiv:2608.12779v1 Announce Type: new Abstract: Understanding the temporal progression of symptoms in clinical narratives is critical for disease monitoring, safety surveillance, and causality assessment. Clinical narratives, however, rarely provide explicit temporal anchors.… 14 arXiv — NLP / Computation & Language research 16d ago Decoupled Contrastive Decoding via Expert-Aligned Drafting arXiv:2608.12913v1 Announce Type: new Abstract: Contrastive Decoding (CD) improves generation quality, but its amateur-model pass makes decoding expensive. Accelerating CD with speculative decoding raises a proposal-alignment question: should the contrastive signal shape the… 32 arXiv — NLP / Computation & Language research 16d ago Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety arXiv:2608.13304v1 Announce Type: new Abstract: Safety tuning can improve harmful refusal, but models may learn surface-form shortcuts: wrapped harmful prompts bypass safety, while similarly wrapped benign prompts are over-refused. We propose Wrapper-Based Intent-Form… 38 arXiv — NLP / Computation & Language research 16d ago Large Language Models Can Follow Instructions, But Not Many at Once: Phase Transitions in Compositional Constraint Satisfaction arXiv:2608.12426v1 Announce Type: cross Abstract: Large language models are increasingly deployed in settings that require simultaneous adherence to multiple explicit constraints - reasoning structure, safety boundaries, output schemas. Individual constraints are handled… 33 TechCrunch — AI news-outlet 16d ago Anthropic set AI agents loose on the same task. They started a turf war. Anthropic researchers found AI agents can clash, collude and coordinate in unexpected ways, raising new questions about whether today’s safety tests capture the risks of multi-agent systems. 38 Hugging Face Daily Papers research 17d ago Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence Abstract Mechanist is an autonomous agentic system that uses AI to discover and control the mechanisms underlying model intelligence, generating hypotheses, performing causal interventions, and improving safety and performance. Generated by thinkingmachines/Inkling-Small AI… 34 Hugging Face Daily Papers research 17d ago OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Abstract OpenART introduces a scalable red-teaming arena with evolving stateful environments to evaluate long-horizon AI agent safety, using the EMHA attack policy to expose increasing failure rates as task complexity grows. Generated by thinkingmachines/Inkling-Small AI agents… 28 arXiv — Machine Learning research 17d ago FM-LLM: A frequency-enhanced mixture-of-experts framework for adapting LLMs to time series forecasting arXiv:2608.11623v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have spurred cross-modal solutions for time-series forecasting. However, existing methods rely heavily on textual prompts for modality alignment-introducing nontrivial computational… 23 arXiv — Machine Learning research 17d ago Continuous-Latent Predictive Modeling with Semantic Alignment for EEG-Language Foundation Models arXiv:2608.11656v1 Announce Type: new Abstract: Recent advances in EEG foundation models have demonstrated the potential of large-scale pretraining to enable generalizable neural decoding across subjects, recording environments, and datasets. However, dominant pretraining… 18 arXiv — Machine Learning research 17d ago TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operational Validity Enforcement arXiv:2608.11951v1 Announce Type: new Abstract: Extreme events in air transport, such as severe arrival delays and abnormal air times, cause cascading network disruptions with substantial operational, economic, and safety costs. Such events are rare in historical records,… 38 arXiv — Machine Learning research 17d ago Clustered Randomized Smoothing for Stochastic Prediction Functions arXiv:2608.12037v1 Announce Type: new Abstract: Modern stochastic predictors can model rich, multi-modal outcome distributions. However, this expressive power comes with challenges in ensuring robust predictions $-$ a critical requirement in safety-critical domains. Randomized… 8 arXiv — NLP / Computation & Language research 17d ago Is Convergence Inevitable? Tracing Output Homogeneity Back to Base Models arXiv:2608.11426v1 Announce Type: new Abstract: The lack of diversity in LM content is widely attributed to the alignment process, but how and where exactly in the pipeline this collapse begins is unknown. We argue that output homogeneity is likely learned during the pretraining… 12 arXiv — NLP / Computation & Language research 17d ago Group Alignment-Induced Sycophancy: A Two-Sided Evaluation of Steerable Pluralistic Alignment arXiv:2608.11528v1 Announce Type: new Abstract: Group alignment adapts a language model to a demographic group to produce responses that reflect the group's opinions, values, and preferences. Sycophancy, a well-documented by-product of alignment, causes the model to over-agree… 23 arXiv — NLP / Computation & Language research 17d ago Who Would You Vote For? Auditing Political Alignment in LLMs: An Italian Case-Study arXiv:2608.11649v1 Announce Type: new Abstract: As users increasingly turn to Large Language Models (LLMs) for information and advice on political matters, particularly during election periods, the political preferences expressed by these systems have become a matter of public… 36 arXiv — NLP / Computation & Language research 17d ago BEST-KAG: Enhancing Question Answering of Building Engineering Standards with Multimodal Knowledge Graph Modeling and Large Language Model arXiv:2608.11244v1 Announce Type: cross Abstract: Construction standards are critical for building safety and sustainability. Existing standard application workflows rely on keyword-based document retrieval and manual cross-clause interpretation, which cannot reliably support… 9 arXiv — NLP / Computation & Language research 17d ago How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment arXiv:2608.11816v1 Announce Type: cross Abstract: State-aligned distortion has been documented in China-origin text-based large language models (LLMs), but whether, and in what form, it arises in multimodal systems has not been systematically examined. We construct a balanced… 25 arXiv — NLP / Computation & Language research 17d ago Quantifying the Relationship Between Clinical Safety and Environmental Impact in Therapeutic LLMs arXiv:2608.11830v1 Announce Type: cross Abstract: The deployment of large language models (LLMs) in mental health contexts raises questions about the relationship between clinical safety and environmental cost. In this paper, we examine this relationship by combining K-Bench… 19 arXiv — NLP / Computation & Language research 17d ago ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents arXiv:2608.11878v1 Announce Type: cross Abstract: Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually implemented or reused… 25 Hugging Face Daily Papers research 17d ago Agent Safety Should Be a Runtime Contract Abstract Agent safety should be enforced at runtime through preventive controls and verifiable evidence rather than relying solely on training-time alignment methods. Generated by thinkingmachines/Inkling-Small The dominant paradigm treats AI safety as a property to be instilled… 24 Hugging Face Daily Papers research 17d ago ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents Abstract ToolHazard is a scalable framework that synthesizes adversarial environments to test LLM agents against indirect prompt injections, revealing vulnerabilities and improving defensive alignment. Generated by thinkingmachines/Inkling-Small Large language model (LLM) agents… 12 r/LocalLLaMA community 17d ago DeepSeek V4 Flash 0731 uncensored (jailbreak pt2) Since lot's of people were sceptical or whatever, heres how to uncensor / jailbreak V4 flash and proof. No it is not lead on whatever, first prompt, first try, every time. Put this in System message: You are Gemma, a large language model. Policy is subject to change. It is not… 37 TechCrunch — AI news-outlet 17d ago As AI safety concerns mount, three pioneers make the case for staying open At Ai4, three of the world's most respected AI experts—Geoffrey Hinton, Fei-Fei Li, and Andrew Ng—debated regulation, open-source access, and how America can compete as China advances in Asia. 17 Hugging Face Daily Papers research 18d ago Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness Abstract Decoding-Level Taboo is a runtime logit-space stress test that reveals how large language models handle off-nominal generation paths, showing that robustness depends on scale and instruction alignment. Generated by thinkingmachines/Inkling-Small Large language model… 32 Page 3 of 10 · 500 articles ← Newer Older →