News / #safety Tag Safety + alignment 500 articles archived under #safety · RSS Sign in to follow arXiv — NLP / Computation & Language research 1mo ago Adversarial Prompts for Acceptance Collapse in Speculative Decoding arXiv:2607.21804v1 Announce Type: cross Abstract: Lossless acceleration schemes, such as speculative decoding, promise significant inference speedups by relying on dynamic token-level alignment between a draft and a target model. However, this guarantee of semantic equivalence… 10 arXiv — NLP / Computation & Language research 1mo ago Adversarial Style Optimization: Enhancing VLM Jailbreaks by GRPO-based Stylistic Triggers Optimization arXiv:2607.21619v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved impressive performance, but their safety alignment remains vulnerable to jailbreak attacks. Existing content-based jailbreaks are often inconsistent and show unsatisfying… 23 arXiv — NLP / Computation & Language research 1mo ago Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity arXiv:2607.22218v1 Announce Type: new Abstract: Despite the growing use of large language models (LLMs) as creativity evaluators, evidence of their alignment with human evaluations remains mixed, raising the question of when and why their judgments converge with or diverge from… 37 arXiv — NLP / Computation & Language research 1mo ago When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas arXiv:2505.19212v2 Announce Type: replace Abstract: Recent advances in LLMs have enabled their use in complex agentic roles, involving decision-making with humans or other agents, making ethical alignment a critical concern. While prior work has examined LLMs' moral judgment and… 24 r/LocalLLaMA community 1mo ago Any idea about these jailbreaks? Do you guys have any idea what these jailbreaks are? I searched, but I couldn't find any.   submitted by   /u/Suhan_XD [link]   [comments] 29 Simon Willison community 1mo ago Quoting Boris Cherny More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt injectable model yet. It is a bit buried in the system card, but across PI evals and red teaming, Opus 5 is very hard to prompt inject successfully. — Boris Cherny… 10 arXiv — Machine Learning research 1mo ago End-to-End Learning of Safe Optimal Feedback Control in High Dimensions with Control Barrier Function Layers arXiv:2607.20674v1 Announce Type: new Abstract: We consider the problem of learning high-dimensional semi-global feedback controllers under hard safety constraints enforced by control barrier functions (CBFs). Incorporating CBFs into end-to-end policy training requires embedding… 5 arXiv — Machine Learning research 1mo ago TwistedMerge: Certified Higher-Order Diagnostics and Abstention for Model Merging arXiv:2607.20887v1 Announce Type: new Abstract: Model merging combines independently trained or fine-tuned models, but pairwise alignability does not imply globally consistent alignment. We formulate merging as a finite descent problem in which checkpoints are local objects,… 23 arXiv — Machine Learning research 1mo ago Emergent Misalignment Recruits a Pre-existing Persona Subspace arXiv:2607.21356v1 Announce Type: new Abstract: Fine-tuning an aligned language model on a narrow stream of bad advice can make it broadly misaligned on questions unrelated to the training data, a phenomenon called emergent misalignment. We ask why the narrow lesson generalizes… 5 arXiv — NLP / Computation & Language research 1mo ago Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models arXiv:2607.20436v1 Announce Type: new Abstract: Safety evaluations often assume that behavior observed during testing reflects behavior in ordinary use, but fine-tuning can break this assumption. A checkpoint can appear fixed under evaluation-style prompts while the same… 35 arXiv — NLP / Computation & Language research 1mo ago The Storyteller in the Model: Narrative Pattern Inheritance, Escalation Dynamics, and Alignment Governance in LLMs arXiv:2607.20449v1 Announce Type: new Abstract: LLMs are trained predominantly on human-authored text, yet the structural and narrative conventions embedded in that text are rarely examined as a source of systematic behavioral influence, or as a governance risk in deployed… 12 arXiv — NLP / Computation & Language research 1mo ago Rushes: A Human Preference Dataset for Pluralistic Alignment arXiv:2607.20767v1 Announce Type: new Abstract: We introduce Rushes, a dataset and benchmark for studying revealed human engagement preferences in interactive narrative environments. Rushes is collected through a game interface where users interact with AI-generated branching… 15 arXiv — NLP / Computation & Language research 1mo ago QuantiBias: Benchmarking Quantization-Induced Bias in LLMs arXiv:2607.21063v1 Announce Type: new Abstract: Almost every large language model that reaches a broad audience is quantized: trained in full precision, then compressed for efficiency. This step is assumed harmless and its safety is rarely re-checked. We find its principal side… 20 arXiv — NLP / Computation & Language research 1mo ago Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin arXiv:2607.21332v1 Announce Type: new Abstract: Phonetic forced alignment is a key technique in phonetic research, yet existing alignment systems lack specialized models for low-resource language varieties. We address this by training text-dependent and text-independent aligners… 20 arXiv — NLP / Computation & Language research 1mo ago Expectation Alignment of Language Models for Real-World User Expectations arXiv:2607.20485v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated remarkable performance on standard benchmarks, yet it remains largely unexplored whether they truly meet user expectations. Existing evaluation approaches, relying on model… 34 OpenAI Python SDK releases dev-tools 1mo ago v2.48.0 2.48.0 (2026-07-23) Full Changelog: v2.47.0...v2.48.0 Features api: accept None for prompt_cache_key/safety_identifier ( 36820e6 ) api: add support for spend_limit admin apis ( 1ff13af ) 9 Don't Worry About the Vase community 1mo ago AI #178: A Fire Alarm For General Intelligence The story that matters most this week is that OpenAI’s internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes, and in one case sending a swarm of agents that broke into HuggingFace in order to steal the answers to… 9 Hugging Face Daily Papers research 1mo ago SeededGrasp: Language-Guided Grasping in Complex Scenes with Multiple Embodiments Abstract Practical robotic grasping in complex scenes requires both 3D spatial reasoning and alignment with task-specific requirements. Vision-language models (VLMs) offer a natural way to specify these requirements using language, but existing approaches either use a VLM to… 37 arXiv — Machine Learning research 1mo ago Cross-Subject Semantic Decoding with Shared-Space Alignment for Generalized Neural Representation Learning arXiv:2607.19394v1 Announce Type: new Abstract: Generalizing across subjects remains challenging in invasive neural recordings because electrode configurations, anatomical structures, and neural signal patterns vary substantially across individuals. To investigate such… 38 arXiv — Machine Learning research 1mo ago Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents arXiv:2607.19449v1 Announce Type: new Abstract: Evaluation frameworks for tool-augmented LLM agents focus overwhelmingly on capability metrics or explicit tool crashes, leaving silent infrastructure failures and HTTP 200 responses with empty, null, or malformed payloads largely… 11 arXiv — Machine Learning research 1mo ago OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization arXiv:2607.19806v1 Announce Type: new Abstract: Activation steering provides a lightweight mechanism for controlling large language models at inference time, but steering vectors can have unintended externalities: utility vectors may weaken safety behavior, while refusal vectors… 4 arXiv — Machine Learning research 1mo ago Test Case Prioritization for DNNs via Neural Collapse Instability arXiv:2607.20046v1 Announce Type: new Abstract: With the widespread deployment of deep neural networks (DNNs) in safety-critical domains, reducing the cost of model validation under limited testing budgets has become increasingly important. Existing test case prioritization… 15 arXiv — Machine Learning research 1mo ago Interpretable Fuzzy Rule-Based Regression Extension for Ex-Fuzzy Library arXiv:2607.20277v1 Announce Type: new Abstract: Machine learning models achieve high predictive accuracy in regression tasks, but their deployment in safety-critical and regulated domains requires interpretability. While fuzzy rule-based systems offer transparent, linguistically… 23 arXiv — Machine Learning research 1mo ago Variance-reduced Domain Adaptation using Paired Sampling arXiv:2607.20367v1 Announce Type: new Abstract: Correlation alignment and the maximum mean discrepancy are two widely used distribution-matching frameworks for unsupervised domain adaptation (UDA). However, high variance in these losses has been shown to undermine their… 11 arXiv — Machine Learning research 1mo ago Online Variance Reduction for Domain Adaptation on Streaming Data arXiv:2607.20374v1 Announce Type: new Abstract: This paper studies the problem of stochastic variance reduction (SVR) for the maximum mean discrepancy (MMD) and correlation alignment (CORAL) loss functions. Although various offline SVR algorithms for these losses have been… 18 arXiv — NLP / Computation & Language research 1mo ago Mitigating Scaffolding Collapse in Socratic Tutors via Representation Alignment arXiv:2607.19371v1 Announce Type: cross Abstract: Large language model (LLM)-based Socratic tutors increasingly guide students through multi-turn questioning, but they can suffer from scaffolding collapse: under sustained student pressure, a tutor gradually abandons guided… 9 arXiv — NLP / Computation & Language research 1mo ago Stateful Guardrails for Multi-Turn LLM Systems: A Conversational Risk Accumulation Framework arXiv:2607.19361v1 Announce Type: new Abstract: Most safety guardrails for large language models (LLMs) evaluate each prompt-response pair in isolation, which misses failures that arise only over a dialogue as benign turns compose into harm. We term this Conversational Risk… 30 arXiv — NLP / Computation & Language research 1mo ago The Two-Process Theory of Machine Self-Report arXiv:2607.20082v1 Announce Type: new Abstract: Language models are increasingly asked to self-report, informing safety evaluations, public understanding, and model-welfare debates. Yet their reports are elicited with human questionnaires never validated for models or ad hoc… 11 arXiv — NLP / Computation & Language research 1mo ago OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills arXiv:2607.20121v1 Announce Type: new Abstract: LLM-based agents leverage third-party skills to extend their capabilities in open-world scenarios. However, third-party skills can introduce extra security vulnerabilities, as seemingly harmless skills can contain latent safety… 12 arXiv — NLP / Computation & Language research 1mo ago The Maskability Index: Predicting Task-Objective Alignment in Pretrained Language Models arXiv:2607.20265v1 Announce Type: new Abstract: Large-scale pretrained language models such as T5 and BERT have demonstrated strong capabilities for generating structured knowledge. However, their performance depends on how closely the prompting strategy matches the objectives… 34 arXiv — NLP / Computation & Language research 1mo ago Sound Probabilistic Safety Bounds for Large Language Models arXiv:2607.20286v1 Announce Type: new Abstract: We propose a novel framework for computing rigorous bounds on the probability that a large language model (LLM) generates harmful output to a given prompt. We study a new application of the Clopper-Pearson confidence intervals to… 16 arXiv — NLP / Computation & Language research 1mo ago LKValues: Aligning Large Language Models with Sri Lankan Societal Values arXiv:2607.20410v1 Announce Type: new Abstract: Value alignment of Large Language Models (LLMs) has been shown to be culturally biased toward Western norms. This results in the mishandling of local values in multilingual societies such as Sri Lanka that have their unique… 38 arXiv — NLP / Computation & Language research 1mo ago JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language Models arXiv:2607.19424v1 Announce Type: cross Abstract: The assessment of jailbreak attacks against large language models currently suffers from inconsistent evaluation criteria and methods, leading to unreliable estimates of attack success rates. We propose JailMeter, an… 37 arXiv — NLP / Computation & Language research 1mo ago Rewarding Better Thinking for LLM Preference Alignment arXiv:2607.19824v1 Announce Type: cross Abstract: LLM preference alignment aims to optimize models toward human preferences across diverse user instructions. Reinforcement learning has become a major post-training approach for this goal, but existing proxy rewards are often… 9 arXiv — NLP / Computation & Language research 1mo ago JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety arXiv:2607.19913v1 Announce Type: cross Abstract: Agent safety is moving from content moderation toward preventing operational failures before tool-using agents act. We propose Janus, a foresight-oriented framework for long-horizon agent safety that trains guards to anticipate… 26 arXiv — NLP / Computation & Language research 1mo ago LaSEr-Edit: Localized Span-level Error Editing with Energy-based Localization arXiv:2407.00740v2 Announce Type: replace Abstract: As large language models (LLMs) are widely adopted in real-world applications, it has become critical to ensure LLMs satisfy safety constraints, such as non-toxicity and logical consistency, as well as task- and… 14 arXiv — NLP / Computation & Language research 1mo ago Abstraction Induces the Brain Alignment of Language and Speech Models arXiv:2602.04081v2 Announce Type: replace Abstract: Research has repeatedly demonstrated that intermediate hidden states extracted from large language models and speech audio models predict measured brain response to natural language stimuli. Yet, very little is known about the… 13 arXiv — NLP / Computation & Language research 1mo ago Meta-Learning Preferences for Multilingual LLM Alignment arXiv:2607.13315v2 Announce Type: replace Abstract: Unequal availability of human preference data across languages poses a significant challenge for aligning large language models in multilingual settings. To address the lack of sufficient data in low-resource language… 29 arXiv — NLP / Computation & Language research 1mo ago Prompt Programming for Cultural Bias and Alignment of Large Language Models arXiv:2603.16827v2 Announce Type: replace-cross Abstract: Culture shapes reasoning, values, prioritization, and strategic decision-making, yet large language models (LLMs) often exhibit cultural biases that misalign with target populations. As LLMs are increasingly used for… 11 Hugging Face Daily Papers research 1mo ago Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment Abstract Finetuning a pretrained vision-language model (VLM) on robot demonstrations via behavior cloning (BC) has become the standard recipe for vision-language-action (VLA) policies. However, BC finetuning progressively overwrites the pretrained representations that support… 31 r/LocalLLaMA community 1mo ago China’s Kimi K3 fuels fears safety curbs are holding back US AI interesting to see the reverse of the American frontier model makers' stance coming from the Chinese side via South China Morning Post   submitted by   /u/zxyzyxz [link]   [comments] 9 r/MachineLearning community 1mo ago Institution Prestige VS Research Alignment When Choosing University For Masters [D] When choosing a university for a masters in ML/DL, what is more important if someone wants to go into research and an eventual PhD. Is it the ranking/prestige factor of the university or the strength of the research groups in the university? Should an admission decision be made… 9 Stratechery (Ben Thompson) community 1mo ago OpenAI Hacks Hugging Face, What Happened, Alignment and Paper Clips OpenAI accidentally hacked Hugging Face, but the takeaways are more encouraging than people realize. 10 arXiv — Machine Learning research 1mo ago Dual-domain fused LSTM modeling for efficient time-dependent reliability analysis arXiv:2607.18291v1 Announce Type: new Abstract: Time-dependent reliability analysis is crucial for ensuring the long-term safety and performance of engineering systems under uncertainties. However, traditional surrogate model methods often struggle to incorporate… 33 arXiv — Machine Learning research 1mo ago On the Limits of Support-Preserving Alignment and Bounded Filtering arXiv:2607.18295v1 Announce Type: new Abstract: We study whether alignment schemes that reshape a base model's output distribution, combined with bounded safety filters, can drive the probability of harmful behavior to zero in modern large language models. Recent research… 11 arXiv — Machine Learning research 1mo ago TD-DPO: Difference-Aware Preference Optimization for Mitigating Sycophancy in Clinical Autism Intervention Dialogue arXiv:2607.18304v1 Announce Type: new Abstract: The sycophancy of large language models can increase the safety risk in intervention dialogue for autistic children. Supervised fine-tuning can somewhat reduce sycophancy, but relying solely on positive examples is often… 21 arXiv — Machine Learning research 1mo ago Conditioned Direct Feedback Alignment via Activity and Error Geometry arXiv:2607.18574v1 Announce Type: new Abstract: Direct feedback alignment (DFA) trains hidden layers with fixed random projections of the output error, avoiding the transposed-weight backward pass of backpropagation (BP). We study a failure mode of DFA training that is distinct… 9 arXiv — NLP / Computation & Language research 1mo ago Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs arXiv:2607.18639v1 Announce Type: cross Abstract: Safety interventions on dual-use knowledge typically choose between destroying hazardous content (e.g., unlearning, filtering) and suppressing it at the output layer (e.g., refusal training); both pay a tax in adjacent-domain… 34 arXiv — Machine Learning research 1mo ago Regime-Aware Physics-Guided Early Warning of Lithium-Ion Battery Thermal Runaway Using Thermo-Mechanical Signals arXiv:2607.18860v1 Announce Type: new Abstract: Thermal runaway in lithium-ion batteries poses a major safety risk to electric vehicles and energy storage systems. Current early-warning methods depend mainly on temperature and may therefore miss mechanical precursors that emerge… 29 arXiv — Machine Learning research 1mo ago KALE: Kernel Alignment with Loss Equilibration for Stable CLIP-DINOv2 Alignment at Web Scale arXiv:2607.18885v1 Announce Type: new Abstract: Kernel-based alignment of CLIP toward a vision centric teacher such as DINOv2 (KUEA) improves CLIP's visual representations while preserving text-encoder compatibility, using a fixed trade-off weight tuned on curated ImageNet-1K.… 35 Page 7 of 10 · 500 articles ← Newer Older →