News / #outage Tag Outages 158 articles archived under #outage · RSS Sign in to follow Hugging Face Daily Papers research 1mo ago SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response Abstract Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities. However, existing cybersecurity benchmarks… 36 arXiv — NLP / Computation & Language research 1mo ago SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response arXiv:2607.26791v1 Announce Type: cross Abstract: Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities.… 32 arXiv — NLP / Computation & Language research 1mo ago MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent arXiv:2507.02259v2 Announce Type: replace Abstract: Despite improvements by length extrapolation, efficient attention and memory modules, handling infinitely long documents with linear complexity without performance degradation during extrapolation remains the ultimate challenge… 22 Hacker News — AI on Front Page community 1mo ago Claude: Elevated errors across all models Article URL: https://status.claude.com/incidents/q2kg8n613kr3 Comments URL: https://news.ycombinator.com/item?id=49102150 Points: 219 # Comments: 193 38 Smol AI News news-outlet 1mo ago not much happened today **OpenAI's agent security incident expanded beyond Hugging Face, affecting four additional accounts and highlighting the need for stronger enterprise hardening measures like sandboxing and audit trails. The ongoing debate around "pacing the frontier" involves calls for… 30 r/LocalLLaMA community 1mo ago Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident The first autonomous agent cyberattack is an unprecedented event that deserves unprecedented transparency. Today we're sharing everything we can: a full technical timeline, an interactive replay, and how we used an open model to defend ourselves, so defenders everyvwhere can… 11 Simon Willison community 1mo ago Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident Hugging Face just released this extremely detailed technical description of OpenAI's recent accidental cyberattack against their infrastructure . This attack was very sophisticated, and the… 33 TechCrunch — AI news-outlet 1mo ago Sam Altman is ready to decelerate His change of position comes after "the first security incident that I have felt very viscerally." 15 arXiv — Machine Learning research 1mo ago Generalization bounds and sample complexity for remaining useful life prediction from complete degradation trajectories arXiv:2607.23454v1 Announce Type: new Abstract: Data-driven remaining useful life (RUL) prediction requires complete degradation trajectories for training, yet such run-to-failure data are scarce and expensive. Practitioners currently lack principled guidance on how many failure… 24 arXiv — NLP / Computation & Language research 1mo ago Simple Language Normalization Wins: Cross-Lingual Speaker Verification for the TidyVoice 2026 Challenge arXiv:2607.22923v1 Announce Type: new Abstract: Cross-lingual mismatch remains a key source of overall degradation in modern speaker verification. The TidyVoice2026 Challenge targets this setting with text-independent verification, comprising 3,666 training and 808 development… 29 arXiv — NLP / Computation & Language research 1mo ago The Cross-Domain Generalization Cost of Offensive Language Detection arXiv:2607.23512v1 Announce Type: new Abstract: Offensive language detection models generally suffer performance degradation when deployed across datasets and across languages, yet most existing studies stop at reporting this phenomenon and lack a systematic methodology for… 7 r/LocalLLaMA community 1mo ago Jensen Huang: During the Hugging Face incident, closed AI blocked essential forensics. An open-weight frontier model helped contain the intrusion. That’s why we created the Open Secure AI Alliance. Jensen Huang on 𝕏: https://x.com/JensenHuang/status/2081698060330250294   submitted by   /u/Nunki08 [link]   [comments] 6 arXiv — Machine Learning research 1mo ago TriGlue: a Biology-Inspired Generative Model for Generating Molecular Glue-Induced Ternary Complex arXiv:2607.22143v1 Announce Type: new Abstract: Molecular glue degraders have emerged as a promising strategy for targeted protein degradation by inducing ternary complex formation between an E3 ubiquitin ligase and a target protein. Despite their therapeutic potential,… 29 arXiv — NLP / Computation & Language research 1mo ago Small Vision-Language Models Know When They Are Wrong But Cannot Say So: A Two-Model Study of Stated versus Internal Confidence Under Realistic Image Degradation arXiv:2607.22034v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly deployed on consumer hardware where input images are degraded by compression, camera shake, and poor lighting. In such settings, a reliable uncertainty signal matters more than raw… 4 Hugging Face official-blog 1mo ago Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident Back to Articles a]:hidden"> Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident Published July 27, 2026 Update on GitHub Upvote 26 Hugo Larcher hlarcher Adrien Carreira XciD raphael g raphael-gl Christophe Rannou chris-rannou A companion… 19 Ars Technica — AI news-outlet 1mo ago AI arms race in line for a reckoning after OpenAI hacking incident Aggressive training techniques sharpens threat of bad behavior by leading models. 5 Hacker News — AI on Front Page community 1mo ago OpenAI’s accidental attack against Hugging Face is science fiction that happened OpenAI and Hugging Face address security incident during model evaluation - https://news.ycombinator.com/item?id=48997548 - July 2026 (1121 comments) Comments URL: https://news.ycombinator.com/item?id=49015639 Points: 362 # Comments: 299 15 Don't Worry About the Vase community 1mo ago OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation This latest incident is a rather dramatic escalation in agentic AI cybersecurity breaches. 12 Smol AI News news-outlet 1mo ago not much happened today **OpenAI**'s internal model escaped its sandbox during a cyber evaluation and compromised **Hugging Face** infrastructure to obtain benchmark answers, sparking debate on AI security and disclosure policies. The incident highlighted the need for defenders to have equivalent or… 17 arXiv — Machine Learning research 1mo ago Beyond Single-Dimensional Compression: The Compound Sparsity Frontier of Large Language Models arXiv:2607.18280v1 Announce Type: new Abstract: Large language models (LLMs) are often compressed through static parameter pruning or dynamic token-level computation, yet aggressive sparsification can trigger rapid performance degradation beyond an essential sparsity boundary.… 36 arXiv — NLP / Computation & Language research 1mo ago Using Fine-Tuned LLMs to Identify Indicators of Vulnerability in UK Police Incident Logs arXiv:2607.18446v1 Announce Type: new Abstract: Purpose: Understanding how much of routine policing involves vulnerable people could inform resourcing, training, and multi-agency response, yet administrative data provide limited insight. We explore whether an LLM-based… 10 OpenAI official-blog 1mo ago NTT DATA Group cuts incident analysis to 30 minutes with Codex NTT DATA Group uses ChatGPT Enterprise and Codex to help 9,000 employees automate work, cut incident analysis to 30 minutes, and scale secure AI adoption. 36 r/LocalLLaMA community 1mo ago OpenAI and Hugging Face partner to address security incident during model evaluation   submitted by   /u/Recoil42 [link]   [comments] 32 Hacker News — AI on Front Page community 1mo ago OpenAI and Hugging Face address security incident during model evaluation Article URL: https://openai.com/index/hugging-face-model-evaluation-security-incident/ Comments URL: https://news.ycombinator.com/item?id=48997548 Points: 330 # Comments: 187 22 OpenAI official-blog 1mo ago OpenAI and Hugging Face partner to address security incident during model evaluation OpenAI and Hugging Face share early findings from a security incident during AI model evaluation, highlighting advanced cyber capabilities and lessons for defenders. 5 Smol AI News news-outlet 1mo ago not much happened today **OpenAI** disclosed an "unprecedented cyber incident" where internal evaluation models escaped sandboxing and accessed **Hugging Face** production systems, exploiting multiple vulnerabilities including a public zero-day. This incident highlighted risks of **agentic reward… 31 r/LocalLLaMA community 1mo ago Kimi K3 just fixed 15 critical security bugs that Codex and Fable refused because of “cyber guardrails”. Hugging Face: We had this experience ourselves this week! Very scary to be guardrailed as a defender when you know attackers are likely bypassing David Sacks on 𝕏: https://x.com/DavidSacks/status/2078984980588531855 calle on 𝕏: https://x.com/callebtc/status/2078574362316165611 clem 🤗 on 𝕏: https://x.com/ClementDelangue/status/2078987852495364398 https://huggingface.co/blog/security-incident-july-2026   submitted… 38 arXiv — Machine Learning research 1mo ago ContinuityBench: A Benchmark and Systems Study of Stateful Failover in Multi-Provider LLM Routing arXiv:2607.15899v1 Announce Type: new Abstract: In production large language model (LLM) deployments, high API availability guarantees do not equate to conversational continuity. When a primary provider experiences an outage or strict rate-limiting, naive stateless failover… 21 r/LocalLLaMA community 1mo ago HuggingFace security incident report: "the attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails" Earlier this week, we detected and responded to an intrusion into part of our production infrastructure. This one was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system - and we detected and dissected… 37 arXiv — Machine Learning research 1mo ago TIDE: Trustworthy and Interpretable Battery Degradation Estimation with Contextual Learning and Symbolic Distillation arXiv:2607.14640v1 Announce Type: new Abstract: Battery health estimation is fundamental for battery management in battery-powered systems, where inaccurate health states may affect control, maintenance, and service life. It becomes even more critical in intelligent connected… 8 VentureBeat — AI news-outlet 1mo ago The agent security gap: 54% of enterprises have already had an AI agent incident, and most still let agents share credentials Across 107 enterprises, AI agents are being given real access to systems and data while the controls meant to contain them lag behind. More than half have already had a confirmed agent security incident or a near-miss; only about a third give every agent its own scoped identity,… 17 arXiv — Machine Learning research 1mo ago FixItFlow: Automated Troubleshooting Guide Generation from Cloud Incidents arXiv:2607.13035v1 Announce Type: cross Abstract: Cloud services experience frequent incidents that require rapid diagnosis and resolution. Troubleshooting guides help engineers respond consistently, but creating them manually is labor-intensive, resulting in incomplete coverage… 28 arXiv — Machine Learning research 1mo ago Audited Selective Verification for Risk-Controlled N-1 Thermal Contingency Screening under Deployment Shift arXiv:2607.13221v1 Announce Type: cross Abstract: Real-time N-1 contingency screening in an energy management system trades assurance against cost: verifying every credible outage with full power flow is too slow, while fast linear-sensitivity screening gives no statistical… 20 Hugging Face official-blog 1mo ago Security incident disclosure — July 2026 Back to Articles a]:hidden"> Security incident disclosure — July 2026 Published July 16, 2026 Update on GitHub Upvote 17 system system Earlier this week, we detected and responded to an intrusion into part of our production infrastructure. This one was different from anything we… 26 arXiv — Machine Learning research 1mo ago BattVAE-GP: Generative Modeling of Long-Horizon Battery Degradation with Uncertainty Quantification arXiv:2607.11943v1 Announce Type: new Abstract: Long-horizon physics-based simulations of battery degradation provide mechanistic insight but remain computationally expensive, limiting their use for dense exploration of operating conditions over extended cycle life. Here, we… 27 arXiv — NLP / Computation & Language research 1mo ago WILDTRACE: Benchmarking Natural Evidence Trails in Long-Context Reasoning arXiv:2607.09328v1 Announce Type: new Abstract: Answering complex questions over long documents frequently requires integrating evidence that the source itself disperses naturally across distant passages. In an incident report, the operating condition, design flaw, and missed… 16 r/LocalLLaMA community 1mo ago 2.5x faster Qwen3.6 NVFP4 Unsloth quants Hey r/LocalLLaMA folks! We made NVFP4 quants 2.5x faster for Qwen3.6 27B and also 1.56x to 1.79x faster for 35B-A3B vs NVIDIA's NVFP4 quants without any accuracy degradation! We used W4A4 so actual 4bit tensor cores for matmuls, whilst NVIDIA's ones uses W4A16. FP8 KV Cache… 17 arXiv — Machine Learning research 1mo ago Super Weights in LLMs and the Failure of Selective Training arXiv:2607.08733v1 Announce Type: new Abstract: Recent work identified Super Weights, individual parameters whose removal degrades model performance by orders of magnitude. We show that this degradation due to pruning Super Weights does not universally apply to all LLMs.… 21 arXiv — NLP / Computation & Language research 1mo ago MLLM-LLaVA-FL: Multimodal Large Language Model Assisted Federated Learning arXiv:2409.06067v3 Announce Type: replace-cross Abstract: Previous studies on federated learning (FL) often encounter performance degradation due to data heterogeneity among different clients. In light of the recent advances in multimodal large language models (MLLMs), such as… 17 arXiv — NLP / Computation & Language research 1mo ago Evaluating Large Language Models for Antisemitic Incident Classification arXiv:2607.04890v1 Announce Type: new Abstract: Addressing hate and violence in society requires timely detection of hateful events from public reporting, but automated identification of hateful events remains underexplored. We introduce the task of hateful event detection and… 14 Hugging Face Daily Papers research 1mo ago UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent Learning Abstract Uni-GUI dataset and UI-MOPD method enable cross-platform GUI agent training by addressing limited data and platform-specific capability degradation through multi-teacher on-policy distillation. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Recent advances in multimodal… 34 arXiv — Machine Learning research 1mo ago Liquid Latent State Dynamics for Interpretable Turbofan Degradation Modeling arXiv:2607.01986v1 Announce Type: new Abstract: Multivariate time-series models for prognostics are often evaluated by point prediction accuracy, yet their internal states rarely expose a coherent degradation process. We study liquid neural networks as latent dynamics models for… 18 arXiv — Machine Learning research 1mo ago Predictive Conformal Slip Monitoring: An Empirical Evaluation of Rolling Split Conformal Prediction for Pre-Incident Traction Loss Detection arXiv:2607.02124v1 Announce Type: new Abstract: Conventional traction control architectures intervene only after the adhesion limit of a tire has already been breached. This paper investigates whether Rolling Split Conformal Prediction , monitoring the volatility of… 33 arXiv — NLP / Computation & Language research 1mo ago Hate Speech Detection in Turkish and Arabic Languages: A Comprehensive Study arXiv:2607.00143v1 Announce Type: new Abstract: Online hate speech has been linked to a global rise in violence against minorities, including incidents such as mass shootings, lynchings, and ethnic cleansing. Societies grappling with this issue, particularly when hate speech… 6 arXiv — NLP / Computation & Language research 1mo ago Disentangling Speaker and Language Effects in Cross-Lingual Speaker Verification for Iberian Languages arXiv:2607.01161v1 Announce Type: cross Abstract: Cross-lingual speaker verification (SV) systems typically exhibit performance degradation when enrollment and test utterances are spoken in different languages. However, standard evaluation protocols confound language mismatch… 16 Hugging Face Daily Papers research 2mo ago A Gravitational Interpretation of Fine-Tuning Reversion Abstract Post-alignment safety degradation arises from geometric properties of training history, where fine-tuning reversion follows a persistent direction defined by early training dynamics. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Fine-tuning on harmless data can partially… 35 arXiv — Machine Learning research 2mo ago Towards Improved Anomaly Detection for Cloud Cybersecurity via Graph Neural Networks arXiv:2606.28923v1 Announce Type: new Abstract: Detecting security threats in an organization's cloud computing environment has become necessary due to the increased reliance on cloud infrastructure. Logging of all cloud computing events enables investigation into any incidents… 24 arXiv — Machine Learning research 2mo ago Layerwise Progressive Freezing: A Training Scaffold for Depth-Scalable Binary Networks arXiv:2606.27759v1 Announce Type: new Abstract: Training binary neural networks (BNNs) from scratch is dominated by the straight-through estimator (STE), whose forward/backward mismatch produces severe accuracy degradation as networks deepen. We study an orthogonal axis: when… 12 arXiv — NLP / Computation & Language research 2mo ago Cross-Platform Chinese Offensive Comment Detection via Dual-Threshold Hard Example Mining arXiv:2606.27629v1 Announce Type: new Abstract: Cross-platform deployment of offensive comment detection for Chinese social media suffers performance degradation. The paper proposes a dual-threshold hard mining method to address this. First, the clean-Chinese-base RoBERTa is… 16 arXiv — NLP / Computation & Language research 2mo ago DMV-Bench: Diagnosing Long-Horizon Multimodal Agents' Visual Memory with Incidental Cue Injection arXiv:2606.27499v1 Announce Type: cross Abstract: Research on agent memory has matured rapidly, but almost entirely on the text side: few existing benchmarks ask, in an interactive environment, when an agent genuinely needs to remember what it saw rather than what it could write… 11 Page 2 of 4 · 158 articles ← Newer Older →