News / #security Tag Security 500 articles archived under #security · RSS Sign in to follow arXiv — Machine Learning research 9d ago Active Spiking Perception: The Membrane Potential as a Belief State for Anytime 3D Point Cloud Recognition arXiv:2608.19232v1 Announce Type: cross Abstract: Spiking point cloud networks usually scan space in a fixed, input-agnostic order, which leaves the most distinctive resource of spiking computation, the temporal evolution of the membrane potential, unused as a locus of… 14 arXiv — Machine Learning research 9d ago skchange: Fast and Flexible Algorithms for Changepoint Detection arXiv:2608.19767v1 Announce Type: cross Abstract: Skchange is an open-source Python library for detecting structural changes in time series. It implements modern change detection algorithms within a unified and extensible framework. The algorithms are modular and composable, and… 12 arXiv — NLP / Computation & Language research 9d ago A Speech Corpus for Mizo Automatic Speech Recognition: Whisper and SraVaani 1.0 Fine-Tuning with Morphology-Aware Evaluation arXiv:2608.19361v1 Announce Type: new Abstract: This study reports the development of an Automatic Speech Recognition (ASR) system in Mizo, a low-resource language. The development included collecting 17.62 hours of speech data, curating it, and fine-tuning the Mizo ASR system… 4 arXiv — NLP / Computation & Language research 9d ago Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization arXiv:2608.20281v1 Announce Type: new Abstract: Large language models often fail to answer questions about a bounded document collection when the source documents are not retrieved at inference time. We study this setting as document knowledge internalization: converting a fixed… 21 arXiv — NLP / Computation & Language research 9d ago Inducing Task Models from Computer-Use Traces arXiv:2608.20319v1 Announce Type: new Abstract: Naturalistic computer-use traces, passively recorded screenshots and mouse or keyboard actions, are a valuable resource for deriving symbolic, auditable, and reusable models of how everyday work is done. Such models matter as… 8 arXiv — NLP / Computation & Language research 9d ago HARP: Hierarchical Adaptive Ranking with Preference-Adaptive Fusion for Query-Based CVE Prioritization arXiv:2608.19430v1 Announce Type: cross Abstract: Vulnerability prioritization is inherently preference dependent, since the same CVE can receive different remediation priority under different operational preference scenarios. Existing scoring systems and ranking methods… 35 r/LocalLLaMA community 9d ago Is there any interest in a self hosted open source version of manus/perplexity? I have been working on a project for about 3 months because I found Vane to be awful and I wanted something like Manus for myself. I have a pretty solid app that's probably about ready to release, but I'm not sure if there is any desire there. It does research at different… 9 The Information — AI news-outlet 9d ago Walmart’s Slowdown Spotlights Amazon’s Strength This was not the day to be a Walmart shareholder. Stock of the venerable retailer tumbled 9% after the company reported a dip in comparable sales for its U.S. stores in its fiscal second quarter ending July, even as the e-commerce part of its business grew 23%. The results were… 16 r/LocalLLaMA community 9d ago Gonna be huge for US open source   submitted by   /u/pmv143 [link]   [comments] 32 r/LocalLLaMA community 9d ago DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation Their summary: For over a decade, we’ve accepted that end-to-end backprop is the only way to train deep networks. But holding the entire network in memory all at once is why AI training is hitting a resource wall. . We found a new way to break the network into blocks and train… 31 r/LocalLLaMA community 9d ago Any speculation on whether or not Google will announce a new Gemma model at the Gemma SF Celebration tonight? From the Digg article ( https://digg.com/tech/3pf3046j ) “Google Gemma posted that the family of open models has achieved 1 billion downloads. The account is hosting an exclusive evening in San Francisco on August 20 to honor open-source builders, researchers, and contributors.… 12 TechCrunch — AI news-outlet 9d ago Google gives publishers a new way to fight AI-driven traffic losses Google is giving publishers a new button that lets readers make them a preferred source across Search, Discover, and Google News, potentially boosting their traffic as AI search sends fewer clicks to the web. 4 The Information — AI news-outlet 9d ago AT&T is Using Open Source Models to Curb Anthropic Bills Anthropic and OpenAI had better hope more companies don’t follow the example of AT&T. The telecommunications firm plans to keep its employees’ spending on Anthropic and OpenAI models flat in the coming years by using more open-source models such as Nvidia’s Nemotron, according… 5 arXiv — Machine Learning research 10d ago SingularClip: Preventing Spectral Collapse to Maintain Plasticity in Continual and Reinforcement Learning arXiv:2608.18319v1 Announce Type: new Abstract: Neural networks trained on nonstationary tasks frequently lose the ability to fit new targets, a phenomenon referred to as loss of plasticity. We identify a novel source of plasticity loss due to the growing anisotropy of weight… 31 arXiv — Machine Learning research 10d ago LLM-Powered Predictive Decision-Making for Sustainable Data Center Operations arXiv:2608.18503v1 Announce Type: new Abstract: The growing demand for AI-driven workloads, particularly from Large Language Models (LLMs), has raised concerns about the significant energy and resource consumption in data centers. This work introduces a novel LLM-based… 11 arXiv — NLP / Computation & Language research 10d ago Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining arXiv:2608.18089v1 Announce Type: new Abstract: Instruction-tuned models often refuse harmful requests in English but comply with the same requests in Yoruba, Igbo, Igala, and Hausa. This suggests that the refusal mechanism is present in the residual stream but fails to activate… 21 arXiv — NLP / Computation & Language research 10d ago NE-BERT: A Multilingual Language Model for Nine Northeast Indian Languages arXiv:2608.18094v1 Announce Type: new Abstract: Large pretrained language models have demonstrated remarkable capabilities across diverse languages, yet critically underrepresented low-resource languages remain marginalized. We present NE-BERT, a domain-specific multilingual… 15 arXiv — NLP / Computation & Language research 10d ago When Do LLMs Actually Help? Evaluating LLMs as Data Quality Annotators arXiv:2608.18158v1 Announce Type: new Abstract: LLMs have been increasingly used to catch data quality issues automatically, but we know very little about how consistent these judgments actually are. This study tests an LLM on two e-commerce data quality tasks, entity matching… 31 arXiv — NLP / Computation & Language research 10d ago Tangut Word Segmentation under Extreme Resource Scarcity: Integrating Traditional Lexicons and Unlabeled Text arXiv:2608.18437v1 Announce Type: new Abstract: Tangut is an extinct language whose script does not explicitly mark word boundaries. We present the first systematic study of Tangut word segmentation using 2,750 expert-annotated segments(31,893 tokens), traditional lexicons, and… 6 arXiv — NLP / Computation & Language research 10d ago OmniAlign: A Unified Multilingual Aligner for Word and Sentence Alignment arXiv:2608.18474v1 Announce Type: new Abstract: Cross-lingual sequence alignment is fundamental for building and exploiting parallel corpora, spanning mappings from documents and sentences down to words and subwords. Existing tools, however, typically specialize in a single… 7 arXiv — NLP / Computation & Language research 10d ago TranslatePsy-AfriSLM: High-Quality Data Scaling For Low-Resource Machine Translation arXiv:2608.18655v1 Announce Type: new Abstract: The rapid progress in Artificial Intelligence has largely bypassed African languages, creating a digital divide that limits AI adoption on the continent. Recent open-source LLMs systematically underperform on African machine… 20 arXiv — NLP / Computation & Language research 10d ago MemFuse: Multi-Source Memory Fusion from Fragmented Observations arXiv:2608.18704v1 Announce Type: new Abstract: Long-term memory is essential for agents that operate across extended interactions, yet existing memory systems and benchmarks predominantly focus on single-source textual histories. In realistic settings, however, relevant… 10 arXiv — NLP / Computation & Language research 10d ago Do Large Language Models Hallucinate Electric Fata Morganas? arXiv:2608.18816v1 Announce Type: new Abstract: AI hallucinations - that is, outputs which are made up, cannot be verified, or contradict the source material - are generally regarded as an engineering flaw to be dealt with. This paper contends that they also have philosophical… 19 arXiv — NLP / Computation & Language research 10d ago Test-Time Scaling in the Wild: Why Exploitation, Not Exploration, Is the Bottleneck arXiv:2608.18931v1 Announce Type: new Abstract: Test-time scaling (TTS) improves language model outputs by spending additional inference compute - generating multiple candidates, searching over partial sequences, or iteratively refining drafts. These techniques yield large gains… 25 arXiv — NLP / Computation & Language research 10d ago Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale arXiv:2608.19026v1 Announce Type: new Abstract: Released in 2025, Institutional Books: Harvard Library (IB-HL) is a collection of 983,004 volumes (242B o200k_base tokens), originally digitized through Harvard Library's participation in the Google Books Library project. As… 17 arXiv — NLP / Computation & Language research 10d ago Efficient Adaptation of LLMs for Hate Speech Detection in Low-Resource Languages: A Comparative Study on Roman Urdu arXiv:2608.18142v1 Announce Type: cross Abstract: It is challenging to detect hate speech in Low Resource Languages (LRLs) because of the absence of annotated data, the informality of its language structure, and the lack of standardized grammar. A good example of such a… 19 arXiv — NLP / Computation & Language research 10d ago When Readability and Source Retention Diverge: An Evaluability Gap in AI Translation arXiv:2608.19083v1 Announce Type: cross Abstract: Readable AI output can leave an evaluability gap: even when the source is shown, an overall-quality judgment may not reflect what an output preserves. We investigated how source-text condition and output rendering relate to… 5 Hugging Face Daily Papers research 10d ago FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis Abstract FACET constructs executable terminal tasks by preserving source intent and grounding instructions, solutions, and verifiers in a shared repaired environment to enable scalable agent training. Generated by thinkingmachines/Inkling-Small Training terminal agents requires… 15 Hacker News — AI on Front Page community 10d ago Unlocking a locked/deactivated e-waste Cricut Maker Article URL: https://sprocketfox.io/xssfox/2026/07/01/cricut-unlock/ Comments URL: https://news.ycombinator.com/item?id=49365841 Points: 207 # Comments: 51 14 The Information — AI news-outlet 10d ago Nvidia Discusses Funding Its AI Data Supplier Mercor at a $20 Billion Valuation Nvidia has discussed an investment in Mercor, a data labeling provider that helps the chip designer develop its open-source AI models, according to a person with knowledge of the process. The investment would be part of a $20 billion-valuation round. Existing investor General… 13 Hacker News — AI on Front Page community 10d ago Google replaced Git tags for certain source code with obtaining via Google Drive Article URL: https://grapheneos.social/@GrapheneOS/117057099753905023 Comments URL: https://news.ycombinator.com/item?id=49364745 Points: 400 # Comments: 167 21 r/MachineLearning community 10d ago Pandas API for DuckDB, PostgreSQL & ClickHouse — keeping computation inside the database[P] I've been building memFrame — an open-source dataframe API that compiles operations to SQL. The idea: **Python/DataFrame API → SQL → DuckDB / PostgreSQL / ClickHouse** Instead of pulling data into Python and doing everything in pandas, memFrame tries to keep computation inside… 14 r/LocalLLaMA community 10d ago Ornith-1.5 (397B [DeepSWE 56], 35B-A3B, 9B) Aloha! 🌺Introducing Ornith-1.5, a family of open-source LLMs spanning 9B Dense, 35B MoE, and 397B MoE, trained with self-improving strategies. It achieves state-of-the-art performance among open-source models of comparable size and delivers performance comparable to Claude Opus… 28 Hacker News — AI on Front Page community 10d ago A joke domain purchase turned in geopolitical warfare Article URL: https://sprocketfox.io/xssfox/2026/08/19/sondehub-and-war/ Comments URL: https://news.ycombinator.com/item?id=49360015 Points: 327 # Comments: 43 8 Hugging Face Daily Papers research 11d ago EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing Abstract EditBridge enables efficient ultra high-resolution image editing via a diffusion bridge that translates low-resolution edits to high-resolution outputs while preserving source details through sparse attention. Generated by thinkingmachines/Inkling-Small High-resolution… 11 arXiv — Machine Learning research 11d ago Benchmarking Classical and Transformer-Based Models for Document Sensitivity Classification arXiv:2608.16928v1 Announce Type: new Abstract: Automatic sensitivity classification of organizational documents is a critical yet underserved problem, where the consequences of misclassification range from regulatory violations to security breaches. While AI-based approaches… 9 arXiv — Machine Learning research 11d ago Study-Strategy Clusters from EdNet Logs Track Engagement, Not Mastery arXiv:2608.16963v1 Announce Type: new Abstract: Learning analytics often treats unsupervised clusters of intelligent tutoring system (ITS) logs as learner types that should predict learning. We test that assumption on EdNet-KT3. Clustering study-strategy features (resource use,… 17 arXiv — Machine Learning research 11d ago Deep Learning for Cross-Border Electricity Price Forecasting: A Comparative Study arXiv:2608.17091v1 Announce Type: new Abstract: While publicly available electricity market data presents a valuable resource for forecasting research, the field lacks established benchmark datasets for standardized comparison. As a result, many studies have relied on different… 12 arXiv — Machine Learning research 11d ago OOD Detection for EEG-based Machine Learning in High-Risk Environments arXiv:2608.17620v1 Announce Type: new Abstract: Machine learning models for electroencephalography (EEG) analysis show great promise across a wide range of applications, but their deployment in high-risk domains is hindered by their vulnerability to distribution shifts.… 25 arXiv — Machine Learning research 11d ago rl-triton: High-Performance Triton GPU Kernels for Reinforcement Learning Credit Assignment arXiv:2608.17641v1 Announce Type: new Abstract: We present rl-triton, an open-source library of high-performance GPU kernels for reinforcement learning credit assignment, implemented in Triton. The core contribution is a unified associative scan framework that recasts seven… 38 arXiv — Machine Learning research 11d ago Efficient Resource Optimization for Split Federated Learning arXiv:2608.17849v1 Announce Type: new Abstract: Split federated learning (SFL) has emerged as a powerful paradigm for model training at the edge. However, SFL inherently involves discrete decision variables for model splitting and resource allocation, resulting in a challenging… 36 arXiv — Machine Learning research 11d ago Hybrid ML for Lightweight Pre-Route Delay Estimation in Open-Source IC Design arXiv:2608.17914v1 Announce Type: new Abstract: Static Timing Analysis (STA) is a critical step in the design flow of digital integrated circuits, however, obtaining accurate delay estimations can represent a challenge when limited information regarding physical design is… 34 arXiv — NLP / Computation & Language research 11d ago Which Source Wins? Task-Dependent Reliance in Vision-Language Models arXiv:2608.17205v1 Announce Type: new Abstract: Vision-language models (VLMs) combine images and text, but when the two conflict and one becomes harder to read, it is unclear how a model shifts its reliance between them. We study this modality reallocation with a controlled… 29 arXiv — NLP / Computation & Language research 11d ago Q-Interference: Memory-Efficient Phase-Aware Quantum-Inspired Attention arXiv:2608.17288v1 Announce Type: new Abstract: GPT attention measures token compatibility through dot-product similarity. This mechanism is simple, effective, and memory-efficient. But it does not explicitly model whether strong token features should reinforce or suppress one… 11 arXiv — NLP / Computation & Language research 11d ago ArguLens: An Open-Source System for Automated Essay Scoring and Label-Aware Feedback Generation arXiv:2608.17356v1 Announce Type: new Abstract: Most automated essay scoring (AES) systems output a single holistic score without interpretable evidence and rely on closed APIs that introduce data privacy and cost barriers. We present ArguLens, an opensource, locally deployable… 10 arXiv — NLP / Computation & Language research 11d ago Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See arXiv:2608.17744v1 Announce Type: new Abstract: Take three frontier mixture-of-experts models (Alibaba, OpenAI, NVIDIA; 3.6-4.0B active parameters each) and fine-tune them to reason in a low-resource language. On accuracy benchmarks almost nothing happens, and the benchmark… 17 arXiv — NLP / Computation & Language research 11d ago Decomposition Attacks Across Unlinkable Identities: Limits of Stateful Defenses for LLM Services arXiv:2608.17445v1 Announce Type: cross Abstract: Most large language model services use stateless defenses, which judge only the current request, to refuse harmful tasks. Decomposition attacks exploit this limitation by splitting a harmful task into individually permissible… 18 arXiv — NLP / Computation & Language research 11d ago An Analysis of Language Frequency and Error Correction for Esperanto arXiv:2402.09696v3 Announce Type: replace Abstract: Current Grammar Error Correction (GEC) initiatives tend to focus on major languages, with less attention given to low-resource languages like Esperanto. In this article, we begin to bridge this gap by first conducting a… 4 Hugging Face Daily Papers research 11d ago Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection Abstract Researchers evaluate indirect prompt injection risks in DeepSeek Harness using controlled taint and dual judges, finding notable success rates across text and file channels and recommending controls between untrusted content and sensitive actions. Generated by… 16 r/LocalLLaMA community 11d ago What is the best Qwen3.8 27b Abliterated version out there? I'm trying to get a model to reverse engineer / decompile or otherwise reverse to source some of my old c,c++, pascal, and asm demo programs I made from decades ago and I'm constantly met with refusals. It is highly annoying. Does anybody know of a good abliterated/uncensored… 24 Page 3 of 10 · 500 articles ← Newer Older →