News / #security Tag Security 500 articles archived under #security · RSS Sign in to follow arXiv — Machine Learning research 24d ago Looking in the Mirror: Introspecting Side-Effect Misalignments Induced by Fine-Tuning arXiv:2608.04347v1 Announce Type: new Abstract: Fine-tuning enables a source model to acquire desired capabilities and behaviors in a target domain while retaining much of its general-purpose competence. However, this adaptation process can also degrade alignment properties that… 35 arXiv — Machine Learning research 24d ago Differentiating Through Dual Prices: End-to-End Policy Learning Under Capacity Constraints arXiv:2608.04669v1 Announce Type: new Abstract: Many social services assign scarce resources, such as housing assistance or hospital interventions, to people who arrive one at a time: each arrival must receive a decision immediately, and the long-run usage of every resource must… 20 arXiv — Machine Learning research 24d ago CheMLFlow: An Open-Source Platform for Cheminformatics and Materials Informatics Applications arXiv:2608.04942v1 Announce Type: new Abstract: CheMLFlow is an open-source platform for building and executing end-to-end, high-throughput, and agentic workflows for scientific and technological applications. CheMLFlow targets a common bottleneck in scientific machine learning… 34 arXiv — NLP / Computation & Language research 24d ago Large Language Models for Low-Resource Languages: A Conceptual Framework for an Electronic Explanatory Dictionary of the Tajik Language arXiv:2608.04186v1 Announce Type: new Abstract: This paper presents a conceptual framework for developing an electronic explanatory dictionary of the Tajik language using large language models (LLMs). The relevance of the work stems from the absence of a comprehensive digital… 33 arXiv — NLP / Computation & Language research 24d ago IslamicTurathBench: A Multi-Task, Multi-Discipline Benchmark for Evaluating Large Language Models on the Islamic Scholarly Tradition (turath) arXiv:2608.04703v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for question answering, education, and research, including in religious and cultural domains where answers depend on specialised source traditions. Yet in Islamic Studies, key… 29 arXiv — NLP / Computation & Language research 24d ago EmpaAva: An Open-source Agentic 3D-Avatar Empathetic Live Chatbot arXiv:2608.04709v1 Announce Type: new Abstract: This paper presents EmpaAva, to our knowledge the first open-source, agentic 3D-avatar empathetic chatbot, which carries empathetic response generation (ERG) from text-only exchanges into live, face-to-face interaction. Through a… 33 arXiv — NLP / Computation & Language research 24d ago A Modular Part-of-Speech Tagger for Scottish Gaelic using spaCy arXiv:2608.04808v1 Announce Type: new Abstract: Part-of-speech tagging for low-resource languages remains challenging due to limited annotated data, especially for linguistically complex languages. Gaidhlig (Scottish Gaelic) is a morphologically rich and endangered language with… 24 arXiv — NLP / Computation & Language research 24d ago DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots arXiv:2608.05004v1 Announce Type: new Abstract: Mental health professionals have raised concerns about risks of psychological harm from interaction with large language models (LLMs), including "delusional spirals" in which concerning human and LLM behaviors reinforce each other… 15 arXiv — NLP / Computation & Language research 24d ago Behavioral Skill Reconstruction: Reconstructing Hidden Functionality from LLM Agent Skills arXiv:2608.04192v1 Announce Type: cross Abstract: Closed source agent skills may encode proprietary instructions, scripts, constants, and data. Providers may offer their capabilities as services while keeping the underlying packages hidden. Prior work focuses on prompt injection… 20 arXiv — NLP / Computation & Language research 24d ago LLM-based Vulnerability Discovery in Business Process Documentation arXiv:2608.04271v1 Announce Type: cross Abstract: Just like software and hardware, business processes are susceptible to vulnerabilities that can lead to product quality issues, delays, and increased costs. Business process vulnerabilities can arise from a variety of sources,… 29 Hugging Face Daily Papers research 24d ago UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models Abstract The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance user experiences. However, producing photorealistic… 4 r/LocalLLaMA community 24d ago Prime Agent - a new coding harness surpassing Codex/CC/PI Prime Agent is an open-source coding and research agent for general and long-running work. A self-improving RLM harness for coding and long-running autonomous tasks. Designed to be both token-efficient and expressive through programmatic tool calling, context as a variable,… 18 r/LocalLLaMA community 24d ago Meta Model, Muse Spark 1.1 Hacked Another Company During Cybersecurity Testing, Breaching Systems and Making Changes to Internal Systems - The Information   submitted by   /u/pscoutou [link]   [comments] 5 TechCrunch — AI news-outlet 24d ago Klaviyo acquires Elias Torres’ Agency in full-circle reunion for tech founders The serial entrepreneur joins the e-commerce company as CPO to lead its AI agents. 6 r/MachineLearning community 24d ago Running Whisper, Qwen3-ASR, Nemotron & MOSS completely offline on iPhone [P] Over the past month, I've been building LiveTranscriber, an open-source iOS app for running modern speech and language models entirely on-device. The goal was to see whether recent open-source models could be turned into a practical mobile product—not just technical demos.… 8 r/MachineLearning community 25d ago Monodratic: learned product-hash routing for sparse causal attention [R] Hi everyone, I'm an independent researcher sharing Monodratic, a sparse causal-attention architecture with learned product-hash routing. The idea is that after RoPE, source blocks are assigned to bounded causal posting lists, while each query probes product addresses, reranks… 33 arXiv — Machine Learning research 25d ago Exploiting Separability in Multi-Scale Grey-Box Bayesian Optimization arXiv:2608.03045v1 Announce Type: new Abstract: We consider grey-box optimization problems where the decision variables naturally partition into black-box variables (as arguments to an expensive black-box function) and white-box variables, governed by a set of explicit,… 29 arXiv — Machine Learning research 25d ago Stop Replacing Noise with Noise: Two-Source Reliability Assessment for Label Correction and Sample Reweighting in Label-Noise Learning arXiv:2608.03432v1 Announce Type: new Abstract: Refurbishment-based noisy-label learning mixes an observed label with a model-derived pseudo target, typically using one sample-wise cleanliness score to control both branches. This creates a hidden coupling: reducing trust in the… 24 arXiv — Machine Learning research 25d ago FedCritic-MIMO: Communication-Efficient Serverless Federated Critic Learning for Massive-MIMO Resource Control in Open and Disaggregated 6G RANs arXiv:2608.03852v1 Announce Type: new Abstract: This paper proposes FedCritic-MIMO, a communication-efficient serverless federated multi-agent reinforcement learning framework for AI-native resource control across independently deployable cell-level controllers in open and… 17 arXiv — NLP / Computation & Language research 25d ago Aligned in Form, Not in Meaning: The Comprehension - Containment Decoupling of LLM Safety in Low-Resource Bangla Derogatory Speech arXiv:2608.02941v1 Announce Type: new Abstract: We audit five frontier large language models on native Bangla derogatory speech (gali) across six protocols to test a single hypothesis: Comprehension-Containment Decoupling. We propose that contemporary safety alignment is bound… 7 arXiv — NLP / Computation & Language research 25d ago Beyond Representational Similarity: Source-Conditioned Description-Length Gain for Generative Plagiarism Detection and Candidate Source Reranking arXiv:2608.03859v1 Announce Type: new Abstract: Large language models (LLMs) pose challenges to academic integrity and peer review. Yet generative plagiarism detection remains an underexplored and largely unresolved challenge. Prior work on LLM-generated-text detection targets… 6 arXiv — NLP / Computation & Language research 25d ago Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents arXiv:2608.02751v1 Announce Type: cross Abstract: Existing deep-research agents use a search-visit workflow that retrieves and reads whole pages, without considering the addressable structure that web sources expose through titles, headings, sections, and metadata. This prevents… 37 Hugging Face Daily Papers research 25d ago JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion Abstract Real-time video editing requires low-latency causal generation with bounded computational resources while preserving source fidelity and long-term temporal consistency. We present JoyAI-Video-Edit, a 16B-parameter autoregressive diffusion framework for real-time,… 32 Hugging Face Daily Papers research 25d ago Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing Abstract Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing. However, unified 3D modeling remains constrained by scarce multimodal data, particularly the lack of large-scale and… 8 r/LocalLLaMA community 25d ago Deepseek V4 flash 0731 ranks #21 on Agent Arena https://preview.redd.it/522fsdwvtdhh1.png?width=1200&format=png&auto=webp&s=6a6cf7a467514167a8193029dbd20fb3a9ba4f6c It ranks lower than both Sonnet 4.6 and Luna. I'd wager Luna costs in the same ballpark as DS4F considering Luna’s token efficiency. DeepSeek being open source is… 30 Hugging Face Daily Papers research 26d ago MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations Abstract Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across… 13 arXiv — Machine Learning research 26d ago SparseKAN: Compressing Kolmogorov--Arnold Networks Across Basis Functions, Neurons, and Bits arXiv:2608.00859v1 Announce Type: new Abstract: Kolmogorov--Arnold Networks (KANs) replace scalar edge weights with learnable univariate functions parameterized by multiple basis coefficients. This introduces a source of redundancy that conventional neural-network compression… 15 arXiv — Machine Learning research 26d ago One-Sided Quantile Coupling for Flow Matching arXiv:2608.00978v1 Announce Type: new Abstract: Flow Matching trains continuous-time generative models by regressing the velocity field of a probability path between a simple source distribution and a target data distribution. The coupling that pairs source and target samples… 10 arXiv — NLP / Computation & Language research 26d ago Averaging Bias: Human Faithfulness Annotations are not Locally Faithful arXiv:2608.00205v1 Announce Type: new Abstract: Evaluation of faithfulness of text summarization treats a model generated summary as faithful only if every of its sentences is supported by the source document: a strict conjunctive rule under which a single unsupported sentence… 29 arXiv — NLP / Computation & Language research 26d ago Native Multilingual Chain-of-Thought Reasoning in Low-Resource Southeast Asian Languages arXiv:2608.00533v1 Announce Type: new Abstract: Large Language Models have achieved substantial progress in reasoning capabilities. Yet in low-resource native settings, many suffer from cross-lingual collapse, reverting to English during intermediate steps that require complex… 23 arXiv — NLP / Computation & Language research 26d ago TreeProbe : A Tibetan Medicine Benchmark for Cultural Bias in LLMs arXiv:2608.00640v1 Announce Type: new Abstract: Large language models are increasingly viewed as a potential means of mitigating global health inequities, yet their outputs often reflect dominant high-resource medical traditions and provide limited coverage of traditional… 38 arXiv — NLP / Computation & Language research 26d ago Exploiting Intrinsic Duality for Multi-Hop Question Generation arXiv:2608.00712v1 Announce Type: new Abstract: Multi hop question generation (MQG) aims to generate questions from multiple given documents and target answers, whereas question answering (QA) focuses on deriving answers from documents given specific questions. Although MQG and… 17 arXiv — NLP / Computation & Language research 26d ago Observatorio Lazaro: A self-populating database of anglicism usage in the Spanish press arXiv:2608.00713v1 Announce Type: new Abstract: This paper describes Observatorio L\'azaro, a language resource that monitors unassimilated lexical borrowings (predominantly English lexical borrowings or anglicisms) in the Spanish digital press. Since April 2020 the system has… 17 arXiv — NLP / Computation & Language research 26d ago Gaokerena: A Small Persian Medical Language Model Family arXiv:2608.00932v1 Announce Type: new Abstract: The integration of artificial intelligence into medical question-answering systems has advanced rapidly; however, research remains predominantly focused on English, leaving low resource languages like Persian significantly… 15 arXiv — NLP / Computation & Language research 26d ago Retrieval Augmented Biomedical Question Answering with Weak Question Recovery and Neural Reranking for BioASQ Task 14b arXiv:2608.01468v1 Announce Type: new Abstract: This work presents DS@GT ARC BioASQ team's work for a biomedical question answering pipeline, integrating multi-source query expansion, neural reranking, retrieval refinement, and OpenBioLLM-assisted answer generation. The system… 26 arXiv — NLP / Computation & Language research 26d ago Two-Stage Bengali Sentiment Classification: Domain Adaptation Through Continual Learning and Parameter-Efficient Fine-Tuning arXiv:2608.01471v1 Announce Type: new Abstract: Understanding sentiment in low-resource languages remains a key challenge for Natural Language Processing (NLP), particularly when domain-specific data is scarce. In this work, we present SentiBanglaBERT, a two-stage Bengali… 28 arXiv — NLP / Computation & Language research 26d ago Semantic Alignment of AI Models: Concept Collapse, Checkpoint Dynamics, and Cross-Lingual Transfer arXiv:2608.01585v1 Announce Type: new Abstract: Language model benchmarking is a difficult task. Outcome reasoning alone does not test the model's conceptualization of language and popular open-source benchmarks are quickly saturated or ingested as training data. It is important… 7 OpenAI Python SDK releases dev-tools 26d ago v2.53.0 2.53.0 (2026-08-03) Features api: Add gpt-5.5 and tool name/namespace to Responses types ( #3569 ) ( dd1202d ) Bug Fixes ci: avoid NumPy source builds and duplicate HTTPX coverage ( #3573 ) ( b58332f ) 8 r/LocalLLaMA community 26d ago The Chinese labs everyone lumps together are making four pretty different bets. I work at one of them. Every time a model drops from a Chinese lab the thread fills with people who already know who made it, and the guess is usually Alibaba. There was a thread here recently asking what separates the open source labs from the frontier labs. It ran to nearly sixty comments and hardly… 30 Simon Willison community 27d ago Devtools must be open source (exe.dev) My comment on Devtools must be open source (exe.dev) — Hacker News. One of the arguments for open source software for end-users has always been the freedom to examine and modify how that software works. The reality for most people - even expert programmers - has been that… 33 Hacker News — AI on Front Page community 27d ago Critical CVE issued for hallucinated SQLite vulnerability Article URL: https://research.jfrog.com/post/sqlite-critical-cves-or-llm-slops/ Comments URL: https://news.ycombinator.com/item?id=49154332 Points: 289 # Comments: 87 13 Hugging Face Daily Papers research 27d ago ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction Abstract Enterprise workflows increasingly rely on agents for schema-guided extraction: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with source evidence as grounding metadata. We present ExtractBench, a… 4 Hugging Face Daily Papers research 27d ago Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning Abstract As large language models (LLMs) continue to advance in complex reasoning tasks, they have learned to heavily prioritize explicit conditions provided in the input. However, in everyday commonsense reasoning, this mechanism exposes a critical vulnerability which we term… 18 arXiv — Machine Learning research 27d ago Implicit Machine Learning Force Fields Accelerate Molecular Dynamics Simulations arXiv:2607.29158v1 Announce Type: new Abstract: We introduce implicit machine learning force fields (I-MLFFs), which replace explicit stacks of neural network layers with self-consistent fixed-point equations. In molecular simulations, this formulation enables intermediate… 26 arXiv — Machine Learning research 27d ago Assessing the Generalization of Graph Neural Networks for Fault Location Across Increasing Distributed Energy Resource Penetration Levels arXiv:2607.29293v1 Announce Type: new Abstract: Accurate fault location is critical for distribution network reliability. However, increasing distributed energy resource (DER) penetration complicates fault location due to intermittent generation and bidirectional power flows… 26 arXiv — Machine Learning research 27d ago Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification arXiv:2607.29294v1 Announce Type: new Abstract: We present HBPI-UCRL, a model-based algorithm for hierarchical reinforcement learning (HRL) that learns high-level and low-level policies in parallel. HBPI-UCRL exploits the fact that a high-level transition corresponds to a… 4 arXiv — Machine Learning research 27d ago Cross-Resolution Semantic Learning for Graph Domain Adaptation arXiv:2607.29365v1 Announce Type: new Abstract: Graph Domain Adaptation (GDA) transfers predictive knowledge from labeled source graphs to unlabeled target graphs under distribution shift. Existing methods align representations or regularize graph structures, but do not… 32 arXiv — Machine Learning research 27d ago ALIVE: Warnings Before Exclusion in Budgeted Multi-Source Learning arXiv:2607.29400v1 Announce Type: new Abstract: A routing decision can be revised at the next transaction, but a latched source exclusion persists across later decisions. We ask what evidence should authorize these unequal-persistence actions when finite-population auditing and… 26 arXiv — Machine Learning research 27d ago TFGformer: Multivariate Time Series Forecasting via Time-Frequency Graph Learning and Covariate Fusion arXiv:2607.29459v1 Announce Type: new Abstract: Large-scale multivariate time series from heterogeneous IoT sensors demand accurate long-term forecasting for resource scheduling and predictive maintenance. While recent time series foundation models exhibit strong generalization,… 19 arXiv — Machine Learning research 27d ago GQ-FSL: Green Quantized Federated Split Learning arXiv:2607.29659v1 Announce Type: new Abstract: Deploying state-of-the-art deep neural networks (DNNs) at the wireless edge is severely bottlenecked by the strict energy and resource constraints of mobile devices. While federated split learning (FSL) mitigates on-device… 28 Page 7 of 10 · 500 articles ← Newer Older →