News / #security Tag Security 500 articles archived under #security · RSS Sign in to follow Simon Willison community 11d ago Mojo🔥 is now open source Mojo🔥 is now open source Mojo🔥 is now open source The Mojo programming language has been promising an open source release since May 2023 . Last week they shipped their 1.0 and today they have followed through on that original promise, releasing the compiler and toolchain under… 7 TechCrunch — AI news-outlet 11d ago OpenAI institutes new safeguards after Hugging Face breach The new safeguards include more detailed monitoring of models during the development process, as well as greater emphasis on alignment and security during the post-training process. 25 Vercel — AI dev-tools 11d ago Vercel for Platforms can now deploy from your users' GitHub repositories Teams building on Vercel for Platforms can now create deployments directly from their users' GitHub repositories, without requiring them to install the Vercel GitHub App. When creating a deployment , pass a gitAccessToken alongside gitSource . Vercel uses the token to retrieve… 22 Hacker News — AI on Front Page community 11d ago Mojo is now open source Article URL: https://www.modular.com/blog/mojo-open-source Comments URL: https://news.ycombinator.com/item?id=49348079 Points: 292 # Comments: 63 15 MIT Technology Review — AI news-outlet 12d ago We still don’t know how people are really using AI AI companies like Anthropic and OpenAI regularly publish reports on how people are using products like Claude and ChatGPT, but they only release the data they want us to see, AI researchers say.  “There is no independent source to corroborate it,” says Anka Reuel, a… 30 Hugging Face Daily Papers research 12d ago Advancing Open and Reproducible Relational Learning: RelArena-α, TabPFN-Rel and RPI Abstract Prior Labs released open-source tools including a unified relational benchmark framework, a TabPFN-based relational model, and a model-agnostic predictive interface to advance reproducible relational learning. Generated by thinkingmachines/Inkling-Small This first… 21 Hugging Face Daily Papers research 12d ago GRNEdit: Efficient General Video Editing from a New Binary-Evidence Perspective in Generative Refinement Networks Abstract GRNEdit is a lightweight two-stage framework that models video editing intent via binary semantic decisions and source evidence, achieving strong results with minimal parameters. Generated by thinkingmachines/Inkling-Small Instruction-based general video editing seeks… 16 r/LocalLLaMA community 12d ago Qwen 3.8 27b vs Deepseek Flash Hey Guys, What amazing weeks it has been for open source releases. I was really impresssed by DS flash final checkpoint and i have been playing around with it until qwen 3.8 released. I checked the benckmarks, and I dont know what to think anymore how can such a small model… 6 arXiv — Machine Learning research 12d ago Training and Evaluating Ethical Reinforcement Learning Agents on Per-Episode Distributions arXiv:2608.14642v1 Announce Type: new Abstract: Reinforcement Learning (RL) agents trained on a single reward signal exploit the gap between the designed reward and the intended behavior. This is particularly a problem when we are trying to imbue ethical behavior into RL agents.… 11 arXiv — Machine Learning research 12d ago Iterative Refinement Diffusion for Super-Resolved Data Assimilation of Multiscale Physical Systems arXiv:2608.14744v1 Announce Type: new Abstract: Recovering high-resolution states from sparse, low-resolution observations is a central challenge in scientific machine learning and data assimilation. Classical data assimilation exploits temporal information through… 5 arXiv — Machine Learning research 12d ago ER-KANs: Efficient and Robust Kolmogorov-Arnold Networks for Data-Scarce Scientific Machine Learning arXiv:2608.14773v1 Announce Type: new Abstract: The efficient-KAN literature---covering Chebyshev, wavelet, and radial-basis-function variants of the original Kolmogorov-Arnold Network---has been benchmarked almost entirely on clean data. We show that this choice conceals a… 32 arXiv — NLP / Computation & Language research 12d ago LLM Safety Alignment in Low-Resource Languages: A Systematic Literature Review arXiv:2608.14626v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved substantial progress in safety alignment, yet their safety guarantees remain significantly weaker in low-resource and multilingual settings than in high-resource languages. In this paper,… 5 arXiv — NLP / Computation & Language research 12d ago BengaliMCQ: Automatic Generation and Answer Prediction of Academic Multiple-Choice Questions in a Low-Resource Language arXiv:2608.15547v1 Announce Type: new Abstract: Traditional retrieval-augmented generation (RAG) frameworks process documents without attending to their hierarchical structure, leading to poor performance, especially in low-resource languages such as Bengali. To address this, we… 28 arXiv — NLP / Computation & Language research 12d ago TaoLive Digital Avatar Agent Technical Report: Training Agents to Evolve with Their Harness arXiv:2608.15763v1 Announce Type: new Abstract: AI-powered digital-avatar streamers in live e-commerce must answer product questions, engage viewers, and execute changing business strategies in real time. This requires low latency, factual and effective replies, and rapid… 35 arXiv — NLP / Computation & Language research 12d ago $R^3$-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets arXiv:2608.16033v1 Announce Type: new Abstract: In cognitive science, resource rationality asks how an agent should allocate limited computation to maximize expected value. Most reasoning and agent benchmarks use independent per-task budgets; existing shared-budget studies do… 12 arXiv — NLP / Computation & Language research 12d ago DuplexGen: Decoupling Content, Timing, and Acoustics for Synthetic Dialogue Speech arXiv:2608.16053v1 Announce Type: new Abstract: Synthetic conversational speech has become an important resource for developing and evaluating conversational speech systems. However, existing dialogue synthesis pipelines typically generate dialogue content first and then insert… 28 arXiv — NLP / Computation & Language research 12d ago Skill2Query: Exploiting Skill Structure to Generate Pseudo-Queries for Agent Skill Retrieval arXiv:2608.16071v1 Announce Type: new Abstract: Pseudo-query generation can alleviate the supervision bottleneck for agent skill retrieval, but existing document-level approaches typically leave the rich internal relations among capabilities, parameters, and usage examples… 31 arXiv — NLP / Computation & Language research 12d ago IndicQE-APE: A Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages arXiv:2608.16344v1 Announce Type: new Abstract: Indic quality estimation (QE) and automatic post-editing (APE) data is spread across separate releases, so no single resource supports training and evaluation across tasks and language pairs on one footing. We consolidate the WMT… 18 arXiv — NLP / Computation & Language research 12d ago When Context Misleads: Intent-Guided Decoding for Robust Retrieval-Augmented Generation arXiv:2608.16515v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) improves large language models by grounding generation in external evidence, but it also introduces a source trust problem: retrieved context may be useful, irrelevant, or even misleading.… 37 arXiv — NLP / Computation & Language research 12d ago PCA-guided Activation Scaling for Monotonic Bidirectional Control over LLM Sycophancy arXiv:2608.16650v1 Announce Type: new Abstract: Large language models (LLMs) exhibit sycophancy, a tendency to agree with user beliefs regardless of factual accuracy. This can reinforce misconceptions, but eliminating it entirely risks over-correction against valid opinions.… 34 arXiv — NLP / Computation & Language research 12d ago Semantic Bandits: In-Context Exploration-Exploitation is Biased by Semantic Priors arXiv:2608.16707v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as decision-making agents in settings that require sophisticated environmental exploration. However, existing work has raised questions about how LLMs actually balance… 19 r/LocalLLaMA community 12d ago Optimizing Qwen3.6 / Qwen3.8-27B on 16GB VRAM: Complete Benchmark Results and Setup Guide (~30-50tps at 32k to 72k context) This post was made with AI. I tried to remove as much slop as possible and keep it straight to the point to save your time as I know how annoying AI slop posts can be, but I still wanted to retain all the details so it can be used as a resource for comparison with other future… 38 Hugging Face Daily Papers research 12d ago VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End? Abstract A unified framework benchmarks and trains multimodal agents that infer intent, plan 3D scenes, invoke tools, and reflect on feedback, revealing that reinforcement learning improves open-source models beyond closed-source frontiers. Generated by… 14 arXiv — Machine Learning research 13d ago SAGE: Surrogate-gradient Adaptation via Attention-Guided Entropy for Spiking Transformers arXiv:2608.13702v1 Announce Type: new Abstract: Spiking neural networks (SNNs) offer an energy-efficient alternative to conventional deep neural networks by exploiting sparse event-driven computation, but their training remains challenging because the non-differentiable spike… 30 arXiv — Machine Learning research 13d ago Polar Code Based Federated Learning: Convergence Analysis and Resource Allocation arXiv:2608.13961v1 Announce Type: new Abstract: Federated learning (FL) enables collaborative model training across distributed devices without sharing raw data; however, it faces significant communication bottlenecks and channel impairments in practice. Conventional network… 8 arXiv — Machine Learning research 13d ago Overcoming Shortcut Learning in Graph Neural Networks through Active Explanation Guidance arXiv:2608.14121v1 Announce Type: new Abstract: Graph Neural Networks (GNNs) can solve prediction tasks by unintentionally exploiting shortcuts---that is, edges, nodes, and features that correlate with but are not causal for the prediction---which compromise their reliability in… 23 arXiv — Machine Learning research 13d ago AutoSchema: Live Schema Grounding for Agentic Text-to-Sparql over Heterogeneous Knowledge Graphs arXiv:2608.14228v1 Announce Type: new Abstract: Life science knowledge graphs make large collections of structured data available through SPARQL, but each resource uses its own schema, identifiers, and links. TogoMCP helps language model agents query these resources by providing… 11 arXiv — Machine Learning research 13d ago Convex losses and their applications to SVM, SVR, and Shallow Neural Networks arXiv:2608.14288v1 Announce Type: new Abstract: We propose multiple new convex losses for SVM and Neural Networks, applied to binary classification tasks. While there are practical limitations in exploiting them with the dual SVM models, we are able to use them with SVM primal… 7 arXiv — Machine Learning research 13d ago Detecting Contaminated Code-Generation Prompt Batches via Influence Functions arXiv:2608.14303v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for code generation, yet they remain vulnerable to prompts that elicit insecure implementations. Existing defenses typically rely on predefined threat models or known vulnerability… 12 arXiv — NLP / Computation & Language research 13d ago A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure arXiv:2608.13626v1 Announce Type: cross Abstract: A hidden state signal can be decodable or causally usable without supporting a reusable action map. We test whether action maps fitted without a source reach its natural post-action activation and compose. We organize the tests… 26 arXiv — NLP / Computation & Language research 13d ago BM25-Augmented Many-Shot Translation for Low-Resource North-Eastern Indian Languages arXiv:2608.13722v1 Announce Type: new Abstract: This paper describes the University of Florida Gators submission to the WMT26 Low-Resource Indic Language Translation shared task. We adapt the retrieval-augmented many-shot translation pipeline from our AmericasNLP 2026 system to… 9 arXiv — NLP / Computation & Language research 13d ago A Survey of Large Models in Sports arXiv:2608.14377v1 Announce Type: new Abstract: Sports have witnessed growing global enthusiasm in recent years, serving as a vital force for physical health, cultural exchange, social connection, and economic growth. The rapid advancement of large models, particularly… 24 arXiv — NLP / Computation & Language research 13d ago Split the Labor: Separating Evidence Interpretation from Decision Aggregation arXiv:2608.14509v1 Announce Type: cross Abstract: Systems that ask a language model to reach a conclusion from many sources usually concatenate them into one prompt. This conflates two operations with different requirements. Interpreting a source rewards capacity and context.… 15 arXiv — NLP / Computation & Language research 13d ago Seeing is Coding: On the Effectiveness of Vision Language Models in Code Understanding arXiv:2602.01785v3 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved remarkable success in source code understanding, yet as software systems grow in scale, computational efficiency has become a critical bottleneck. Currently, these models rely on a… 23 r/LocalLLaMA community 15d ago Building an open-source control plane for self-hosted vLLM, what would you want in it? Every time I self-host a model I rebuild the same stuff: start the container, set up a route, check why it died overnight, remember to shut the GPU off before it burns money. So I'm building a panel that handles it. Start/stop models, OpenAI-compatible endpoint, health checks… 25 r/LocalLLaMA community 15d ago NInfer day0 support for Qwen3.8 27b: ~200 tok/s generation, with tons of engine improvments Qwen3.8-27B is finally here, and NInfer already has Day-0 support! Weights: https://huggingface.co/neroued/Qwen3.8-27B-NInfer Just update to the latest source and give it a try. On a single RTX 5090, NInfer can still reach around 200 tok/s generation with speculative decoding.… 5 r/MachineLearning community 15d ago Open-source Python library + no-code web dashboard for evaluating oncology AI models at clinical decision thresholds. [P] Most classification metrics for oncology AI models (AUC, ICC, MAE) measure global agreement. They don't answer the question that actually matters at the point of care: how reliable is this model at the exact cutoff that decides whether a patient gets flagged, biopsied, or… 34 llama.cpp releases dev-tools 16d ago b10426 ggml: force single thread on wasi ( #25686 ) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64… 27 Hugging Face Daily Papers research 16d ago SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models Abstract SKILLER is a reinforcement learning framework that automatically generates tailored skills for small open-source models to reduce inference costs while maintaining high task performance. Generated by thinkingmachines/Inkling-Small Agent skills represent a standardized… 36 arXiv — Machine Learning research 16d ago CoMedBench: A Multi-Source Benchmark of Synthetic Medical Data Fidelity and Downstream Utility arXiv:2608.12805v1 Announce Type: new Abstract: Access to clinical data is essential for developing reliable healthcare machine learning systems, but direct use of electronic health records is constrained by privacy regulation, institutional review, data-use agreements, and the… 38 arXiv — Machine Learning research 16d ago Doubly Robust Estimation of Causal Effect on CVR with Targeted Regularization arXiv:2608.13461v1 Announce Type: new Abstract: Post-click conversion rate (CVR) is a key metric in various scenarios including e-commerce and advertising, reflecting the efficiency and user experience in the second stage of the conversion process. Estimating the causal effect… 28 arXiv — NLP / Computation & Language research 16d ago AnchorSIPS: A Synthetic Dataset and Evaluation Resource for Evidence-Supported Psychosis-Risk Symptom Measurement arXiv:2608.12329v1 Announce Type: new Abstract: Progress on AI for psychosis-risk assessment is limited by a data-access bottleneck. Real clinical interviews are difficult to share because of privacy, governance, and consent constraints. We present AnchorSIPS, a synthetic… 32 arXiv — NLP / Computation & Language research 16d ago When Explanations Betray Backdoors: Black-Box Auditing for Language Model Classifiers arXiv:2608.12623v1 Announce Type: new Abstract: Language model classifiers with explanations are used for moderation, routing, topic triage, and low-resource annotation. We study black-box auditing when the defender has only clean calibration data without trigger information but… 38 arXiv — NLP / Computation & Language research 16d ago BavGround: A Benchmark for Regional Cultural Grounding and Dialect Competence in Bavarian arXiv:2608.12894v1 Announce Type: new Abstract: Cultural evaluation of large language models (LLMs) often focuses on high-resource standard languages, leaving regional culture and dialect communities underrepresented. We introduce BavGround, a benchmark for evaluating Bavarian… 35 arXiv — NLP / Computation & Language research 16d ago DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data arXiv:2608.13517v1 Announce Type: new Abstract: Current large language model development relies on massive, often non-permissible datasets, creating a high barrier for researchers committed to open-source and ethically sourced data. We introduce Mimir v1, a 1-billion-parameter… 32 arXiv — NLP / Computation & Language research 16d ago MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination arXiv:2608.13476v1 Announce Type: cross Abstract: We present Multi-Agent Reasoning and Coordination (MARC), an open-source framework that replaces monolithic LLM prompting with deterministic multi-agent orchestration for clinical reasoning. MARC coordinates role-specialized… 17 TechCrunch — AI news-outlet 16d ago Writer introduces new AI model and upgraded harness to contain token costs Built as a post-training variation on Z.ai's open source model GLM-5.2, Writer says the new system should provide deployment-ready capabilities at a much lower price. 28 r/MachineLearning community 16d ago worldproof: diagnosing where world-model predictions break and a measurement of when pixel metrics stop being able to rank models at all [P] I've been building an open-source tool for diagnosing world models, the kind that predict future frames from a starting context and a sequence of actions. It compares a rollout against ground truth and against physical invariants, then tells you where and why the prediction… 30 r/MachineLearning community 16d ago Neurips 2026: Modified date on reviews [D] Reviews modified dates are public, and some are recent. I’m a bit confused as to how to interpret this. In other conferences, reviewers were required to provide a final justification, which would practically force them to modify their reviews during the AC discussion phase lest… 37 r/LocalLLaMA community 16d ago Deepseek Harness is Up! DeepSeek Harness (dsh) is an open-source agent harness developed by DeepSeek AI. It uses an architecture where everything is a plugin, and is powered by Cordis, whose design is described in A Programming Paradigm for Spatiotemporal Composability. DeepSeek Harness is currently in… 13 Page 4 of 10 · 500 articles ← Newer Older →