arXiv — Machine Learning
500 articles archived · Visit source ↗ · RSS
-
-
-
-
arXiv — Machine Learning research 6d ago
BF1: A Causal Dyadic Sparse-Attention Retrofit for Efficient Long-Context Transformers
arXiv:2608.20427v1 Announce Type: new Abstract: Dense causal attention remains expensive at long context even when implemented with highly optimized exact kernels. We study BF1, a deterministic block-aligned dyadic sparse-attention route that combines a small exact local…
16 -
arXiv — Machine Learning research 6d ago
Approximate Homomorphisms and Convergent Representations in Transducers
arXiv:2608.20428v1 Announce Type: new Abstract: We study the stability of minimal representations of controlled stochastic processes (in particular, transducers) under perturbations. This question is motivated by recent experiments finding predictive-state structure in the…
11 -
arXiv — Machine Learning research 6d ago
Wrong-Physics Backdoors in Neural PDE Operators
arXiv:2608.20439v1 Announce Type: new Abstract: Neural PDE operators are increasingly trained on reusable solver archives, yet validation often relies on clean prediction error and parameter-agnostic plausibility checks. We introduce cross-parameter relinking, a data-poisoning…
22 -
arXiv — Machine Learning research 6d ago
Decision Tree and K-Means Analysis of Raman Spectra for Edible Oils: A Physics-Informed AI Approach
arXiv:2608.20440v1 Announce Type: new Abstract: Authentication of edible oils in processed foods is important for food quality, fraud prevention, and regulatory compliance. This study establishes an integrated Raman spectroscopy and machine-learning framework that links…
8 -
arXiv — Machine Learning research 6d ago
Shared Physics Responses Recover Hidden Rankings in Neural Operator Libraries
arXiv:2608.20441v1 Announce Type: new Abstract: Selecting the optimal neural-operator prediction during deployment is challenging when high-fidelity reference solutions are unavailable. We demonstrate that under a squared Hilbert-space loss, ranking a finite model library…
7 -
arXiv — Machine Learning research 6d ago
Stored in Optimizer State, Valued by Later Training: A Causal Account of Subliminal Trait Transfer
arXiv:2608.20442v1 Announce Type: new Abstract: Subliminal trait transfer allows a student model to acquire behavioral dispositions from teacher-generated data in which the trait is not semantically expressed. Recent work explains how such signals enter gradients, but not how…
14 -
arXiv — Machine Learning research 6d ago
Amortized Bandwidth Learning for Kernel Density Estimation under Logarithmic Score
arXiv:2608.20445v1 Announce Type: new Abstract: Kernel density estimation converts finite samples into probability densities, but its performance depends critically on bandwidth selection. Classical selectors prescribe the sample-to-bandwidth rule analytically or asymptotically,…
10 -
arXiv — Machine Learning research 6d ago
Mutual information and sensitivity analysis for feature selection in customer targeting: a comparative study
arXiv:2608.20447v1 Announce Type: new Abstract: Feature selection is a highly relevant task in a data-driven knowledge discovery project. Several techniques have been developed aiming at finding the features that influence most an outcome to predict, including mutual information…
21 -
arXiv — Machine Learning research 6d ago
When Clean Data Hurts: Learning with Monotone Corruptions Beyond Binary Classification
arXiv:2608.20480v1 Announce Type: new Abstract: Optimal learners are tailored to exploit the i.i.d.\ data assumption underlying the classic PAC model. What if an i.i.d.\ training sample were corrupted with correctly labeled examples drawn from an otherwise unrelated, even…
4 -
arXiv — Machine Learning research 6d ago
Metag: A dataset to build agentic meta-reviewing capabilities
arXiv:2608.20488v1 Announce Type: new Abstract: AI tools increasingly support tasks across the scientific research cycle, from experiment design and manuscript preparation to peer review. At the same time, the continuing growth in conference submissions has increased the burden…
7 -
arXiv — Machine Learning research 6d ago
Bern2Edge: A Neurosymbolic Compiler for Edge Deployment via Bernstein Polynomial Networks
arXiv:2608.20497v1 Announce Type: new Abstract: Deploying high-accuracy neural networks on resource-constrained edge devices remains challenging, as existing approaches treat training, compression, and hardware synthesis as separate stages, leaving a gap between software-trained…
5 -
arXiv — Machine Learning research 6d ago
When Graph-JEPA Learns the Wrong Thing: Diagnosing and Repairing Category-Conditional Collapse
arXiv:2608.20516v1 Announce Type: new Abstract: Joint-embedding predictive architectures are selected almost universally by linear probing and effective rank. We report a case where both read healthily while the representation carries zero usable instance information. We repair…
37 -
arXiv — Machine Learning research 6d ago
Learning Exact NVIDIA SASS Encoders with $\mathbb{F}_2$ Linear Algebra
arXiv:2608.20532v1 Announce Type: new Abstract: NVIDIA provides a SASS disassembler but no public SASS assembler for recent data-center GPUs, limiting controlled machine-code rewriting. We present F2Asm, which learns exact 128-bit SASS encoders from paired disassembly and…
25 -
arXiv — Machine Learning research 6d ago
AgentDecarbonizer: Carbon-Aware Execution for AI Agents
arXiv:2608.20566v1 Announce Type: new Abstract: AI agents extend large language models from single prompt-response interactions to long-running, goaldirected workflows that issue many model calls, invoke tools, and interact with external environments. These workflows enable…
32 -
arXiv — Machine Learning research 6d ago
Faults That Fortify: CNN Adversarial Robustness via GPU Undervolting
arXiv:2608.20572v1 Announce Type: new Abstract: Convolutional Neural Networks (CNNs) face a dual challenge: vulnerability to adversarial attacks and prohibitive training cost. Adversarial training is effective but expensive, a burden that grows as learning shifts to the…
15 -
arXiv — Machine Learning research 6d ago
Provable Edge-of-Stability for Adam on a One-Dimensional Quadratic
arXiv:2608.20638v1 Announce Type: new Abstract: The edge-of-stability (EoS) phenomenon of Adam has been widely observed, while its underlying dynamical mechanism is not yet fully understood. We study uncorrected Adam on a one-dimensional quadratic, a clean setting where constant…
20 -
-
arXiv — Machine Learning research 6d ago
RiskTraf: Risk-Extrapolated Residual Learning for Multi-Variate Traffic Flow Prediction
arXiv:2608.20656v1 Announce Type: new Abstract: Traffic sensors commonly record flow, speed, and occupancy, but standard traffic flow forecasting benchmarks and models rarely exploit all three raw measurements reliably. Although speed and occupancy provide sensor-native…
28 -
arXiv — Machine Learning research 6d ago
C-Score: Beyond Accuracy for Robustness Assessment in Semi-Supervised Learning under Open-World Unlabeled Contamination
arXiv:2608.20667v1 Announce Type: new Abstract: Pseudo-label-based semi-supervised learning has achieved strong performance due to its simplicity and scalability. However, it is typically developed under a closed-world assumption that unlabeled data are drawn from the same…
22 -
arXiv — Machine Learning research 6d ago
Lightweight Adaptive ReduNet via Hyperspherical Manifold Learning
arXiv:2608.20668v1 Announce Type: new Abstract: In recent years, a white-box neural network called ReduNet has been proposed, which employs the maximal coding rate reduction (MCR$^2$) principle to transform raw data into low-dimensional discriminative features via a forward…
36 -
-
arXiv — Machine Learning research 6d ago
Geometric Regularization for Long-Tailed Semi-Supervised Learning via Gaussian Feature Bridges
arXiv:2608.20710v1 Announce Type: new Abstract: Real-world semi-supervised learning (SSL) often encounters significant challenges with long-tailed label distributions and noisy pseudo-labels, which hinder generalization and amplify confirmation bias. In this work, we introduce a…
7 -
arXiv — Machine Learning research 6d ago
Hidden Axis of Uncertainty: Latent-Posterior Alignment in Graph Neural Networks with Bayesian Output Layers
arXiv:2608.20758v1 Announce Type: new Abstract: Bayesian Neural Networks (BNNs) with Bayesian output layers provide a principled and tractable framework for quantifying predictive uncertainty, yet the mechanisms shaping that uncertainty remain unclear. While conventional theory…
31 -
-
arXiv — Machine Learning research 6d ago
Resolution-Consistent Greedy Neural Approximation on Infinite-Dimensional Spaces
arXiv:2608.20812v1 Announce Type: new Abstract: We develop constructive approximation and learning guarantees for shallow neural models with infinite-dimensional inputs observed through finitely many coordinates. The analysis is based on a parameter-normalized neural dictionary…
37 -
arXiv — Machine Learning research 6d ago
Scaling Muon for Diffusion Transformers
arXiv:2608.20818v1 Announce Type: new Abstract: The matrix-aware optimizer Muon improves large model training by balancing updates across singular directions, yet its scaling behavior and end-to-end efficiency on large Diffusion Transformers (DiTs) remain unclear. We first…
7 -
-
arXiv — Machine Learning research 6d ago
Decoupling Policy Extraction for Offline Reinforcement Learning
arXiv:2608.20909v1 Announce Type: new Abstract: Offline RL methods commonly jointly train the actor and critic, where the critic is used to guide the actor toward higher-value actions. This coupled learning process is well motivated in online RL, where an improved actor collects…
14 -
arXiv — Machine Learning research 6d ago
Training, learning and inference: unified dynamics of neural systems
arXiv:2608.20965v1 Announce Type: new Abstract: We define an atomic generation fact f=(u,tau,omega,z;rho), recording the origin, realized transformation, concrete occurrence, generated result and relation role. Compiled into a Generation-Fact Graph (GFG), these facts provide an…
26 -
arXiv — Machine Learning research 6d ago
A Critical Audit of Spatiotemporal Forecasting Benchmark Datasets and Baselines
arXiv:2608.20980v1 Announce Type: new Abstract: Graph neural networks (GNNs) are routinely employed for short-range forecasting on multivariate time series with a spatial graph structure. Despite the availability of many alternative datasets, method innovations within this…
14 -
arXiv — Machine Learning research 6d ago
Jacobian-guided Noise Injection for Quantization Robustness in Large Language Models
arXiv:2608.20988v1 Announce Type: new Abstract: Quantization of Large Language Models (LLMs) is often hindered by the sensitivity of the self-attention mechanism to discretization errors. We identify the softmax operator as a bottleneck for quantization stability due to its…
21 -
arXiv — Machine Learning research 6d ago
Trojaning the Alignment: Stealthy Backdoor Attacks against Graph Foundation Models
arXiv:2608.20991v1 Announce Type: new Abstract: Graph Foundation Models (GFMs) on text-attributed graphs (TAGs) align graph representations with language semantics to support transferable graph learning. Despite these advantages, the backdoor vulnerability of GFMs on TAGs…
20 -
arXiv — Machine Learning research 6d ago
Free-Probability Kernels for Zero-Rollout Hyperparameter Selection in Reservoir Computing
arXiv:2608.20998v1 Announce Type: new Abstract: Reservoir computing (RC) couples a fixed recurrent dynamical system with a trained lightweight readout, but this efficiency is partly lost during hyperparameter selection: the recurrent gain, input scale, and leakage rate determine…
9 -
arXiv — Machine Learning research 6d ago
RODE: A Radial-Orthogonal Decoupled Engine for Optimization
arXiv:2608.21024v1 Announce Type: new Abstract: Modern neural network training increasingly uses matrix-aware optimizers, yet their conditioned matrix step is typically added directly to the weight, jointly changing its norm and direction. This interaction matters because the…
12 -
arXiv — Machine Learning research 6d ago
Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment
arXiv:2608.21057v1 Announce Type: new Abstract: Agentic large language model (LLM) systems are reshaping scientific workflows in chemistry and drug discovery, but evaluating their open-ended, tool-augmented outputs remains a fundamental bottleneck. Reference-based metrics such…
12 -
arXiv — Machine Learning research 6d ago
TracingFlow: A Simulation-Free Trajectory Inference Framework Based on Second-Order Dynamics
arXiv:2608.21070v1 Announce Type: new Abstract: Inferring continuous system evolution from sparse temporal snapshots is a key challenge in generative modeling and single-cell omics. While Optimal Transport (OT) is popular, existing frameworks are largely restricted to…
17 -
arXiv — Machine Learning research 6d ago
Causal Modeling of Adverse Pregnancy Outcomes via Adaptive LLM Proposals
arXiv:2608.21079v1 Announce Type: new Abstract: Adverse Pregnancy Outcomes (APOs) such as preterm birth and gestational diabetes can have long-term consequences for both the mother and child, yet an understanding of their causes remains elusive. Causal discovery in this domain…
12 -
arXiv — Machine Learning research 6d ago
FlatLand: Personalized Graph Federated Learning via Tailored Lorentz Space
arXiv:2608.21096v1 Announce Type: new Abstract: Federated learning enables privacy-preserving collaborative training, but highly heterogeneous client data remain challenging, especially in graph federated learning where clients possess structurally diverse graphs. Existing…
11 -
arXiv — Machine Learning research 6d ago
BackDFL: A Unified Benchmark For Backdoor Attacks and Defenses In Decentralized Federated Learning
arXiv:2608.21137v1 Announce Type: new Abstract: Decentralized Federated Learning (DFL) promises trust-free collaborative learning by replacing the centralized parameter server with peer-to-peer model exchange. However, this architectural shift fundamentally reshapes the threat…
7 -
arXiv — Machine Learning research 6d ago
COEC: Calibrated Orthogonal-Equivalence Compensation for Structured Pruning of Large Language Models
arXiv:2608.21142v1 Announce Type: new Abstract: Structured pruning reduces the size and inference cost of large language models (LLMs) by removing weight columns, but the resulting output error can degrade accuracy. Existing training-free compensation methods use an additive…
16 -
arXiv — Machine Learning research 6d ago
Capturing Cardiac Cyclicity through Phase-Equivariant Self-Supervised Learning
arXiv:2608.21147v1 Announce Type: new Abstract: The cyclic structure of physiological processes offers a natural prior for self-supervised representation learning, and the cardiac cycle provides a particularly well-defined setting in which to exploit it. We derive a…
24 -
arXiv — Machine Learning research 6d ago
Thermo-FL: Thermal-Aware Robust Federated Fine-Tuning of Large Language Models for Edge AI
arXiv:2608.21172v1 Announce Type: new Abstract: Federated fine-tuning enables large language models to adapt on edge devices without centralizing private data, but practical deployments must address hardware instability and adversarial update corruption together. Thermally…
4 -
arXiv — Machine Learning research 6d ago
A Neurosymbolic Approach for Constructing Planning Domain Models from Clinical Narratives
arXiv:2608.21186v1 Announce Type: new Abstract: Surgical procedures such as laparoscopic appendectomy are complex, high-stakes processes, yet formalizing their workflows for decision support remains a significant challenge. Inducing probabilistic planning domain models in this…
13 -
arXiv — Machine Learning research 6d ago
Tydra: An Efficient Hybrid Model for Tabular Data
arXiv:2608.21199v1 Announce Type: new Abstract: Transformer-based tabular foundation models such as TabPFN achieve strong predictive performance but incur quadratic computational cost with context length. On the other hand, subquadratic SSM-based alternatives such as Hydra trade…
27 -
arXiv — Machine Learning research 6d ago
Curriculum-Aware Interpolate-then-Refine: Learned Physiological Time-Series Imputation under Realistic Missingness
arXiv:2608.21207v1 Announce Type: new Abstract: Imputing physiological time series (arterial blood pressure, blood glucose, etc.) is essential for addressing the missingness that pervades clinical data. Yet modern imputation methods perform poorly in this domain: a recent…
23 -
arXiv — Machine Learning research 6d ago
TRACE-C: Rank-Calibrated Relational Anomaly Detection for Multi-Stream Operational Telemetry
arXiv:2608.21251v1 Announce Type: new Abstract: Operational telemetry can be jointly anomalous while every individual stream stays inside its familiar range. TRACE-C is an auditable strictly-prior rank-calibrated detector for aligned multi-stream telemetry: same-regime rolling…
35 -
arXiv — Machine Learning research 6d ago
ConceptTS: LLM-Guided Concept Bottlenecks for Interpretable Multivariate Time-Series Forecasting
arXiv:2608.21277v1 Announce Type: new Abstract: State-of-the-art multivariate time-series forecasters can model complex temporal and cross-variable dependencies, yet their opaque representations provide limited insight into why a particular forecast is produced. This lack of…
11