arXiv — Machine Learning
500 articles archived · Visit source ↗ · RSS
-
-
arXiv — Machine Learning research 4d ago
Data Predictability Shapes Weibull Weight-Scale Growth in Transformer Training
arXiv:2608.23573v1 Announce Type: new Abstract: A trained transformer's weight magnitudes can be summarized by a two-parameter Weibull distribution whose shape $k \approx 1.2$ is stable across layers and models, so the scale $\lambda$ carries most training-induced movement. What…
10 -
arXiv — Machine Learning research 4d ago
From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers
arXiv:2608.23660v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to provide prior causal knowledge for structural causal discovery, yet whether their direct-edge judgments and confidence can be trusted remains unclear. We systematically evaluate…
25 -
arXiv — Machine Learning research 4d ago
Renormalization Group Flow Matching for Scalable Local Generative Modeling
arXiv:2608.23696v1 Announce Type: new Abstract: Despite their remarkable success in modeling complex data, generative models face a fundamental tradeoff. Global approaches can capture full structural coherence but suffer from high computational costs, while local models are…
26 -
arXiv — Machine Learning research 4d ago
Response Renormalization for Critical Deep Equilibrium Models
arXiv:2608.23725v1 Announce Type: new Abstract: Deep Equilibrium Models (DEQs) compute predictions from a hidden representation unchanged by the model update. Training through this equilibrium uses implicit differentiation and requires solving an adjoint system built from the…
8 -
arXiv — Machine Learning research 4d ago
Calibration-Preserving Pruning: Compression as a Reliability Contract
arXiv:2608.23744v1 Announce Type: new Abstract: Split conformal prediction, not the pruning rule, supplies finite-sample marginal coverage once a pruned model is fixed independently of the conformal calibration split. We study the separate efficiency problem: can pruning…
31 -
arXiv — Machine Learning research 4d ago
Tight Majorizations and Convergence Rates of Nuclear Norm Minimization IRLS
arXiv:2608.23765v1 Announce Type: new Abstract: Iteratively reweighted least squares (IRLS) methods constitute a natural approach to nuclear norm minimization, but their convergence rates and the role of the weight operator have remained poorly understood. This paper establishes…
7 -
arXiv — Machine Learning research 4d ago
Disentangled Skill Representations for Predictive Human Modeling
arXiv:2608.23776v1 Announce Type: new Abstract: Understanding human skill is important for AI systems that collaborate with, coach, or assist people. Unlike typical latent variable estimation problems which rely on single observations, skill is a persistent, compositional, and…
21 -
arXiv — Machine Learning research 4d ago
GAP-Prompt: Gated Adaptive Prompting for Efficient Continual Learning
arXiv:2608.23782v1 Announce Type: new Abstract: Continual learning faces the persistent challenge of catastrophic forgetting, where sequential task updates degrade previously acquired knowledge. While prompt-based methods integrated with pre-trained models offer a compelling…
15 -
arXiv — Machine Learning research 4d ago
Mixture of Channel Experts: Static Sparse Supports with Input-Adaptive Mixing for Pointwise Projections
arXiv:2608.23794v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) scales language models by routing each input through a small set of independently parameterized experts. We show that copying this design into convolutional networks fails for a structural reason: parallel…
15 -
arXiv — Machine Learning research 4d ago
A Theory of Speciation in Generative Diffusion Models on Compact Riemannian Manifolds
arXiv:2608.23798v1 Announce Type: new Abstract: Speciation in generative diffusion models denotes the emergence of distinct stable branches during denoising, through which initially undifferentiated trajectories progressively commit to different data classes. In this work we…
6 -
arXiv — Machine Learning research 4d ago
Discovering Cross-Language Reasoning Invariance in LLMs with Geometry-Invariant Sparse Autoencoders
arXiv:2608.23809v1 Announce Type: new Abstract: Multilingual language models can solve the same mathematical problem in different languages, but it remains unclear whether they rely on shared features or on language-specific computations that only produce similar outputs. We…
31 -
arXiv — Machine Learning research 4d ago
Learning to Grade Efficiently: A Bandit-Driven Prompt-Selection Framework for Low-Cost LLM Essay Scoring
arXiv:2608.23814v1 Announce Type: new Abstract: Large Language Models (LLMs) demonstrate strong capabilities in automated essay scoring (AES), but contemporary approaches typically employ fixed prompt selection, failing to address operational cost concerns and evolving optimal…
26 -
arXiv — Machine Learning research 4d ago
AQLoRA: A Zero-Search Recipe for Fast Quantized LoRA Fine-Tuning
arXiv:2608.23816v1 Announce Type: new Abstract: Quantized fine-tuning (QLoRA) saves memory but not time. It dequantizes every 4-bit weight on the fly, so it trains more slowly than fp16 LoRA. We present AQLoRA (Adaptive-Quantization LoRA), a recipe that buys part of that time…
10 -
arXiv — Machine Learning research 4d ago
Generating Intervention Hypotheses using Explainable Explanations on Graphs: G2I, a Two-Stage Greedy Framework
arXiv:2608.23835v1 Announce Type: new Abstract: Real-world decision-making in public health and social science can greatly benefit from predictive models, yet translating predictions into effective interventions requires explaining the model behavior. While Graph Neural Networks…
36 -
arXiv — Machine Learning research 4d ago
PuzzleKV: Page-Wise Low-Rank Decomposition for KV Cache Compression
arXiv:2608.23843v1 Announce Type: new Abstract: Long-context inference in large language models (LLMs) is increasingly limited by the memory required for the key-value (KV) cache. KV cache compression addresses this problem by reducing the storage cost of previous tokens. Among…
9 -
arXiv — Machine Learning research 4d ago
FlowNeg: GFlowNet-Guided Diverse Hard Negative Sampling for Knowledge Graph Embedding
arXiv:2608.23849v1 Announce Type: new Abstract: Negative sampling determines whether a knowledge graph embedding (KGE) model learns from informative counterexamples or wastes updates on implausible corruptions. Uniform negatives are diverse but easy, whereas hard-negative miners…
15 -
arXiv — Machine Learning research 4d ago
UHI-Bench: Benchmarking Dual-Source Urban Heat Island Modeling Across Cities in Diverse Climate Regimes
arXiv:2608.23857v1 Announce Type: new Abstract: Urban heat islands (UHIs) are intensifying under climate change, exacerbating thermal exposure risks. Their two primary observations, land surface temperature UHI (LST-UHI) and near-surface air temperature UHI (AirT-UHI), capture…
21 -
arXiv — Machine Learning research 4d ago
Revelation Control
arXiv:2608.23860v1 Announce Type: new Abstract: Revelation Control is the problem of choosing priced interventions that reveal hidden state only insofar as the revealed distinctions can change a consequential decision, while accounting separately for any useful progress created…
30 -
arXiv — Machine Learning research 4d ago
Every Layer Counts: An Exponential $L_2$ Depth Hierarchy for ReLU Networks
arXiv:2608.23877v1 Announce Type: new Abstract: We prove a depth hierarchy for ReLU neural networks in which every additional ReLU layer can save exponentially many neurons. For every $\ell\geq 3$, a globally $[0,1]$-valued, $1$-Lipschitz function is realized by a depth-$\ell$…
15 -
arXiv — Machine Learning research 4d ago
Partial Optimal Transport on the Circle for All Transported Masses in O(N log N)
arXiv:2608.23910v1 Announce Type: new Abstract: Partial optimal transport compares two measures while leaving part of the mass unmatched, which is what makes it robust to outliers, occlusion, and clutter. The quantity of interest is usually the whole profile - the optimal cost…
29 -
arXiv — Machine Learning research 4d ago
The Loss Floor of Denoising Score Matching: Fisher Geometry from Schr\"odinger Bridges
arXiv:2608.23916v1 Announce Type: new Abstract: Denoising score matching trains diffusion models by regressing onto a conditional score, although generation ultimately requires the marginal score. The two objectives share the same population minimizer, but the conditional target…
6 -
arXiv — Machine Learning research 4d ago
GATNextHop: A GAT for Shortest Path Routing with Cross-Topology Generalization
arXiv:2608.23917v1 Announce Type: new Abstract: Common shortest-path algorithms, such as Dijkstra's (SPF), that OSPF uses, provide exact routing solutions but must be recomputed for each network topology, limiting scalability in dynamic or large-scale networks. This paper…
7 -
arXiv — Machine Learning research 4d ago
MnemoDyn: Learning Resting State Dynamics from 40K FMRI sequences
arXiv:2608.23936v1 Announce Type: new Abstract: We present a dynamical-systems based model for resting-state functional magnetic resonance imaging (rs-fMRI), trained on a dataset of roughly 40K rs-fMRI sequences covering a wide variety of public and available-by-permission…
25 -
arXiv — Machine Learning research 4d ago
CoDrift: Compositional Drifting for Offline Reinforcement Learning
arXiv:2608.23939v1 Announce Type: new Abstract: Offline reinforcement learning is intrinsically multi-objective: a policy must remain compatible with the behavioral support of a fixed dataset while preferentially selecting high-value actions. We recast these objectives in a…
18 -
-
arXiv — Machine Learning research 4d ago
Revenge of Monosemanticity: Specialized Neurons Improve Data Efficiency in MLPs
arXiv:2608.24007v1 Announce Type: new Abstract: Understanding how neural networks learn and organize features is central to understanding their behavior. Much existing theory of feature learning has focused on the emergence of a global low-dimensional predictive geometry. We…
10 -
arXiv — Machine Learning research 4d ago
ChorusTIC: Training-Free Multivariate Time Series Classification via Chorus In-Context Learning
arXiv:2608.24033v1 Announce Type: new Abstract: Time series classification underpins applications in healthcare, sensing, and industrial monitoring. Although time series foundation models support forecasting and transferable representation learning, classification still…
28 -
arXiv — Machine Learning research 4d ago
PinSieve: Production Selective VLM Serving and a Governed Memory Flywheel for Enterprise Content-Quality Triage
arXiv:2608.24040v1 Announce Type: new Abstract: Enterprise AI agents in production often need to be bounded, stateful, observable, and governable rather than fully autonomous. We present PinSieve, a production case study in a large-scale content-quality pipeline. Its deployed…
21 -
arXiv — Machine Learning research 4d ago
XP-JEPA: Cross-Predictive Physics Grounding for Forecastable Latent Dynamics
arXiv:2608.24044v1 Announce Type: new Abstract: Latent world models plan by predicting how candidate actions transform learned representations. In self-predictive models, however, the encoder and predictor are optimized jointly and can co-adapt to latent transitions that are…
7 -
arXiv — Machine Learning research 4d ago
Physics-Integrated Operator Learning via Gaussian Splatting Representations
arXiv:2608.24049v1 Announce Type: new Abstract: Neural operators provide efficient surrogates for spatiotemporal PDE systems, but purely data-driven formulations often accumulate substantial errors during long-horizon autoregressive prediction and may fail to exploit available…
38 -
arXiv — Machine Learning research 4d ago
ALPHABET: A Laplace-Pole History Aggregator with Banked Exponential Transport
arXiv:2608.24051v1 Announce Type: new Abstract: Can a sequence model remain competitive with only a few thousand parameters and an explicitly auditable prediction interface? We introduce ALPHABET, a compact linear-time model that compresses temporal history into stable complex…
18 -
-
arXiv — Machine Learning research 4d ago
Mechanistic Circuit Identification for Controllable Data Generation
arXiv:2608.24065v1 Announce Type: new Abstract: While recent advances in data synthesis aim to curate high-quality datasets, most generation pipelines still rely on heuristic prompt-based control. This black-box paradigm provides limited insight into how individual samples…
15 -
-
arXiv — Machine Learning research 4d ago
Knowing When to Ask for Help: Bayesian Self-Escalation in Hierarchical LLM Agents
arXiv:2608.24087v1 Announce Type: new Abstract: Current LLM agent systems decide delegation before reasoning begins (a router picks a model) or after a response is complete (a verifier scores it and may retry). We study a third regime: an agent that recognises, during its own…
30 -
arXiv — Machine Learning research 4d ago
The Sharp Tail of Uniform Stability
arXiv:2608.24098v1 Announce Type: new Abstract: Uniform stability controls how much one training example can change the loss at any test point. A new logarithmic-free upper bound shows that a $\gamma$-uniformly stable algorithm with loss in $[0,L]$ has generalization gap at most…
6 -
arXiv — Machine Learning research 4d ago
Structured Frequency-Domain Evidence for LLM-Based Time-Series Anomaly Detection
arXiv:2608.24113v1 Announce Type: new Abstract: Time-series anomalies can appear not only as pointwise deviations but also as changes in recurring temporal structure, such as shifted periodicity or localized oscillatory fluctuations. However, existing LLM-based time-series…
23 -
arXiv — Machine Learning research 4d ago
A mesh-free multiresolution deep energy method with phase-field modeling of brittle fracture
arXiv:2608.24126v1 Announce Type: new Abstract: Phase-field modeling of brittle fracture removes the need to track cracks explicitly by recasting their evolution as the minimization of an energy functional. In return it requires a discretization dense enough to resolve a…
25 -
-
arXiv — Machine Learning research 4d ago
Steering Recurrent Reasoners at Inference Time with Readout Feedback
arXiv:2608.24136v1 Announce Type: new Abstract: Recurrent models, which repeatedly update latent states with shared computation blocks, have emerged as powerful architectures for solving complex reasoning tasks. Existing inference-time methods scale computation by running more…
12 -
arXiv — Machine Learning research 4d ago
Robust Data-Collection Policy Learning for Low-Variance Online Policy Evaluation
arXiv:2608.24146v1 Announce Type: new Abstract: In reinforcement learning policy evaluation, classic on-policy methods often suffer from high variance when estimating policy performance. To mitigate this issue, behavior policy search has been proposed to learn data-collecting…
12 -
arXiv — Machine Learning research 4d ago
From Relaxed Indexability to Exact Indexability: A $t$-Step Approach for Partially Observable Restless Bandits
arXiv:2608.24167v1 Announce Type: new Abstract: Whittle index policies offer a scalable method for restless multi-armed bandits, but under partial observability even determining the indifference subsidy at a single belief requires solving an infinite-horizon belief-state problem…
37 -
arXiv — Machine Learning research 4d ago
PRQ-KMeans: Projection Residual Quantization for Semantic ID Tokenization
arXiv:2608.24207v1 Announce Type: new Abstract: Semantic identifiers (SIDs) represent entities as hierarchical token sequences for generative retrieval and recommendation. Residual-quantization tokenizers construct these sequences by selecting a codeword at each level and…
17 -
arXiv — Machine Learning research 4d ago
A Data-dependent Early Stopping Rule using Rademacher Complexity with L1-norm
arXiv:2608.24210v1 Announce Type: new Abstract: Training neural networks requires balancing the trade-off between fitting the training data and achieving robust performance on unseen inputs. This ability, commonly referred to as generalizability, is determined by the gap between…
22 -
arXiv — Machine Learning research 4d ago
Contrastive Branch Policy Optimization
arXiv:2608.24300v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) enables language models to learn multi-turn interaction with external tools, yet its sparse outcome rewards provide no signal for identifying which intermediate decisions are…
27 -
arXiv — Machine Learning research 4d ago
Causal Analysis for Time Series Foundation Models
arXiv:2608.24303v1 Announce Type: new Abstract: Transitioning from bespoke time series models towards time series foundation models changes the relationship of model and application from one-to-one to one-to-many. This shift introduces concentration risk as many, potentially…
18 -
arXiv — Machine Learning research 4d ago
A Structural FHMM for Interpretable Disease Trajectories in T2DM
arXiv:2608.24328v1 Announce Type: new Abstract: In this work, we propose a structural variant of the Factorial Hidden Markov Model (FHMM) for the analysis of disease trajectories in patients with Type 2 diabetes mellitus (T2DM). The model represents a patient's latent health…
38 -
arXiv — Machine Learning research 4d ago
When Does Self-Supervised Pretraining Help Tabular Models? A Study of Label Scarcity and Missing Data
arXiv:2608.24381v1 Announce Type: new Abstract: Self-supervised learning (SSL) has emerged as a promising approach for tabular data, yet its efficacy under extreme label scarcity and test-time missingness remains under-explored. In this paper, we evaluate a mask-and-recover SSL…
24 -
arXiv — Machine Learning research 4d ago
Equivariant Covariance Tensors: Guaranteed SPD Uncertainty for Tensor-Valued Geometric Learning
arXiv:2608.24386v1 Announce Type: new Abstract: Tensor-valued prediction is fundamental to geometric deep learning, yet uncertainty quantification (UQ) for such outputs remains an open challenge. While E(3)-equivariant neural networks excel at point estimates, they lack rigorous…
36