Hugging Face Daily Papers
500 articles archived · Visit source ↗ · RSS
-
Hugging Face Daily Papers research 10d ago
SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents
Abstract SkillGate fixes selector credit starvation in agent skill selection by separating outcome credit for execution tokens from local advantage for skill-naming tokens, improving success rates and reducing misleading skill exposure. Generated by…
13 -
Hugging Face Daily Papers research 10d ago
The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning
Abstract AdaPop adapts gradient pressure by fact popularity and automates forget-retain balance to reduce leakage of unlearned content. Generated by thinkingmachines/Inkling-Small Popular facts are memorised more deeply during pretraining and resist removal longer than rare…
12 -
Hugging Face Daily Papers research 10d ago
FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents
Abstract FM-Bench evaluates long-horizon decision-making of LLM agents managing a football club over 20 years, revealing that managerial behavior rather than scale or token spend drives performance. Generated by thinkingmachines/Inkling-Small Language model agents now execute…
31 -
Hugging Face Daily Papers research 10d ago
Looped Language Models Improve Compositional Tool Calling
Abstract Looped language models improve compositional, multi-step tool use through recurrent computation, with adaptive inference balancing accuracy and compute cost. Generated by thinkingmachines/Inkling-Small Looped language models have shown promising results on reasoning…
14 -
Hugging Face Daily Papers research 10d ago
Temporal Multi-Signal Fusion for Token-Level Hallucination Detection
Abstract Hallucination is detected as temporally extended spans via sequence labeling over fused external features, achieving robust cross-model performance without internal model access. Generated by thinkingmachines/Inkling-Small Token-level hallucination detectors score each…
33 -
Hugging Face Daily Papers research 10d ago
SPADE: Self-Play in Adaptive Synthetic Executable Environments
Abstract SPADE is a self-play reinforcement learning framework where a language model designs adaptive executable training environments and learns to solve them, improving reasoning and tool-use performance through regret-based environment targeting. Generated by…
7 -
Hugging Face Daily Papers research 10d ago
Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL
Abstract Co-RL enables unsupervised reasoning via cooperative multi-agent reinforcement learning with peer-derived rewards, improving performance across text and vision tasks without ground-truth labels. Generated by thinkingmachines/Inkling-Small Reinforcement learning (RL) has…
36 -
Hugging Face Daily Papers research 10d ago
Scaling Creative Writing Beyond Story-Centric Data with Attribute-Guided Genre Expansion
Abstract A framework that separates thematic seeds from genre-form controls generates diverse, high-quality creative writing data across 13 genres and improves LLM creative writing performance. Generated by thinkingmachines/Inkling-Small High-quality creative writing data for…
8 -
Hugging Face Daily Papers research 10d ago
Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence
Abstract Zetta is a closed-loop embodied harness that evolves runtime critics and recovery skills online to govern physical execution at action frequency, achieving high success on robot benchmarks with faster inference and scaling self-exploration. Generated by…
36 -
Hugging Face Daily Papers research 10d ago
Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning
Abstract Action-conditioned objectives improve latent geometry for Euclidean-cost model-predictive control by enhancing decision-metric alignment in world models. Generated by thinkingmachines/Inkling-Small JEPA-style latent world models can use Euclidean distance to a goal…
16 -
Hugging Face Daily Papers research 10d ago
SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation
Abstract SoftVTBench introduces a synchronized visuo-tactile dataset and deformation-aware benchmark for evaluating physical interaction quality during deformable-object manipulation. Generated by thinkingmachines/Inkling-Small Physical interaction quality is central to…
24 -
Hugging Face Daily Papers research 10d ago
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation
Abstract Semantic task completion video generation evaluates whether generated videos achieve intended outcomes with semantic grounding, supported by a curated dataset and vision-language model-based benchmark. Generated by thinkingmachines/Inkling-Small We introduce Semantic…
22 -
Hugging Face Daily Papers research 10d ago
Training Leaves Traces: Centered Residual Signatures for Language Model Lineage Verification
Abstract Compatible open-weight language model checkpoints share detectable weight-space ancestry signals that distinguish true lineage from independent or distilled models without requiring data. Generated by thinkingmachines/Inkling-Small Open-weight language models are…
30 -
Hugging Face Daily Papers research 10d ago
Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis
Abstract Top-K prompting and plausibility-aware training improve diverse reaction prediction in single-step retrosynthesis, yielding state-of-the-art results on a large verified reaction dataset and motivating ensemble systems. Generated by thinkingmachines/Inkling-Small…
34 -
Hugging Face Daily Papers research 10d ago
SemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code Generation
Abstract SemaPLC is a verification-gated agent harness that validates generated PLC logic through external compilation and live runtime execution, achieving higher verified pass rates than baseline methods. Generated by thinkingmachines/Inkling-Small Programmable logic…
20 -
Hugging Face Daily Papers research 10d ago
FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis
Abstract FACET constructs executable terminal tasks by preserving source intent and grounding instructions, solutions, and verifiers in a shared repaired environment to enable scalable agent training. Generated by thinkingmachines/Inkling-Small Training terminal agents requires…
15 -
Hugging Face Daily Papers research 10d ago
The Problem Is the Problem: Towards Scalable Mathematical Discovery
Abstract A literature-to-review pipeline automates problem discovery and triage to focus expert review on promising mathematical conjectures within a chosen research direction. Generated by thinkingmachines/Inkling-Small AI systems are increasingly capable of contributing to…
34 -
Hugging Face Daily Papers research 10d ago
LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents
Abstract LEGO-RL connects native coding-agent harnesses to scalable policy-gradient training via in-process LLM proxying, sandbox orchestration, and integrated monitoring, improving sparse MoE model performance across multiple harnesses. Generated by…
33 -
Hugging Face Daily Papers research 10d ago
PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX
Abstract PTXBench evaluates large language models on architecture-specific GPU kernel optimization, revealing uneven success and performance gaps that supervised fine-tuning only partially addresses. Generated by thinkingmachines/Inkling-Small We introduce PTXBench, a benchmark…
27 -
Hugging Face Daily Papers research 11d ago
CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation
Abstract CardioState-JEPA learns a unified cardiac representation across ECG, PPG, and PCG by predicting masked latent physiological states with cross-modal delay alignment, improving downstream classification across all three modalities. Generated by…
6 -
Hugging Face Daily Papers research 11d ago
V-RAE: Rethinking Video Latent Spaces for Generation
Abstract V-RAE constructs semantically organized video latents from frozen vision representations to improve generation quality, convergence speed, and predictive modeling. Generated by thinkingmachines/Inkling-Small Latent video generation relies on autoencoders to define a…
32 -
Hugging Face Daily Papers research 11d ago
DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization
Abstract DiSCO is a black-box, zero-shot prompt-level defense that uses distribution-guided suffix expansion and contrastive scoring to reduce harmful image generation without altering the model. Generated by thinkingmachines/Inkling-Small As text-to-image generative models…
27 -
Hugging Face Daily Papers research 11d ago
PixRestore: Unified Image Restoration via Pixel Diffusion Transformer
Abstract PixRestore is a compact, VAE-free pixel-space diffusion transformer trained from scratch for unified image restoration, using flow matching on patchified pixels, DINO-based reliability-guided feature fusion, and adversarial fine-tuning to a one-step generator for…
9 -
Hugging Face Daily Papers research 11d ago
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
Abstract StartupBench evaluates end-to-end AI agents on real-world startup workflows and reveals that even top models complete only about 30% of tasks, highlighting gaps in instruction following and domain expertise. Generated by thinkingmachines/Inkling-Small Recent advances in…
7 -
Hugging Face Daily Papers research 11d ago
Abra: Scaling Diffusion Image Training
Abstract Scaling laws for text-to-image diffusion models reveal predictable compute-optimal training requiring far more data per parameter than language models, with robust overtraining behavior and universal curve shapes. Generated by thinkingmachines/Inkling-Small…
22 -
Hugging Face Daily Papers research 11d ago
aDSL: Agentic 3D Creation via Joint Agent-Program Design
Abstract A co-designed domain-specific language and multi-agent system improve LLM-driven 3D program synthesis by using relational operators and iterative execution feedback. Generated by thinkingmachines/Inkling-Small Programmatic representations provide a compelling paradigm…
31 -
Hugging Face Daily Papers research 11d ago
MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding
Abstract Mixture-of-Experts vision encoders with fine-grained topologies, auxiliary-loss-free balancing, and specialized kernels scale efficiently while outperforming larger dense models on image and video tasks. Generated by thinkingmachines/Inkling-Small Vision encoders are a…
21 -
Hugging Face Daily Papers research 11d ago
Demystifying Agent Skills: Why They Work-Until They Don't
Abstract Skills enhance LLM agents primarily by stabilizing execution through procedural anchoring rather than injecting missing knowledge, though retrieval bottlenecks and brittle assumptions limit their effectiveness. Generated by thinkingmachines/Inkling-Small Skills have…
22 -
Hugging Face Daily Papers research 11d ago
Cross-Model Memory Transfer via Target-Side Reader Adaptation
Abstract Cross-model reuse of frozen external memory tables depends primarily on aligning a lightweight target-side reader rather than the memory content alone, enabling reusable knowledge artifacts with optional adaptation. Generated by thinkingmachines/Inkling-Small Methods…
23 -
Hugging Face Daily Papers research 11d ago
EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing
Abstract EditBridge enables efficient ultra high-resolution image editing via a diffusion bridge that translates low-resolution edits to high-resolution outputs while preserving source details through sparse attention. Generated by thinkingmachines/Inkling-Small High-resolution…
11 -
Hugging Face Daily Papers research 11d ago
CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing
Abstract A new dataset, benchmark, and 22B model enable compositional instruction-guided video editing with multi-region attention and temporal coherence. Generated by thinkingmachines/Inkling-Small The quality and diversity of instruction-based video editing datasets are…
11 -
Hugging Face Daily Papers research 11d ago
MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verification-Guided Refinement
Abstract MathForm improves autoformalization by retrieving Mathlib knowledge and iteratively refining outputs with verification feedback, yielding a large verified dataset and a high-performing 8B model. Generated by thinkingmachines/Inkling-Small Autoformalization is commonly…
7 -
Hugging Face Daily Papers research 11d ago
Agent Lightning v1.0: Towards Harnessed Agentic RL
Abstract Agent Lightning v1.0 enables reproducible reinforcement learning for arbitrary agent harnesses, substantially improving coding-agent performance with minimal data and compute. Generated by thinkingmachines/Inkling-Small Modern agents operate inside agent harnesses that…
30 -
Hugging Face Daily Papers research 11d ago
HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety
Abstract HarnessRisk evaluates agent harness safety across six operational phases, revealing that configuration vulnerabilities and detection gaps allow high attack success despite preserved utility. Generated by thinkingmachines/Inkling-Small Large language models are…
31 -
Hugging Face Daily Papers research 11d ago
Unifying Graph Neural Networks Through a Common Layer Equation
Abstract A unified layer equation decomposes graph neural networks into seven components to compare architectures, derive theoretical bounds, and expose design choices linked to oversmoothing and expressivity. Generated by thinkingmachines/Inkling-Small Graph neural networks are…
16 -
Hugging Face Daily Papers research 11d ago
Personalized Auto-Research: Towards a True AI Co-Scientist
Abstract The paper introduces personalized auto-research, a framework that conditions AI-driven hypothesis generation, experimentation, and writing on individual researcher representations to avoid generic outputs. Generated by thinkingmachines/Inkling-Small AI co-scientists…
23 -
Hugging Face Daily Papers research 11d ago
FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution
Abstract FreeToken is an edge-native Mixture-of-Experts serving system that dynamically maps computation and model state onto heterogeneous local hardware to run large open-weight models on personal machines. Generated by thinkingmachines/Inkling-Small Frontier open-weight…
16 -
Hugging Face Daily Papers research 11d ago
Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation
Abstract TAMP-Nav improves embodied navigation by aligning vision-language models with 2D visual prompting, selective reasoning with compressed memory, and dense policy optimization. Generated by thinkingmachines/Inkling-Small Although Large Vision-Language Models (VLMs) have…
26 -
Hugging Face Daily Papers research 11d ago
Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents
Abstract Empirical evaluation of diverse memory substrates for long-horizon LLM agents reveals regime-dependent trade-offs, motivating adaptive substrate routing for reliable agent memory. Generated by thinkingmachines/Inkling-Small Memory is becoming core infrastructure for…
13 -
Hugging Face Daily Papers research 11d ago
Energy-Guided Flow Matching
Abstract Energy-Guided Flow Matching improves generative quality by progressively revealing high-frequency details through a moving endpoint and adaptive scheduling, reducing training cost and achieving state-of-the-art FID scores. Generated by thinkingmachines/Inkling-Small…
7 -
Hugging Face Daily Papers research 11d ago
From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation
Abstract A capability-driven data infrastructure with curriculum scheduling and specialized data engines trains large multimodal diffusion models on curated heterogeneous supervision for diverse generative tasks. Generated by thinkingmachines/Inkling-Small Large-scale image…
13 -
Hugging Face Daily Papers research 11d ago
GS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS Generation
Abstract GS-Voxel converts unstructured 3D Gaussian reconstructions into sparse structured latents to enable scalable aerial scene generation via flow models. Generated by thinkingmachines/Inkling-Small Many scalable latent 3D generators operate on structured tensors, whereas…
19 -
Hugging Face Daily Papers research 11d ago
Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection
Abstract Researchers evaluate indirect prompt injection risks in DeepSeek Harness using controlled taint and dual judges, finding notable success rates across text and file channels and recommending controls between untrusted content and sensitive actions. Generated by…
16 -
Hugging Face Daily Papers research 11d ago
ASI-Bench: At the Dawn of Artificial Superintelligence
Abstract Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on…
24 -
Hugging Face Daily Papers research 11d ago
From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents
Abstract RUPA models agent execution as a dependency graph to propagate uncertainty across long trajectories, improving failure detection and confidence estimation for LLM agents. Generated by thinkingmachines/Inkling-Small Reliable uncertainty quantification (UQ) is essential…
16 -
Hugging Face Daily Papers research 11d ago
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements
Abstract Agentic ESOpt uses evolution strategies for scalable full-parameter fine-tuning of long-horizon LLM agents via trajectory-level reward-weighted updates and parameter-context co-evolution. Generated by thinkingmachines/Inkling-Small Reinforcement Learning (RL) has been…
12 -
Hugging Face Daily Papers research 11d ago
Dynamic Multi-Byte Prediction With Hierarchical Language Models
Abstract Multi-byte prediction accelerates byte-level hierarchical language models by generating parallel bytes via variable-length windows and causal attention masking, improving inference speed with minimal quality loss. Generated by thinkingmachines/Inkling-Small Byte-level…
7 -
Hugging Face Daily Papers research 11d ago
StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
Abstract StateM is a runtime system that improves long-horizon agent execution through durable states, recoverable runbooks, and enforceable procedural controls without altering model weights. Generated by thinkingmachines/Inkling-Small Long-horizon agents can fail even when…
9 -
Hugging Face Daily Papers research 11d ago
StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding
Abstract StreamOPD improves streaming video understanding via on-policy distillation with verifiable rewards and a spatio-temporal cue-gating mechanism, achieving near-teacher performance without inference-time memory. Generated by thinkingmachines/Inkling-Small Streaming video…
10 -
Hugging Face Daily Papers research 11d ago
Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays
Abstract Per-field selective risk control for document extraction requires a validity ladder with fit/val splits and Mondrian PAC certificates, revealing that support-bin provenance outperforms learned fusion only under specific model conditions. Generated by…
11