Hugging Face Daily Papers
500 articles archived · Visit source ↗ · RSS
-
Hugging Face Daily Papers research 24d ago
Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data
Abstract Learning generalizable robot manipulation policies requires large-scale and diverse demonstration data. Egocentric human manipulation videos offer rich scene and task diversity, and prior work has shown that retargeting and rendering such videos into robot-format data…
12 -
Hugging Face Daily Papers research 24d ago
Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance
Abstract Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training large language models (LLMs). While Group Relative Policy Optimization (GRPO) is widely adopted, it suffers from sparse reward signals and loses gradients entirely…
22 -
Hugging Face Daily Papers research 24d ago
BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation
Abstract Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as a promising paradigm for 3D robot manipulation. However, existing 3D VLA methods remain data-hungry, exhibit limited generalization under distribution…
4 -
Hugging Face Daily Papers research 24d ago
WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models
Abstract Interactive video world models are essential for long-horizon planning and exploration, yet they suffer from compounding errors. Post-training methods such as reinforcement learning (RL) can improve these models, but they hit a verification bottleneck: for arbitrary…
26 -
Hugging Face Daily Papers research 24d ago
OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents
Abstract LLM agents are increasingly applied to open-ended everyday requests that span work, study, and life. These tasks are long-horizon, cross-environment, and multimodal, forcing the agent to preserve goals and constraints across many steps while navigating heterogeneous…
6 -
Hugging Face Daily Papers research 24d ago
When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation
Abstract On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense token-level teacher signals. Recent selective OPD methods improve this process by prioritizing signals that are confident, informative, or learnable.…
17 -
Hugging Face Daily Papers research 24d ago
The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads
Abstract Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains unexamined. We study over-inference (OI): the phenomenon where LLMs fabricate user attributes beyond what evidence supports. We introduce MirageBench,…
29 -
Hugging Face Daily Papers research 24d ago
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning
Abstract Long-horizon reasoning in recent LLMs demands that the model switch between distinct skills inside a reasoning chain, such as first doing a math derivation, then using the result to plan a schedule. We call such problems cross-skill long-horizon tasks: multi-step tasks…
30 -
Hugging Face Daily Papers research 24d ago
When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents
Abstract Memory-augmented VLM agents act on persistent spatial knowledge, yet that knowledge silently goes stale as the environment changes. We ask what happens when an agent must reconcile a confident memory claim with a contradicting observation, and whether current models can…
11 -
Hugging Face Daily Papers research 24d ago
NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap
Abstract We introduce NOLLI, a procedurally generated English-Korean puzzle benchmark designed to diagnose where Korean performance gaps arise. It comprises 15 puzzle types (25 tasks; 7,500 items), with every instance seed-regenerable, verified to have a unique solution, and…
18 -
Hugging Face Daily Papers research 24d ago
SKILL-KD: Contrastive Skill Distillation for LLM Agents
Abstract Skill-based prompting has become a practical mechanism for improving large language model (LLM) agents, yet existing skill acquisition methods often treat skills as experience summaries, memory entries, or direct summaries of successful demonstrations. This creates a…
13 -
Hugging Face Daily Papers research 24d ago
UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models
Abstract The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance user experiences. However, producing photorealistic…
4 -
Hugging Face Daily Papers research 24d ago
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
Abstract Vision offers a critical axis for advancing foundation models, driving a shift towards natively unified multimodal pretraining. Despite this momentum, the design space and the fundamental mechanisms of how modalities interact during unified training remain…
12 -
Hugging Face Daily Papers research 24d ago
Poly-OPD: Heterogeneous Multi-Teacher On-Policy Distillation for Capability-Selectable Flow Models
Abstract Leading open text-to-image models often carry complementary strengths: one may lead on preference-aligned aesthetics while another follows compositional instructions more faithfully. However, differences in their autoencoders and noise schedules make it difficult to…
23 -
Hugging Face Daily Papers research 24d ago
TriGlue: a Biology-Inspired Generative Model for Generating Molecular Glue-Induced Ternary Complex
Abstract Molecular glue degraders have emerged as a promising strategy for targeted protein degradation by inducing ternary complex formation between an E3 ubiquitin ligase and a target protein. Despite their therapeutic potential, computational design of molecular glues remains…
24 -
Hugging Face Daily Papers research 25d ago
ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning
Abstract On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories from stronger expert models. However, when the expert fails on harder problems, existing…
9 -
Hugging Face Daily Papers research 25d ago
CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning
Abstract Chart question answering (CQA) requires multimodal large language models (MLLMs) to integrate visual comprehension with logical reasoning, yet current models struggle with accurate visual grounding and coherent reasoning chains. While extrinsic chain-of-thought…
11 -
Hugging Face Daily Papers research 25d ago
RestoreKV: Recovering Full-Cache Behavior Under Aggressive Query-Agnostic KV Cache Eviction
Abstract Query-agnostic KV cache eviction compresses a context once and reuses the resulting cache for arbitrary future queries, but performance can collapse under tight budgets. Existing methods primarily improve which original KV pairs are retained. We introduce RestoreKV,…
8 -
Hugging Face Daily Papers research 25d ago
ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels
Abstract Historical language change affects morphology, syntax, semantics, and pragmatics, yet computational studies typically examine these levels with incompatible representations and therefore cannot determine whether they evolve together across languages. We address this…
10 -
Hugging Face Daily Papers research 25d ago
Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation
Abstract MLLM-based segmentation faces a core segmentation trilemma: high segmentation performance, preserved dialogue ability, and fast inference. Embedding-prediction methods may disrupt language modeling through pixel-level objectives, whereas next-token generation is…
36 -
Hugging Face Daily Papers research 25d ago
When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings
Abstract We identify a previously overlooked failure mode of ALiBi positional encoding: its linear bias scaling underflows floating-point precision, which zeroes out a large fraction of attention weights and renders the affected attention heads partially blind. We analyze this…
27 -
Hugging Face Daily Papers research 25d ago
Multi-Task Multi-Frame Visual Piano Transcription
Abstract Audio-based piano transcription performs well on onset, pitch, and velocity, but the sustain pedal lets sound persist long after key release, so audio systems predict pedal-extended offsets rather than physical key release. Yet existing Visual Piano Transcription (VPT)…
38 -
Hugging Face Daily Papers research 25d ago
Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements
Abstract Do Large Language Models (LLMs) possess genuine structural reasoning, or merely rely on surface-level pattern matching? The financial domain, demanding numerical precision and multi-step logic over long contexts, is an ideal testbed. Existing benchmarks fail to capture…
10 -
Hugging Face Daily Papers research 25d ago
Decoding Children's Gait Behavior
Abstract We introduce a new problem domain for human action recognition: the fine-grained analysis of children's gait behaviors from standard RGB video. We specifically target the ambulatory patterns of children aged 3-17 years. Such behaviors arise naturally in the diagnosis…
37 -
Hugging Face Daily Papers research 25d ago
MiniWorld: Democratizing the Training of Video World Models from Scratch
Abstract Video world models predict future observations conditioned on historical observations and control signals, enabling long-horizon generation through autoregressive state transitions. Unlike conventional video generation models that primarily capture visual appearance and…
18 -
Hugging Face Daily Papers research 25d ago
PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents
Abstract Recursive self-improvement requires agents to turn accumulated experience into better future behavior. Personal AI agents offer a concrete setting for studying this capability because they retain preferences, task histories, tool routines, and learned skills across…
6 -
Hugging Face Daily Papers research 25d ago
ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?
Abstract Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear whether these systems can effectively evolve their skills and whether the resulting skills improve task-solving capabilities. To bridge…
8 -
Hugging Face Daily Papers research 25d ago
PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs
Abstract Scientific poster construction compresses a long multimodal paper into a readable, editable canvas. Existing systems hide request-level failures by scoring only completed outputs; direct image generation is not element-editable, while coding-agent workflows are costly.…
35 -
Hugging Face Daily Papers research 25d ago
LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models
Abstract Diffusion language models (dLLMs) offer an alternative to autoregressive (AR) language modeling, yet the scaling behavior of Mixture-of-Experts (MoE) dLLMs remains poorly understood. We systematically characterize how optimization hyperparameters, compute allocation,…
18 -
Hugging Face Daily Papers research 25d ago
When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills
Abstract Persona skills distill personal interaction histories into portable and executable artifacts for downstream agents. While enabling flexible personalization, this process concentrates fragmented personal signals, amplifies their impact through reuse, and challenges…
29 -
Hugging Face Daily Papers research 25d ago
Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent
Abstract We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open-web exploration. Preliminary evaluations reveal two critical bottlenecks in…
8 -
Hugging Face Daily Papers research 25d ago
Push-Wiper: Toward General-Purpose Robotic Cleaning across Varied Stains and Surfaces with Segmented Pushing Trajectories
Abstract Viscous stains, characterized by high viscosity and complex rheological properties, remain a major challenge for robotic surface cleaning. Conventional wiping often spreads the stain, while scrubbing provides stronger friction but risks damaging the surface. In this…
12 -
Hugging Face Daily Papers research 25d ago
ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts
Abstract World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual dynamics. However, their reliance on pixel-generative future supervision can entangle action-relevant state transitions with task-irrelevant visual…
5 -
Hugging Face Daily Papers research 25d ago
Quo Vadis, World Modeling?
Abstract Continually improving agents require dynamic interaction feedback beyond static supervision, yet direct real-environment interaction is costly, slow, unsafe, and hard to parallelize. World modeling offers a natural intermediate proxy that allows agents to query…
19 -
Hugging Face Daily Papers research 25d ago
ExplainBench: Evaluating Code Explanations from Agents
Abstract Large Language Model (LLM) agents have seen rapid adoption in software engineering. As agents take a greater role in the actual generation of code, they are making larger changes, spanning tens to hundreds of lines. This makes manual review of agent results increasingly…
30 -
Hugging Face Daily Papers research 25d ago
Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation
Abstract Industrial recommenders increasingly adopt the pretrain-then-transfer paradigm, yet behavioral distribution drift raises two questions: what to learn from behavior sequences, and how to transfer the learned knowledge while the pretrained model is continually refreshed.…
29 -
Hugging Face Daily Papers research 25d ago
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion
Abstract Real-time video editing requires low-latency causal generation with bounded computational resources while preserving source fidelity and long-term temporal consistency. We present JoyAI-Video-Edit, a 16B-parameter autoregressive diffusion framework for real-time,…
32 -
Hugging Face Daily Papers research 25d ago
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning
Abstract Large language model agents have shown strong potential in complex interactive tasks, yet their reinforcement learning (RL) is often hindered by sparse rewards, as a long multi-turn trajectory may receive only a single outcome-level signal. On-policy self-distillation…
24 -
Hugging Face Daily Papers research 25d ago
Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing
Abstract Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing. However, unified 3D modeling remains constrained by scarce multimodal data, particularly the lack of large-scale and…
8 -
Hugging Face Daily Papers research 25d ago
OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models
Abstract Omni-modal large language models (Omni-LLMs) have achieved remarkable performance on audio-visual understanding tasks, but processing long and highly redundant visual and audio token sequences incurs substantial computational overhead, demanding aggressive token…
6 -
Hugging Face Daily Papers research 25d ago
AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling
Abstract Language remains an outlier in generative modeling: while images, video, and audio are increasingly modeled in continuous latent spaces, text generation still relies predominantly on discrete tokens. Existing continuous language models either inherit embedding spaces…
32 -
Hugging Face Daily Papers research 25d ago
CAPEval: A Decoupled Caption Evaluation across Understanding and Generation
Abstract Captions serve as a primary supervision signal for both multimodal understanding and text-to-image generation. However, previous evaluations treat the caption quality as a single scalar objective, which conflates two distinct properties: (1) how much visual information…
22 -
Hugging Face Daily Papers research 25d ago
TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning
Abstract Tool-Integrated Reasoning (TIR) enables LLMs to solve complex tasks through iterative tool interactions. However, existing reinforcement learning methods often rely on trajectory-level supervision, limiting fine-grained credit assignment in long-horizon TIR scenarios.…
26 -
Hugging Face Daily Papers research 25d ago
SkillJack: Persistent Skill Backdoors in Self-Evolving Agents
Abstract Self-evolving agents increasingly convert interaction histories into reusable skills that persist beyond individual tasks. While prior work studies memory and retrieval poisoning, such attacks only affect agents when poisoned records are retrieved as context. We uncover…
28 -
Hugging Face Daily Papers research 25d ago
UniWorld-Design: From Pixel Generation to Layer-Native Design
Abstract We introduce UniWorld-Design, a framework that redefines image generation from flat pixel synthesis to structured visual composition, with semantic RGBA layers as the atomic units of generation, understanding, and editing. Our key insight is that pixels define how an…
17 -
Hugging Face Daily Papers research 25d ago
GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience
Abstract A wearable assistant should both answer questions about its visual history and recognize when that history is useful to the present situation. Existing video-memory systems primarily support question-conditioned recall, whereas proactive assistants typically use…
25 -
Hugging Face Daily Papers research 25d ago
Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging
Abstract On-policy distillation, in which a teacher corrects samples that the student itself generates, presupposes that the two models speak the same language: identical VAE latents, matching architectures, and a common timestep grid. We ask what happens when none of this…
28 -
Hugging Face Daily Papers research 26d ago
Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge
Abstract Enterprise question answering requires models to acquire proprietary knowledge without discarding general capabilities. We present Wnuan, a three-stage pipeline that constructs task-oriented supervision from documents, performs supervised fine-tuning with general-data…
24 -
Hugging Face Daily Papers research 26d ago
To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing
Abstract Large language models increasingly write and repair production code, yet evidence is mounting that their test-passing patches leave codebases harder to maintain. We identify one concrete source: deletion avoidance, the systematic tendency to retain code that an intended…
8 -
Hugging Face Daily Papers research 26d ago
SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space
Abstract World Action Models (WAMs) couple action generation with prediction of future states. Their effectiveness depends on whether future dynamics are modeled in a space that is both aligned with action generation and sufficiently geometry-aware to capture where and how…
23