Hugging Face Daily Papers
500 articles archived · Visit source ↗ · RSS
-
Hugging Face Daily Papers research 1d ago
Luce: Relightable Gaussians for 3D Asset Generation
Abstract Luce unifies geometry and PBR materials in a voxelized Gaussian cloud, using a variational autoencoder and rectified-flow transformer to generate relightable 3D assets from single images. Generated by thinkingmachines/Inkling-Small High-fidelity image-to-3D generation…
23 -
Hugging Face Daily Papers research 1d ago
CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes
Abstract CritICL improves LLM reasoning at inference time by using structured failure patterns from weaker models as critique-based guidance, reducing generation and token costs. Generated by thinkingmachines/Inkling-Small Recent advances in inference-time scaling have…
9 -
Hugging Face Daily Papers research 1d ago
What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference in Inspect Evals
Abstract Evaluation artifacts specify a forward computation: a task, scorer, and reported metric. They do not necessarily license the claim attached to that metric because the historical evidence and alternative semantics needed to replay it may be unbound. We formalize this…
25 -
Hugging Face Daily Papers research 1d ago
EditaLive! Unified Character Video Editing for Live Streaming
Abstract EditaLive enables real-time human-centric live-stream video editing by adapting an image animation model to causal streaming generation with distilled two-step sampling and sparse attention. Generated by thinkingmachines/Inkling-Small Conventional video editing…
10 -
Hugging Face Daily Papers research 1d ago
Magpie: Real-Time World Renderer for Interactive Games
Abstract Magpie is a real-time generative rendering system that separates gameplay logic from visual generation to preserve interactive designability while reducing asset requirements for game prototypes. Generated by thinkingmachines/Inkling-Small Modern game development relies…
24 -
Hugging Face Daily Papers research 1d ago
Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher
Abstract Self-OPD eliminates task-specific teachers in flow matching by using self-explored stochastic branches and normalized advantages to optimize the velocity field for multi-objective alignment. Generated by thinkingmachines/Inkling-Small On-policy distillation (OPD), which…
26 -
Hugging Face Daily Papers research 2d ago
TacForcing: Streaming Action Generation with Execution-Time Tactile Feedback
Abstract TacForcing is a streaming action-generation framework that integrates real-time tactile feedback during execution via a streaming action expert and execution-aware tactile attention, improving contact-rich manipulation. Generated by thinkingmachines/Inkling-Small…
9 -
Hugging Face Daily Papers research 2d ago
WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution
Abstract WikiSkill co-evolves reusable agent skills with a persistent knowledge base to systematically accumulate experience and improve performance across models. Generated by thinkingmachines/Inkling-Small Agent skills package specialized knowledge and workflows into reusable…
37 -
Hugging Face Daily Papers research 2d ago
PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents
Abstract PILOT enables live self-improvement by allowing a supervisor to steer active workers and distilling execution experience into reusable skills, improving accuracy and efficiency. Generated by thinkingmachines/Inkling-Small Long-horizon agent runs generate experience that…
26 -
Hugging Face Daily Papers research 2d ago
Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models
Abstract Game engines provide executable verification and long-horizon trajectories for reinforcement learning post-training of spatial world models, motivating a human-engine verification paradigm. Generated by thinkingmachines/Inkling-Small A common strategy for scaling world…
34 -
Hugging Face Daily Papers research 2d ago
Procedura: Agentic 3D Modeling with Procedural Control
Abstract Procedura is a 3D modeling agent that generates editable, part-structured procedural assemblies with sharp geometry and validated articulation from text prompts. Generated by thinkingmachines/Inkling-Small Native 3D generators now recover impressive mesh geometry from a…
8 -
Hugging Face Daily Papers research 2d ago
Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning
Abstract An agentic framework combining LLMs and VLMs enables consistent, multi-instruction editing of long multi-shot videos while preserving spatiotemporal structure. Generated by thinkingmachines/Inkling-Small While generative AI has significantly advanced video editing,…
16 -
Hugging Face Daily Papers research 2d ago
GameWAM: A World Action Model for Video Games
Abstract GameWAM is a unified world-action model for native video-game control that jointly predicts future visuals and executable keyboard-mouse actions using block-causal flow matching, mode-specific distributions, and block-cycle replanning. Generated by…
35 -
Hugging Face Daily Papers research 2d ago
TTPO: Test-Time Policy Optimization
Abstract Test-Time Policy Optimization enables label-free test-time training for mathematical reasoning by asymmetrically distilling agreeing rollouts and penalizing disagreeing ones, matching supervised performance. Generated by thinkingmachines/Inkling-Small Recent prominent…
6 -
Hugging Face Daily Papers research 2d ago
Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning
Abstract Aphanta evaluates when image-editing intermediates improve multimodal reasoning by testing direct, editor-generated, and idealized visual states across tasks. Generated by thinkingmachines/Inkling-Small Explicit visual intermediates can help multimodal large language…
26 -
Hugging Face Daily Papers research 2d ago
CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval
Abstract CaSKG calibrates procedural skill relations via counterfactual-causal graph construction to improve compact, executable retrieval for LLM agents. Generated by thinkingmachines/Inkling-Small Reusable skill libraries allow large language model (LLM) agents to reuse…
10 -
Hugging Face Daily Papers research 2d ago
Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report
Abstract Harness-Aware Training enables compact models to adapt to evolving digital-avatar harness configurations with low latency and high accuracy. Generated by thinkingmachines/Inkling-Small AI-powered digital avatar streamers must answer product questions, engage viewers,…
24 -
Hugging Face Daily Papers research 2d ago
CaRGo-T: Causal Reasoning Graph-of-Thought improves Multimodal Humor Comprehension
Abstract CaRGo-T improves multimodal humor understanding by modeling causal relationships as graph-based reasoning structures interpreted by vision-language models. Generated by thinkingmachines/Inkling-Small Large-scale vision-language models (VLMs) have demonstrated remarkable…
5 -
Hugging Face Daily Papers research 2d ago
PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
Abstract The study formalizes probabilistic alignment for world models, introduces PAWBench and PAWEval to evaluate video generators as stochastic samplers, and finds current models fail to match reference behavior distributions. Generated by thinkingmachines/Inkling-Small…
8 -
Hugging Face Daily Papers research 2d ago
UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City
Abstract UrbanGround evaluates whether multimodal language model agents can sustain reliable navigation and spatial reasoning in a realistic 3D city replica, revealing that local perceptual skills fail to compose into extended goal-directed behavior. Generated by…
37 -
Hugging Face Daily Papers research 2d ago
Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization
Abstract Zero-WAM enables robotic manipulation of unseen tasks by conditioning a causal video-action model on in-context human video guidance, supported by an automatically generated dataset and a future-chunk prediction objective. Generated by thinkingmachines/Inkling-Small…
21 -
Hugging Face Daily Papers research 2d ago
Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO
Abstract Evolution strategies improve reasoning diversity and Pass@K over GRPO through sparse functional updates and population diversity, supporting a hybrid training approach. Generated by thinkingmachines/Inkling-Small Evolution Strategies (ES) have recently emerged as a…
30 -
Hugging Face Daily Papers research 2d ago
What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents
Abstract Agentic data generation is framed as constrained distribution design over factorized experience tuples, emphasizing execution-grounded accuracy, learner-relative complexity, and diversity rather than scale alone. Generated by thinkingmachines/Inkling-Small LLM agents…
30 -
Hugging Face Daily Papers research 2d ago
LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale
Abstract LibriBrain100 is a large-scale MEG speech dataset that demonstrates improved decoding through extensive within-subject recordings and multi-subject supervised fine-tuning of pre-trained models. Generated by thinkingmachines/Inkling-Small We introduce LibriBrain100, a…
12 -
Hugging Face Daily Papers research 3d ago
SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?
Abstract The study introduces a benchmark for evaluating autonomous software migration by coding agents, finding that current models rarely complete migrations correctly. Generated by thinkingmachines/Inkling-Small Modern software systems accumulate technical debt over decades…
33 -
Hugging Face Daily Papers research 3d ago
Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans
Abstract A real-time framework for online co-speech gesture generation uses a causal multimodal autoregressive model with streaming speech and motion history, supported by synthetic dialogue data and continual user-feedback adaptation. Generated by thinkingmachines/Inkling-Small…
13 -
Hugging Face Daily Papers research 3d ago
A Programming Paradigm for Spatiotemporal Composability
Abstract A calculus and framework unify revertible effects and reactive coeffects into a context paradigm enabling spatiotemporal composability for dynamic software components. Generated by thinkingmachines/Inkling-Small Modern software -- from plugin systems to self-evolving…
17 -
Hugging Face Daily Papers research 3d ago
Code World Model: Coding Agent as World Brain
Abstract Code World Model separates persistent world dynamics from visual rendering by using a language model to generate executable state updates and a video model to render observations from proxy representations. Generated by thinkingmachines/Inkling-Small World models aim to…
21 -
Hugging Face Daily Papers research 3d ago
A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans
Abstract A modular medical imaging agent decomposes spatial relation verification into parsing, anatomical localization, and geometric rules to outperform end-to-end vision-language models on CT spatial reasoning. Generated by thinkingmachines/Inkling-Small Reliable spatial…
27 -
Hugging Face Daily Papers research 3d ago
RetrievalRouter: Joint Modality and Architecture Selection for Document Retrieval
Abstract RetrievalRouter adaptively selects retrieval pipelines per query to improve both accuracy and speed across diverse document benchmarks. Generated by thinkingmachines/Inkling-Small Document retrieval increasingly supports high-stakes information access in finance,…
10 -
Hugging Face Daily Papers research 3d ago
Real-TurnTurk: A Multimodal Turkish Corpus for Turn-Taking Prediction
Abstract A multimodal Turkish dialogue dataset and genetic algorithm-optimized interpretable rules are used to predict turn transitions from visual, acoustic, and linguistic cues. Generated by thinkingmachines/Inkling-Small Turn-taking is a basic organizational feature of human…
10 -
Hugging Face Daily Papers research 3d ago
Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data
Abstract Mixed supervised fine-tuning on combined reasoning corpora outperforms next-chunk reinforcement learning in efficiency and final accuracy across mathematical and out-of-domain tasks. Generated by thinkingmachines/Inkling-Small Recent work proposes next-chunk reasoning…
18 -
Hugging Face Daily Papers research 3d ago
The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents
Abstract Coding agents perform long-running tasks spanning dozens of model calls, tool uses, and code edits. As these runs unfold, users face a practical cost-quality trade-off: escalating to a stronger model when a cheaper one struggles, or downshifting once the hard reasoning…
37 -
Hugging Face Daily Papers research 3d ago
Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds
Abstract JoyAI-Echo-1.5 unifies long-form video and interactive world generation through cross-shot memory, geometry-aware camera control, and rollout-aware training to maintain identity and coherence over extended sequences. Generated by thinkingmachines/Inkling-Small Video…
36 -
Hugging Face Daily Papers research 3d ago
Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios
Abstract A new benchmark evaluates how well multimodal language models follow diverse video-based instructions with visual, audio, and structural constraints. Generated by thinkingmachines/Inkling-Small Multimodal Large Language Models (MLLMs) have shown strong performance in…
14 -
Hugging Face Daily Papers research 3d ago
FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling
Abstract FIRM-Video uses checklist-driven verification of temporal visual evidence to build reliable reward models for text-to-video evaluation and alignment. Generated by thinkingmachines/Inkling-Small Reliable reward models are essential for text-to-video evaluation and…
32 -
Hugging Face Daily Papers research 3d ago
Rubrics as Visual-Repair Context for Self-Evolving UI-to-Code Generation
Abstract RubSE improves UI-to-code generation stability by using rubric-guided self-evolution to prevent visual repair coupling and trajectory collapse. Generated by thinkingmachines/Inkling-Small Large vision-language models have shown strong progress in UI-to-code generation,…
26 -
Hugging Face Daily Papers research 3d ago
Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models
Abstract Stream4D improves autoregressive video generation by replacing static 3D critics with a dynamic 4D reconstruction reward and motion prior to preserve coherent motion and reduce geometric drift. Generated by thinkingmachines/Inkling-Small Streaming autoregressive…
38 -
Hugging Face Daily Papers research 3d ago
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
Abstract VoiceMem introduces a dual-brain streaming memory architecture for speech language models that improves retrieval accuracy, emotional personalization, and real-time efficiency. Generated by thinkingmachines/Inkling-Small Conversational systems, such as duplex speech…
15 -
Hugging Face Daily Papers research 3d ago
MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization
Abstract MA-VLA enables multi-arm collaboration by assigning atomic actions to individual arms and using training-time permutations to generalize to unseen coordination patterns. Generated by thinkingmachines/Inkling-Small Multi-arm collaboration is becoming a core capability in…
31 -
Hugging Face Daily Papers research 3d ago
StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models
Abstract StreamPI enhances single-frame vision-language-action models with streaming temporal reasoning via instruction-anchored attention and randomized interval training, improving robot manipulation without extra parameters. Generated by thinkingmachines/Inkling-Small…
6 -
Hugging Face Daily Papers research 3d ago
Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation
Abstract Multi-teacher on-policy distillation suffers from token-level budget misallocation across domains, which is addressed by balancing, dynamic allocation, and reward refresh to recover most of the oracle ensemble's capability. Generated by thinkingmachines/Inkling-Small…
10 -
Hugging Face Daily Papers research 3d ago
V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning
Abstract Visual Rubrics-Based Reinforcement Learning improves vision-language model grounding by scoring answers on visual faithfulness, reasoning consistency, and instruction following using structured partial credit. Generated by thinkingmachines/Inkling-Small Vision-language…
22 -
Hugging Face Daily Papers research 3d ago
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning
Abstract VBBR-Pro introduces a closed-loop testbed that enables scalable, verifiable, and controllable native visual reasoning through generation across diverse visual substrates. Generated by thinkingmachines/Inkling-Small Native visual reasoning treats visual generation as the…
6 -
Hugging Face Daily Papers research 3d ago
JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution
Abstract JIT-Agent is a trainable model that synthesizes adaptive agent harnesses for off-the-shelf LLMs, improving performance across diverse models and tasks. Generated by thinkingmachines/Inkling-Small Agent capability is not determined by the model alone. The agent harness,…
34 -
Hugging Face Daily Papers research 3d ago
VGI-BENCH: Probing Visual Intelligence in Video Generation Models
Abstract VGI-bench evaluates visual reasoning in video generation models through 27 tasks, revealing limited reliability and minimal self-correction during generation. Generated by thinkingmachines/Inkling-Small Recent studies suggest that video generation models can exhibit…
25 -
Hugging Face Daily Papers research 3d ago
Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments
Abstract AnTrap benchmarks GUI agent robustness by injecting dynamic anomalies into execution trajectories, revealing universal vulnerabilities and distinguishing learnable traps from intrinsic reasoning limits. Generated by thinkingmachines/Inkling-Small GUI agents often…
19 -
Hugging Face Daily Papers research 3d ago
D^3-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation
Abstract D³-MOPD dynamically adjusts domain sampling ratios during multi-teacher distillation by monitoring per-domain reverse-KL trajectories, improving convergence efficiency and closing most of the student-to-teacher performance gap. Generated by…
31 -
Hugging Face Daily Papers research 3d ago
FrontierChallenge: Evaluating Scientific Workflow Completion
Abstract FrontierChallenge evaluates end-to-end scientific workflows across domains, revealing that frontier models complete only about 20% of tasks despite high partial scores and frequent claims of completion. Generated by thinkingmachines/Inkling-Small Scientific agents…
23 -
Hugging Face Daily Papers research 3d ago
Agent-G^2: Gaussian Guidance for Agentic Reinforcement Learning
Abstract Agent-G² models hint depth as a Gaussian distribution estimated online from existing rollouts, improving reinforcement learning on long-horizon tasks without extra probing. Generated by thinkingmachines/Inkling-Small Hint-based reinforcement learning addresses reward…
31