Hugging Face Daily Papers
500 articles archived · Visit source ↗ · RSS
-
Hugging Face Daily Papers research 13d ago
PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment
Abstract PRM-as-a-Judge 1.5 provides fine-grained process metrics and reliability tools to evaluate embodied robotic models beyond binary success rates. Generated by thinkingmachines/Inkling-Small Fine-grained robotic evaluation matters for understanding embodied models, going…
33 -
Hugging Face Daily Papers research 13d ago
MobileMem: Learning from a Year of Mobile Experiences
Abstract MobileMem is a benchmark and framework for evaluating on-device long-term memory through year-scale, multimodal mobile experience trajectories that require temporal reasoning, knowledge updating, and preference inference. Generated by thinkingmachines/Inkling-Small The…
5 -
Hugging Face Daily Papers research 13d ago
Multimodal Model Diffing for Feature Discovery and Control
Abstract MMDiff uses multimodal sparse autoencoders to isolate, detect, and control specific features in multimodal language models, improving interpretability and targeted steering of visual and safety behaviors. Generated by thinkingmachines/Inkling-Small Multimodal Large…
20 -
Hugging Face Daily Papers research 13d ago
CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing
Abstract CPI-Bench is a comprehensive benchmark for real-world image editing that evaluates multi-image tasks, practical applications, and reasoning-based editing to better differentiate model performance. Generated by thinkingmachines/Inkling-Small With the rapid advancement of…
36 -
Hugging Face Daily Papers research 13d ago
Verifier-Induced Support Reshaping in On-Policy Optimization
Abstract On-policy reinforcement learning with verifiable rewards can improve immediate task performance while reducing the diversity of successful responses needed for future training, a phenomenon called verifier-induced support reshaping. Generated by…
16 -
Hugging Face Daily Papers research 13d ago
Scaling Domain Data Repetition in LLM Pretraining
Abstract Under proportional scaling of model size and training tokens, optimal repetition of high-quality domain data increases mildly with scale and correlates with domain validation loss rather than unique data volume. Generated by thinkingmachines/Inkling-Small As large…
9 -
Hugging Face Daily Papers research 13d ago
HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark
Abstract HumanTracker introduces a large-scale benchmark and preference-aligned metric to evaluate humanoid motion tracking based on perceptual quality and physical contact stability. Generated by thinkingmachines/Inkling-Small Humanoid motion tracking is central to…
5 -
Hugging Face Daily Papers research 13d ago
Marionette: Predicting World States, Rendering Geometry, Painting Appearance
Abstract Marionette predicts explicit 3D articulated world states for interactive games, uses a fixed renderer for geometry, and synthesizes video via diffusion, enabling direct state-level control and long-horizon consistency repair. Generated by thinkingmachines/Inkling-Small…
6 -
-
Hugging Face Daily Papers research 13d ago
Claim-Level Reliability Assessment for Efficient Test-Time Reasoning
Abstract Claim-Level Reliability Assessment improves reasoning accuracy by verifying critical claims instead of sampling more solutions, reducing token use while boosting performance. Generated by thinkingmachines/Inkling-Small We propose claim-level falsification as a principle…
23 -
Hugging Face Daily Papers research 13d ago
Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
Abstract Frontier autonomous agents excel at engineering optimization but show unstable performance, limited novelty, and variable experience reuse across long-horizon tasks. Generated by thinkingmachines/Inkling-Small Autonomous agents are increasingly capable of improving…
14 -
Hugging Face Daily Papers research 13d ago
Dion3: Full-Stack Orthogonal Updates
Abstract Dion3 accelerates the Muon optimizer by reducing orthogonalization and communication overhead through algorithmic, kernel-level, and update-rule improvements. Generated by thinkingmachines/Inkling-Small The Muon optimizer incurs a significant overhead cost due to its…
9 -
Hugging Face Daily Papers research 14d ago
UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers
Abstract UNMASK automatically discovers and mitigates spurious correlations in text classifiers via causal verification and group-based reweighting without manual annotations. Generated by thinkingmachines/Inkling-Small Neural language models trained on large crowdsourced…
16 -
Hugging Face Daily Papers research 15d ago
Maglev: Sliding Recurrent Memory
Abstract A recurrent Transformer with fixed-size memory and coupled prefiller-decoder training improves long-context modeling while enabling efficient parallel training and reduced inference cost. Generated by thinkingmachines/Inkling-Small We introduce , a recurrent Transformer…
38 -
Hugging Face Daily Papers research 15d ago
Thought-Level Beam Search for Reasoning
Abstract Gambit improves reasoning model efficiency by using thought-level beam search to dynamically allocate compute to promising reasoning traces under fixed hardware budgets. Generated by thinkingmachines/Inkling-Small Test-time compute scaling is a primary driver of…
12 -
Hugging Face Daily Papers research 15d ago
RibAssist 3D: Biplanar Rib-Fracture Detection, Addressing, and Selective 3D Localization from CT-Derived Projections
Abstract Cross-view pairing of rib fractures in orthogonal CT projections enables accurate 3D localization, but correspondence confidence remains the primary bottleneck. Generated by thinkingmachines/Inkling-Small Rib fractures are common and time-consuming to localize on…
17 -
Hugging Face Daily Papers research 15d ago
Context-Matched Distillation: Teacher Causality for Autoregressive Video Distillation
Abstract Context-Matched Distillation aligns teacher supervision with causal generation context for few-step autoregressive video models, improving control adherence and long-video quality. Generated by thinkingmachines/Inkling-Small Interactive autoregressive video generation…
12 -
Hugging Face Daily Papers research 16d ago
From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs
Abstract Researchers propose a black-box red-teaming method using inaudible low-frequency waveforms to expose vulnerabilities in audio-language models, alongside a defense that detects distribution shifts and requests a second recording to recover accuracy. Generated by…
37 -
Hugging Face Daily Papers research 16d ago
Mitigating Gender Bias in English to Romanian Machine Translation
Abstract A hybrid pipeline combining LLM gender classification with tag-aware neural translation improves gender accuracy in English-to-Romanian machine translation. Generated by thinkingmachines/Inkling-Small Machine translation (MT) systems often fail to correctly translate…
35 -
Hugging Face Daily Papers research 16d ago
Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing
Abstract HPSE improves unstructured knowledge editing by distilling from hybrid rollouts that insert missing facts into the model's reasoning paths, enabling composable multi-hop reasoning. Generated by thinkingmachines/Inkling-Small Large language models (LLMs) achieve…
25 -
Hugging Face Daily Papers research 16d ago
OmniScientist: An Omni-Modal Omni-Discipline AI Scientist
Abstract OmniScientist is an end-to-end omni-modal AI scientist that performs multidisciplinary research directly from heterogeneous raw evidence using autonomous agents and lifecycle-wide perception, improving evidence-grounded discovery across diverse scientific modalities.…
23 -
Hugging Face Daily Papers research 16d ago
Intern-S2-Preview: Scientific Agentic Foundation Model
Abstract Intern-S2-Preview is a scientific agentic foundation model series that integrates multimodal pre-training, multi-task reinforcement learning, and memory-augmented extensions to support long-horizon scientific reasoning and forecasting. Generated by…
33 -
-
Hugging Face Daily Papers research 16d ago
PixSDS: Why Latent SDS Makes Noisy Pixels
Abstract PixSDS fixes VAE-induced pixel drift in latent score distillation sampling by guiding optimization with decoded image directions, reducing artifacts in text-to-3D generation. Generated by thinkingmachines/Inkling-Small Score Distillation Sampling (SDS) enables…
14 -
Hugging Face Daily Papers research 16d ago
TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operational Validity Enforcement
Abstract TailBooster uses a dual-layer generative framework with statistical tail extraction and deep autoencoder cleaning to synthesize operationally valid extreme air-transport events, substantially improving extreme-event prediction accuracy. Generated by…
26 -
Hugging Face Daily Papers research 16d ago
Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity
Abstract Instruction tuning changes model confidence and reduces rationale diversity without improving calibration, indicating distinct effects on reasoning and certainty. Generated by thinkingmachines/Inkling-Small Instruction-tuned language models achieve strong performance…
26 -
Hugging Face Daily Papers research 16d ago
Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning
Abstract CaRL uses reinforcement learning with refusal incentives and hindsight augmentation to reduce futile reasoning in large language models while preserving task performance. Generated by thinkingmachines/Inkling-Small Large language models generate computationally…
28 -
Hugging Face Daily Papers research 16d ago
An AI4AI Framework for Visual Token Pruning
Abstract AutoPrune uses large language models to automatically design visual-token pruning policies for multimodal models via a domain-specific language and residual search formulation, achieving high efficiency with minimal performance loss. Generated by…
27 -
Hugging Face Daily Papers research 16d ago
SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models
Abstract SKILLER is a reinforcement learning framework that automatically generates tailored skills for small open-source models to reduce inference costs while maintaining high task performance. Generated by thinkingmachines/Inkling-Small Agent skills represent a standardized…
36 -
Hugging Face Daily Papers research 16d ago
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos
Abstract UniSwap enables synchronized appearance and voice replacement in talking videos through a unified streaming audio-visual diffusion transformer with specialized training and inference adaptations. Generated by thinkingmachines/Inkling-Small Talking-video character…
31 -
Hugging Face Daily Papers research 16d ago
LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time
Abstract LiveAnimate enables real-time, long-form pose-driven human animation via a 14B-parameter video diffusion transformer with specialized training, bounded attention caching, and sequence parallelism. Generated by thinkingmachines/Inkling-Small Pose-driven human animation…
5 -
Hugging Face Daily Papers research 16d ago
SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback
Abstract SkillEvo improves agent skills through multi-turn feedback and active governance to sustain evolution gradients. Generated by thinkingmachines/Inkling-Small Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess…
36 -
Hugging Face Daily Papers research 16d ago
CW-BASS v2: Saturation-Aware Pseudo-Label Selection for Semi-Supervised Segmentation under Foundation-Model Teachers
Abstract CW-BASS v2 selects pseudo-labels by measuring teacher reliability on held-out data and applying either strict filtering or an adaptive floor to avoid confirmation bias under saturated confidence. Generated by thinkingmachines/Inkling-Small Semi-supervised semantic…
4 -
Hugging Face Daily Papers research 16d ago
H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
Abstract H2R-Bench evaluates video generation models on transforming human manipulation videos into robot-centric demonstrations across embodiment constraints and interaction fidelity. Generated by thinkingmachines/Inkling-Small Large-scale manipulation data is essential for…
22 -
Hugging Face Daily Papers research 16d ago
Alaya-EVOKE: From Linear-Scaling Supervision to Endless World
Abstract Evoke is an interactive world model that uses external persistent memory and a redesigned long-horizon teacher to enable responsive, open-ended video generation with bounded context and low latency. Generated by thinkingmachines/Inkling-Small Interactive world models…
34 -
Hugging Face Daily Papers research 16d ago
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers
Abstract LLM routing is formalized as a sequential decision process with a unified benchmark and modular infrastructure to compare and improve cost-effective model selection. Generated by thinkingmachines/Inkling-Small No single large language model (LLM) is optimal across all…
6 -
Hugging Face Daily Papers research 16d ago
DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
Abstract DreamX-Phi 1.0 is an action-conditioned video world model for robotic manipulation that uses geometric attention encoding, depth estimation, object masks with a frozen teacher, and distillation to generate faithful future observations. Generated by…
23 -
Hugging Face Daily Papers research 16d ago
Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence
Abstract A frozen vision-language model improves spatial reasoning by self-evolving through verified experience, reflection, and reusable memory retrieval without parameter updates or external tools. Generated by thinkingmachines/Inkling-Small Spatial intelligence is becoming a…
30 -
Hugging Face Daily Papers research 16d ago
DarwinX: Evolving Agent Harnesses Through Natural Selection
Abstract DarwinX evolves agent harnesses via population selection with frozen models, improving verified performance across benchmarks without benchmark-specific patches. Generated by thinkingmachines/Inkling-Small An LLM agent's capability depends not only on model weights but…
30 -
Hugging Face Daily Papers research 16d ago
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus
Abstract Massive activations in hybrid linear-attention LLMs exhibit pre-attention spikes and inter-spike plateaus governed by cancellation timing, with morphology recovering at full-attention limits. Generated by thinkingmachines/Inkling-Small We present the first systematic…
22 -
Hugging Face Daily Papers research 16d ago
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design
Abstract AutoDesign uses a meta-harness optimizer to recursively improve a code agent for structured media generation, achieving state-of-the-art results on paper-to-poster synthesis. Generated by thinkingmachines/Inkling-Small Transforming multimodal sources into condensed and…
11 -
Hugging Face Daily Papers research 16d ago
LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation
Abstract LycheeMemory V2 improves long-term agent memory by batching interactions into semantic segments for efficient consolidation and structured retrieval, reducing construction costs while maintaining high accuracy. Generated by thinkingmachines/Inkling-Small Long-horizon…
17 -
Hugging Face Daily Papers research 16d ago
PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives
Abstract PlayWorld benchmarks interactive video world models by using multi-modal agents to pursue long-horizon objectives, evaluating geometry consistency, interaction fidelity, and state evolution. Generated by thinkingmachines/Inkling-Small Video world models simulate future…
23 -
Hugging Face Daily Papers research 16d ago
How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review
Abstract Rhetorical framing significantly biases AI scientific review scores in structured ways, with effects shaped by reviewer identity, score range, and evaluation strictness rather than rewriting complexity. Generated by thinkingmachines/Inkling-Small As large language…
4 -
Hugging Face Daily Papers research 16d ago
Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning
Abstract Latent Dynamics Reasoning integrates kinematic dynamics in structured latent space to enable video world models that extrapolate physical laws far beyond training distributions with far fewer parameters and faster inference. Generated by thinkingmachines/Inkling-Small…
35 -
Hugging Face Daily Papers research 16d ago
Full-bandwidth transformer
Abstract Full-bandwidth transformers use latent feedback of top-layer hidden states to improve reasoning and efficiency without altering the core architecture. Generated by thinkingmachines/Inkling-Small Autoregressive transformers compute along two axes: horizontally across…
22 -
Hugging Face Daily Papers research 17d ago
StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization
Abstract StateFlow introduces a persistent 3D world state to enable iterative, controllable previsualization for film and game design by constructing, evolving, and accessing structured scene and camera representations. Generated by thinkingmachines/Inkling-Small…
31 -
Hugging Face Daily Papers research 17d ago
AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research
Abstract The benchmark evaluates autonomous coding agents on open-ended world-model research by having them iteratively improve a starter model across game environments using a shared structured-state format. Generated by thinkingmachines/Inkling-Small World modeling is an…
20 -
Hugging Face Daily Papers research 17d ago
AVA-Encoder: Towards Agent-Native Video Representation Learning
Abstract AVA-Encoder learns structured video representations via agentic auto-encoding using knowledge graphs to enable cinematic video generation and reasoning with reduced token usage. Generated by thinkingmachines/Inkling-Small Creative agents still lack an effective way to…
30 -
Hugging Face Daily Papers research 17d ago
Parameter Exploration for RLVR via Variational Learning
Abstract Parameter-space exploration via perturbed policy sampling improves LLM reinforcement learning by diversifying rollouts and reducing training failures compared to action-space methods. Generated by thinkingmachines/Inkling-Small Exploration has been a focus of…
20