News / #video-gen Tag Video Gen 140 articles archived under #video-gen · RSS Sign in to follow r/MachineLearning community 1d ago WTF is a World Model? [D] I'm trying to understand what a world model is. I understand it has cognitive science and reinforcement learning. I understand at least at the moment what most people are building which they call world models are fancy video generation models. But what actually counts. Does a… 25 The Information — AI news-outlet 3d ago SoftBank Explores Buying Majority Stake in 1X Humanoid Maker SoftBank is in talks to buy a majority stake in 1X Technologies, an OpenAI-backed humanoid robot developer, The Information reported late Wednesday . The investment would support SoftBank’s robotics ambition and give 1X more runway to put its soft-bodied bots in customers’… 20 Hugging Face Daily Papers research 3d ago FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling Abstract FIRM-Video uses checklist-driven verification of temporal visual evidence to build reliable reward models for text-to-video evaluation and alignment. Generated by thinkingmachines/Inkling-Small Reliable reward models are essential for text-to-video evaluation and… 32 Hugging Face Daily Papers research 3d ago Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models Abstract Stream4D improves autoregressive video generation by replacing static 3D critics with a dynamic 4D reconstruction reward and motion prior to preserve coherent motion and reduce geometric drift. Generated by thinkingmachines/Inkling-Small Streaming autoregressive… 38 Hugging Face Daily Papers research 3d ago VGI-BENCH: Probing Visual Intelligence in Video Generation Models Abstract VGI-bench evaluates visual reasoning in video generation models through 27 tasks, revealing limited reliability and minimal self-correction during generation. Generated by thinkingmachines/Inkling-Small Recent studies suggest that video generation models can exhibit… 25 Hacker News — AI on Front Page community 4d ago EVE Online moves to Python 3 Article URL: https://www.eveonline.com/news/view/the-move-to-python-3-begins Comments URL: https://news.ycombinator.com/item?id=49433328 Points: 232 # Comments: 116 11 Vercel — AI dev-tools 5d ago AI Gateway now supports asynchronous video generation Video generation on AI Gateway can now run asynchronously. By default, generateVideo keeps one HTTP request to AI Gateway open until the result is ready. Because video generation can take seconds or minutes, that request can exceed request timeouts. With asynchronous generation,… 27 Hugging Face Daily Papers research 5d ago Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models Abstract SparsePR accelerates video transformers via response-coupled partitioning and probe-fitted residual reconstruction, reducing attention error at low executed-pair densities with substantial speedups. Generated by thinkingmachines/Inkling-Small Training-free block-sparse… 33 Hugging Face Daily Papers research 10d ago SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Abstract Semantic task completion video generation evaluates whether generated videos achieve intended outcomes with semantic grounding, supported by a curated dataset and vision-language model-based benchmark. Generated by thinkingmachines/Inkling-Small We introduce Semantic… 22 Hugging Face Daily Papers research 11d ago V-RAE: Rethinking Video Latent Spaces for Generation Abstract V-RAE constructs semantically organized video latents from frozen vision representations to improve generation quality, convergence speed, and predictive modeling. Generated by thinkingmachines/Inkling-Small Latent video generation relies on autoencoders to define a… 32 Hugging Face Daily Papers research 12d ago AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model Abstract AnyTalk generates 3D speech animations for arbitrary characters without animation data by adapting video diffusion models via character-specific fine-tuning and optimizing blendshape parameters from synthesized talking-head videos, with a distilled real-time variant.… 9 llama.cpp releases dev-tools 12d ago b10472 cuda : skip UMA override for HIP builds ( #27083 ) AMD APUs report accurate memory via hipMemGetInfo. Using MemAvailable over-promises on small-carveout systems. fixes #18159 Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI… 28 OpenAI Python SDK releases dev-tools 15d ago v3.1.0 3.1.0 (2026-08-14) Features api: add WebSocket stream IDs ( #3612 ) ( d9029e3 ) api: add workload identity access token issued event ( #3601 ) ( df274d4 ) api: deprecate Sora video APIs ( #3610 ) ( 721cb1c ) api: Ultrafast tier, structured MCP and websocket errors, separate… 19 Hugging Face Daily Papers research 15d ago Context-Matched Distillation: Teacher Causality for Autoregressive Video Distillation Abstract Context-Matched Distillation aligns teacher supervision with causal generation context for few-step autoregressive video models, improving control adherence and long-video quality. Generated by thinkingmachines/Inkling-Small Interactive autoregressive video generation… 12 Hugging Face Daily Papers research 16d ago H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models Abstract H2R-Bench evaluates video generation models on transforming human manipulation videos into robot-centric demonstrations across embodiment constraints and interaction fidelity. Generated by thinkingmachines/Inkling-Small Large-scale manipulation data is essential for… 22 Hugging Face Daily Papers research 16d ago Alaya-EVOKE: From Linear-Scaling Supervision to Endless World Abstract Evoke is an interactive world model that uses external persistent memory and a redesigned long-horizon teacher to enable responsive, open-ended video generation with bounded context and low latency. Generated by thinkingmachines/Inkling-Small Interactive world models… 34 Hugging Face Daily Papers research 16d ago AVA-Encoder: Towards Agent-Native Video Representation Learning Abstract AVA-Encoder learns structured video representations via agentic auto-encoding using knowledge graphs to enable cinematic video generation and reasoning with reduced token usage. Generated by thinkingmachines/Inkling-Small Creative agents still lack an effective way to… 30 Hugging Face Daily Papers research 19d ago Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains Abstract Sci-VBench evaluates video generation requiring scientific reasoning across disciplines, revealing that visual realism advances have not ensured accurate scientific and causal dynamics. Generated by thinkingmachines/Inkling-Small We introduce Sci-VBench, a comprehensive… 19 r/LocalLLaMA community 20d ago MiniMax H3: A New Open-Weight Video Model, Live in ComfyUI MiniMax H3 is an open-weight, general-purpose multimodal video generation model that works across text, images, video, and audio. In ComfyUI, you can use H3 for text-to-video, image-to-video, first- and last-frame generation, and reference-driven creation. H3 jointly generates… 10 Hugging Face Daily Papers research 20d ago SimWAM: A Simple World Action Model for End-to-End Autonomous Driving Abstract World-Action Models (WAMs) improve end-to-end autonomous driving by transferring video dynamics priors to action prediction, but existing methods require costly future generation at inference. We present SimWAM, a simple yet effective WAM that uses video generation… 19 r/LocalLLaMA community 20d ago Open-weight video gen that actually delivers. Five days with MiniMax H3 on local hardware. H3 weights went live on HuggingFace August 3rd and I started pulling them immediately. An omni-modal video model with native stereo audio in the same forward pass, where audio can actually drive the video generation? On open weights? I had to try it. Five days in, the quality is… 31 Hugging Face Daily Papers research 25d ago MiniWorld: Democratizing the Training of Video World Models from Scratch Abstract Video world models predict future observations conditioned on historical observations and control signals, enabling long-horizon generation through autoregressive state transitions. Unlike conventional video generation models that primarily capture visual appearance and… 18 r/LocalLLaMA community 25d ago [Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding] First of all, obviously I took some help from AI to type this post and this is the topic that enabled me to accomplish all that: https://old.reddit.com/r/LocalLLaMA/comments/1veow4b/deepseek_v4flash_284b_moe_at_33_toks_single_68/ This post of mine is based on the link above. My… 11 Hugging Face Daily Papers research 26d ago WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity Abstract Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the apparent appearance of generated videos to the inherent reactivity of the worlds they depict: the ability to infer from… 6 Hugging Face Daily Papers research 1mo ago VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System Abstract Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt. Existing chain-of-thought… 11 Vercel — AI dev-tools 1mo ago MiniMax H3 now available on AI Gateway MiniMax H3 is now available on AI Gateway. H3 generates 2K video from a text prompt, a starting image, a pair of first and last frames, or reference material. Alongside text-to-video and first-frame image-to-video, the model supports first-to-last keyframe transitions and… 25 Hugging Face Daily Papers research 1mo ago Parallel Decoding Distillation for Fast Image and Video Generation Abstract Generation in video diffusion or flow models is computationally expensive due to the slow and iterative sampling process. Current state-of-the-art (SOTA) acceleration methods heavily rely on variational score distillation (VSD) and adversarial losses to distill… 31 Hugging Face Daily Papers research 1mo ago FilmBench: A Film-Grade Benchmark for Cinematic Video Generation Abstract Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally,… 17 Hugging Face Daily Papers research 1mo ago Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification Abstract Diffusion transformers are essential for high-fidelity video generation, but long token sequences make attention a dominant inference bottleneck. Training-free dynamic sparse attention alleviates this bottleneck by computing only selected key-value blocks, yet existing… 19 Hugging Face Daily Papers research 1mo ago Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative Rendering Abstract Recent conditional video generation models have shown promising potentials to transform 3D engine renderings, such as depth maps and untextured geometry, into photorealistic videos for gaming and immersive content creation. These applications require long-horizon… 38 TechCrunch — AI news-outlet 1mo ago Midjourney acquired the astrology app Co-Star The AI lab Midjourney continues to expand its purview beyond image and video generation. 11 Hugging Face Daily Papers research 1mo ago SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Abstract We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p on a single GPU, SANA-Video 2.0 matches full-softmax video DiTs in quality while… 19 Hugging Face Daily Papers research 1mo ago GraphVid: Interactive Graph-Controllable Video Generation Abstract Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, trajectory-based control often requires users to… 12 TechCrunch — AI news-outlet 1mo ago Runway launches AI model router as generative media gets crowded Runway no longer wants to be just another AI model company. It wants to become the infrastructure layer for generative media. On Thursday, the startup launched Runway Media Router through Runway Dev, its developer platform, released earlier this month, that provides API access… 25 Hugging Face Daily Papers research 1mo ago Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation Abstract Text-to-video generation has advanced significantly over the past five years through scaling of model size, data, and compute. Unlike model architecture, training data is often underexplored. Real-world data curation is complex and non-trivial, involving clip selection… 25 Hugging Face Daily Papers research 1mo ago FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation Abstract Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck in high-resolution video generation. Training-free sparse attention reduces this cost, but adaptive Top-p routing creates uneven per-head workloads under… 16 Hugging Face Daily Papers research 1mo ago Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence Abstract Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without verifying whether the model arrives there through… 16 arXiv — NLP / Computation & Language research 1mo ago Thinking in Video: Can Video Generators Really Reason About the Real World? arXiv:2607.17523v1 Announce Type: cross Abstract: Recent advances in world models and video generation have given rise to an emerging reasoning paradigm that leverages video generative models to simulate, predict, and reason about real-world dynamics. We redefine this paradigm… 25 Hugging Face Daily Papers research 1mo ago HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement Abstract Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most approaches focusing on inter-subject personalization still struggle to strike a balance… 9 Hugging Face Daily Papers research 1mo ago FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications Abstract Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose efficient deployment requires application-specific decisions about placement, streaming, and intra-model parallelism. Existing… 10 arXiv — Machine Learning research 1mo ago Inference-Time Concept Suppression and Video-Centric Evaluation for Text-to-Video Models arXiv:2607.14194v1 Announce Type: cross Abstract: Text-to-video (T2V) generators can synthesize realistic and temporally coherent videos, but controllably removing a target concept from a generator remains difficult. Unlike text-to-image concept erasure, T2V unlearning must… 5 Hugging Face Daily Papers research 1mo ago MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation Abstract Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mainly focus on text-driven generation, single-reference subject preservation, or isolated… 11 Hugging Face Daily Papers research 1mo ago KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation Abstract Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference images to guide generation. Although recent models support multi-keyframe conditioning, it remains unclear whether they can faithfully reproduce the… 14 arXiv — Machine Learning research 1mo ago Text2Sign: A Single-GPU Diffusion Baseline for Text-to-Sign Language Video Generation arXiv:2607.13164v1 Announce Type: cross Abstract: Sign language is a primary communication channel for millions of Deaf and hard-of-hearing people, yet text-to-signer video generation remains costly because video diffusion models are expensive to train and evaluate. This paper… 11 Hugging Face Daily Papers research 1mo ago Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Abstract Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints. Existing… 14 arXiv — Machine Learning research 1mo ago Empowering Long-form Omni-modal Understanding with Robust Audio Perception arXiv:2607.10299v1 Announce Type: new Abstract: Recent advances in large-scale multimodal models have drivenremarkable progress in vision-language tasks; however, comprehensiveomni-modal understanding remains under-explored, largely due to thescarcity of datasets with rich,… 8 TechCrunch — AI news-outlet 1mo ago Video generation startup PixVerse raises $439M, valuation soars past $2B Singapore-based video generation startup PixVerse closed a Series C extension on the strength of 15 million monthly active users, it said. 14 Hugging Face Daily Papers research 1mo ago Video Generation Models are General-Purpose Vision Learners Abstract Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In this paper, we contend that large-scale… 28 arXiv — Machine Learning research 1mo ago GenVid2Robot: From Video Generation to Robot Manipulation via Rigid-Geometric Consistency arXiv:2607.09191v1 Announce Type: cross Abstract: Generated videos provide useful visual motion priors for robot manipulation, but their visual plausibility does not imply physical executability. A generated video usually lacks metric geometry, grasp grounding, robot kinematic… 29 Hugging Face Daily Papers research 1mo ago OpenCoF: Learning to Reason Through Video Generation Abstract OpenCoF framework introduces a reasoning video dataset and model that improve temporal reasoning through diverse supervision and explicit reasoning tokens for visual and textual cues. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Reasoning has become a core capability… 38 Page 1 of 3 · 140 articles Older →