News / #robotics Tag Robotics 419 articles archived under #robotics · RSS Sign in to follow Hacker News — AI on Front Page community 1mo ago Xiaomi-Robotics-1 Article URL: https://robotics.xiaomi.com/xiaomi-robotics-1.html Comments URL: https://news.ycombinator.com/item?id=48974454 Points: 215 # Comments: 150 9 Hugging Face Daily Papers research 1mo ago Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Abstract We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) efficiently adapting to novel… 21 r/LocalLLaMA community 1mo ago How long before Chinese models fully surpass US models? Given the rate at which they have been advancing, I predict we are six months away from a leapfrog moment. EDIT - For those responding “never - they just copy everything”, how is that working out for EV and robotics? The idea that China is still some backwater knock-off empire… 14 r/LocalLLaMA community 1mo ago SigLIP 2 text embedding on CPU with Rust + ONNX We’re building a robotics data platform with a lot of images, video, and text metadata. For search, we use SigLIP 2. GPUs handle batched asynchronous image/video embedding and indexing, while this small Rust + ONNX Runtime service handles live text queries on CPU. Both land in… 37 TechCrunch — AI news-outlet 1mo ago Agility Robotics plants its flag in Tesla’s backyard Agility is opening a new training center for its Digit robots in Fremont, California. 29 TechCrunch — AI news-outlet 1mo ago Patreon stops asking AI bots not to scrape — and starts blocking them Patreon is strengthening its defenses against AI scraping by working with Cloudflare to block bots that train AI models on creators’ content without permission. The move marks a shift away from relying on websites using robots.txt alone to actively block unauthorized AI training. 36 Hugging Face Daily Papers research 1mo ago SUFLECA: Scaling Up Feature Learning for CAD-to-image Alignment Abstract CAD-to-image alignment aims to estimate an object's 9D pose (rotation, translation, and anisotropic scale) from a single RGB image, enabling applications in robotics and augmented reality. Recent zero-shot methods use visual foundation models to match image regions to… 22 Hugging Face Daily Papers research 1mo ago RoboTTT: Context Scaling for Robot Policies Abstract Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond… 29 arXiv — Machine Learning research 1mo ago Active Real-World Factor-Based Evaluation for Generalist Robot Policies arXiv:2607.14439v1 Announce Type: new Abstract: Generalist robot manipulation policies trained on large, diverse datasets have shown remarkable promise across a wide range of tasks. However, rigorously evaluating these policies remains a fundamental challenge. Real-world… 18 arXiv — Machine Learning research 1mo ago Never Too Late for Force: Accelerating VLA Post-Training with Reactive Force Injection arXiv:2607.14236v1 Announce Type: cross Abstract: Pretrained vision-language-action (VLA) policies provide strong language-conditioned manipulation knowledge, but they remain largely vision-driven and can struggle once manipulation enters contact states where the scene is… 27 arXiv — Machine Learning research 1mo ago DiMaS: Distribution Matching for Steering Vision-Language-Action Models arXiv:2607.14280v1 Announce Type: cross Abstract: Flow-matching-based vision-language-action (VLA) models have emerged as powerful policies for robotic manipulation, yet a critical capability remains underexplored: fine-grained behavioral control, the ability to govern how a… 22 arXiv — NLP / Computation & Language research 1mo ago MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning arXiv:2607.14252v1 Announce Type: cross Abstract: Long-horizon robot planning requires more than predicting what actions will do next; it also requires memory of the embodied experience that makes future goals interpretable. People do not plan from the present scene alone: they… 36 Ars Technica — AI news-outlet 1mo ago Fear of humanoid robots spurs human workers to strike at Hyundai auto factory Hyundai aims to deploy 25,000 Atlas robots starting with US factories in 2028. 11 Hugging Face Daily Papers research 1mo ago SPEAR: A Simulator for Photorealistic Embodied AI Research Abstract Interactive simulators have become powerful tools for training embodied agents and generating synthetic visual data, but existing photorealistic simulators suffer from limited generality, programmability, and rendering speed. We address these limitations by introducing… 14 Latent.Space news-outlet 1mo ago 🔬 The Lab of the Future Should Feel Like a Data Center — Andy Beam & Rafa Gómez-Bombarelli, Lila Sciences Lila is betting that science, not the internet, is the last untapped source of training data. We went to find out what that actually looks like in a room full of robots. 30 arXiv — Machine Learning research 1mo ago Where Should RL Post-Training Compute Go? Model Size, Search, Learning, and Feedback arXiv:2607.13389v1 Announce Type: new Abstract: Reinforcement Learning (RL) post-training is increasingly used to adapt foundation models for reasoning, planning, and feedback-driven robot-learning pipelines, but constrained post-training resources are often summarized by a… 17 arXiv — Machine Learning research 1mo ago HRIBench: Benchmarking Interaction-Centric Human-Robot Collaboration arXiv:2607.13056v1 Announce Type: cross Abstract: Current vision-language-action (VLA) benchmarks primarily evaluate isolated manipulation skills while leaving human-robot interaction structure largely unmodeled. However, real-world collaboration fundamentally requires… 12 Hugging Face Daily Papers research 1mo ago GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch Abstract World Action Models (WAMs) improve robot policy learning by jointly modeling actions and future visual observations, using future scene evolution as dense supervision for physically grounded action generation. However, a common design in existing WAMs is to explicitly… 8 NVIDIA Developer Blog official-blog 1mo ago Develop Lightweight USD Runtimes Faster with AI Agents OpenUSD is an open, extensible framework that provides a common scene description language for physical AI. It enables teams to bring CAD data, simulation... 34 r/MachineLearning community 1mo ago Looking for JEPA devil advocates [R] I am currently doing research on world models, specially in tje field of robot learning, and, as probably most of you alredy know, JEPA-like models are mentioned over and over. I read the main recent papers from lecun as well as other research groups, and I personally think the… 12 r/MachineLearning community 1mo ago All major robotics and VLA papers, ranked and benchmarked in a single place [P] Hi folks, There is now a dedicated Robotics page on Papers with Code that lists the major benchmarks, trending papers with linked code, and open-source artifacts. Find it here: https://paperswithcode.co/tasks/robotics… 24 Simon Willison community 1mo ago simonw/pedalican simonw/pedalican Clearly I wasn't paying attention when these were first announced back in May, but today I accidentally activated a "pet" in Codex Desktop - a little animated robot, reminiscent of Clippy - and then learned you can create your own. So I did, and now I have a… 35 Hugging Face Daily Papers research 1mo ago Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Abstract Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints. Existing… 14 Hugging Face Daily Papers research 1mo ago EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos Abstract Steerability is a defining capability of generalist robot policies, yet remains largely absent in dexterous-hand systems for lack of large-scale, language-aligned, and action-accurate demonstration data. To address this bottleneck, we present a full-stack system that… 4 Hugging Face Daily Papers research 1mo ago ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory Abstract Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general runtime layer for reasoning, memory, tool use, verification, and cross-embodiment execution. We present ABot-AgentOS, a general… 14 TechCrunch — AI news-outlet 1mo ago Uber’s product chief on hotels, robotaxis, and why the company doesn’t want to be “everything for everyone” Uber Chief Product Officer Sachin Kansal walks TechCrunch through the company's financial-services ambitions, its increasingly complicated relationship with Waymo, its new AV Labs data operation, and how AI is starting to show up in ways riders and drivers will actually notice. 15 TechCrunch — AI news-outlet 1mo ago Hermes agent maker Nous Research in talks for new funding at $1.5B valuation The company is raising at least $75 million, led by Robot, with significant participation from USV and other prominent investors. 37 MIT News — AI research 1mo ago AI agents create virtual playgrounds to help robots get crucial training data “SceneSmith” system uses collaborative AI agents to create realistic 3D environments of places like kitchens, hotels, and living rooms, where robots can simulate everyday chores. 17 Google DeepMind official-blog 1mo ago Empowering India’s next generation of innovators with ATL Saathi Google and AIM launched ATL Saathi, a Gemini-powered AI tool empowering Indian educators in robotics labs. 15 arXiv — Machine Learning research 1mo ago SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions arXiv:2607.08925v1 Announce Type: new Abstract: Training reinforcement-learning agents directly on physical robots makes every fall costly, since a fall can damage the platform and cannot be undone like a simulator reset; the goal is therefore to minimize falls during training… 8 arXiv — Machine Learning research 1mo ago Learning More from Less: Reinforcement Learning from Hindsight arXiv:2607.09042v1 Announce Type: new Abstract: Reinforcement learning (RL) is increasingly used to post-train vision-language-action (VLA) models, but every update consumes robot rollouts that are slow and costly to collect, making sample efficiency a central concern.… 35 arXiv — Machine Learning research 1mo ago FlowDAgger: Human-in-the-Loop Adaptation of Generative Robot Policies in Latent Space arXiv:2607.08877v1 Announce Type: cross Abstract: Pretrained generative robot policies based on flow matching and diffusion have achieved impressive results across a wide range of manipulation tasks. Yet real-world deployments routinely expose failure modes outside the… 6 arXiv — Machine Learning research 1mo ago GenVid2Robot: From Video Generation to Robot Manipulation via Rigid-Geometric Consistency arXiv:2607.09191v1 Announce Type: cross Abstract: Generated videos provide useful visual motion priors for robot manipulation, but their visual plausibility does not imply physical executability. A generated video usually lacks metric geometry, grasp grounding, robot kinematic… 29 r/MachineLearning community 1mo ago Ph.D. in Operations Research / Big Tech Eng: How to transition into intermediate/advanced ML for high-value industries (Robotics, Defense, Finance)? [D] I hold a Ph.D. in Operations Research, along with a BSc/MSc in Engineering and OR. I previously worked in Big Tech, but I’m currently looking to transition. My primary goal is to upgrade my technical skillset to maximize my industry-related profitability and marketability. I… 32 NVIDIA Developer Blog official-blog 1mo ago How to Evaluate General-Purpose Robot Policies for Real-World Deployment Robotics foundation models have made remarkable progress. Today's best systems can follow natural language instructions to pick, place, sort, and manipulate a... 18 arXiv — Machine Learning research 1mo ago Write-Protected Discrete Bottlenecks for Language-Grounded World Models: A Structural Limitation and Sufficient Fix arXiv:2607.08312v1 Announce Type: new Abstract: How should language interface with a world model's discrete symbol system? The dominant paradigm -- end-to-end injection of LLM/VLM features into robot world models (RT-2, Octo, PaLM-E) -- implicitly assumes that language gradients… 4 Hugging Face Daily Papers research 1mo ago RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures Abstract RoboTALES introduces a two-stage framework that combines LLM-based planning and VLM-based criticism to improve task-aligned video generation and robotic policy training. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Pretrained video generative models are promising… 22 Ars Technica — AI news-outlet 1mo ago Humanoid robots controlled by surgeons did world-first operation on live pigs Preclinical trial is testing the feasibility of humanoid robots in surgery. 17 Hugging Face Daily Papers research 1mo ago OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies Abstract OmniTacTune enables efficient adaptation of tactile feedback to visual robot policies through a two-stage reinforcement learning approach that improves success rates in contact-rich manipulation tasks. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Visual policies learned… 21 MIT News — AI research 1mo ago Tiny robot boats build floating structures MIT researchers developed FloatForm, a swarm of small aquatic robots that snap together like ants forming a raft, assembling into reconfigurable structures on the water. 35 arXiv — Machine Learning research 1mo ago Does Demand Response Increase Vulnerability to Cyber Attacks by Adversarial Data Modifications? arXiv:2607.06632v1 Announce Type: new Abstract: Adversarial attacks are crafted data manipulations that aim to deteriorate the outcomes of prediction or decision-making algorithms. In the energy systems literature, adversarial attacks have been studied with a focus on problems… 4 arXiv — Machine Learning research 1mo ago Safe Reinforcement Learning using Ideas from Model Predictive Control arXiv:2607.07252v1 Announce Type: new Abstract: Reinforcement learning (RL) enables the synthesis of control policies directly from data, making it highly appealing for complex cyber-physical systems (CPSs) and robotics. A persistent challenge, however, is ensuring strict, hard… 16 arXiv — NLP / Computation & Language research 1mo ago Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders arXiv:2607.07294v1 Announce Type: cross Abstract: Turn-taking prediction is a key requirement for social robots involved in human-human interaction, particularly in mediator settings, where the robot must anticipate conversational dynamics rather than merely react to pauses.… 23 Hugging Face Daily Papers research 1mo ago RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies Abstract RoboDojo presents a unified sim-and-real benchmark for evaluating generalist robot manipulation policies across diverse tasks and evaluation dimensions. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Generalist robot manipulation policies have advanced rapidly, yet… 32 Hugging Face Daily Papers research 1mo ago Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Abstract LaMem-VLA introduces a latent-memory-native framework that integrates historical experience into vision-language-action reasoning through coordinated memory components operating in the same latent space. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Mainstream… 30 TechCrunch — AI news-outlet 1mo ago This startup thinks robotics is about to have its ChatGPT moment General Intuition is betting millions of hours of video game data can train the foundation models for physical AI, making it easier to build smarter robots with minimal real-world data. 29 r/MachineLearning community 1mo ago LingBot-Video: sparse-MoE video diffusion transformer (13B total, 1.4B active) post-trained as an action-conditioned world model[R] Single-stream diffusion transformer with a DeepSeek-V3-style sparse MoE (128 experts, top-8 routing, 1.4B active of 13B total). Six-reward RL post-training including a physical-plausibility reward, plus an action-to-video mode that predicts robot rollouts from action and… 25 Hugging Face Daily Papers research 1mo ago RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation Abstract A multi-modal 4D world model generates synchronized RGB, depth, and optical flow data from single RGB-D images and language instructions, enabling efficient robotic manipulation through unified diffusion processes and inverse dynamics policy learning. Generated by… 23 Hugging Face Daily Papers research 1mo ago RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation Abstract Digital teleoperation replaces physical robot interaction with generative world models to create diverse training data for robotics, enabling efficient zero-shot Sim2Real transfer and improved real-world performance. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Scaling… 9 Hacker News — AI on Front Page community 1mo ago Mistral's Robostral Navigate: a state of the art robotics navigation model Article URL: https://mistral.ai/news/robostral-navigate/ Comments URL: https://news.ycombinator.com/item?id=48832212 Points: 230 # Comments: 55 18 Page 4 of 9 · 419 articles ← Newer Older →