News / #robotics Tag Robotics 419 articles archived under #robotics · RSS Sign in to follow arXiv — Machine Learning research 1mo ago MUGEN: A Unified Framework for Efficient Motion Understanding and Generation arXiv:2607.27581v1 Announce Type: new Abstract: Grounding human motion in language, and language in motion, is a central step toward physical AI systems that can understand, generate, and communicate human behavior. Unified motion--language systems first coupled the two… 4 Hugging Face Daily Papers research 1mo ago ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine Abstract Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this… 24 MIT News — AI research 1mo ago Daniela Rus receives Bavarian Minister-President's High-Tech Prize Director of CSAIL and MIT professor honored for her contributions to robotics, artificial intelligence, and autonomous systems. 21 Ars Technica — AI news-outlet 1mo ago Google reveals Gemini Robotics 2.0, promising improved dexterity and safety Gemini Robotics 2 includes three models, but only one is publicly available right now. 22 r/LocalLLaMA community 1mo ago How close are we to local llama robotics for consumer price point? I'm guessing 3 years, what do you think? In other words: many of us will be able to afford a general purpose robot in 3 years to experiment with in the home. Cost roughly $5k? Probably small size, but hopefully still able to do the dishes and operate a vacuum.   submitted by… 13 Hugging Face Daily Papers research 1mo ago πR^2: Reactive Real-time Flow Policies Abstract Generalist manipulation policies increasingly take the form of action-chunking flow policies built on large pretrained backbones. Such chunks run open-loop, so the policy cannot react to sensory input arriving mid-execution, sacrificing reactivity. Replanning more often… 4 Google DeepMind official-blog 1mo ago Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications. 5 arXiv — Machine Learning research 1mo ago Conformal Changepoint Localization and Root Cause Analysis with Corrupted Observations arXiv:2607.26481v1 Announce Type: new Abstract: Detecting when the statistical behavior of an engineered system changes, and identifying which component is responsible, are core problems in the monitoring of telecommunication networks, robotic platforms, security infrastructure,… 10 arXiv — NLP / Computation & Language research 1mo ago The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness arXiv:2503.10647v2 Announce Type: replace Abstract: This study evaluated the diagnostic reliability of two Large Language Models (LLMs), Google Gemini 2.0 Flash and OpenAI ChatGPT-4o, across three dimensions: consistency under rephrased inputs, susceptibility to irrelevant… 10 Hugging Face Daily Papers research 1mo ago TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Abstract Vision-language-action (VLA) models commonly adopt an LLM-centric V to L to A pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Although effective, this design incurs… 36 Hugging Face Daily Papers research 1mo ago Explicit Layer Modeling for Video Object Insertion and Layer Decomposition Abstract Most video editing systems still lack explicit layered video representations, limiting their ability to perform realistic compositing, object reuse, and consistent manipulation. This limitation is especially pronounced in video object insertion and video layer… 26 Ars Technica — AI news-outlet 1mo ago Who wins and who loses after US bans foreign robots? Government ban on foreign-made robots may hinder instead of help US robotics. 18 Hugging Face Daily Papers research 1mo ago HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone Abstract Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scalable. Real-robot teleoperation is accurate but costly to scale; robot-free UMI capture scales readily, and current practice uses the resulting data… 13 r/LocalLLaMA community 1mo ago I got Kimi-k3 running..... Results: prompt eval: 40 tokens / 97.5s → 0.41 tok/s eval: 400 tokens / 1769.9s → 0.23 tok/s total: 440 tokens / 1867s (31 min) Prompt: "Write a C++ function that reverses a linked list in place. Explain the pointer manipulation." How I ran it: Using PR#26185 from llama.cpp… 32 NVIDIA Developer Blog official-blog 1mo ago Developing Healthcare Robotics with GPU-Native Medical Physics Simulation Unlike autonomous driving or industrial robotics, healthcare robotics can’t rely on internet-scale data collection or unlimited real-world experimentation.... 28 r/MachineLearning community 1mo ago NeurIPS-side prompt injection triggering ethics reviewers? [D] Does anyone experience a similar story that some reviewers reporting ethical issue due to NeurIPS-side prompt injection for catching LLM-reviewers? Even ethics reviewers were not informed about this conference-side manipulation…   submitted by   /u/dontknowwhattoplay… 19 Hugging Face Daily Papers research 1mo ago WorldDiT: A Unified Diffusion Architecture for World and Action Modeling Abstract Many recent robot policies pursue stronger control by using large pretrained vision-language models (VLMs) as the action backbone. We introduce WorldDiT, a unified diffusion transformer architecture that couples action generation with visual world modeling and achieves… 35 r/MachineLearning community 1mo ago I built a deep learning library from scratch in C that lets you train language models [P] my goal was to train a Language model (SLM) entirely from scratch so no ML libraries allowed . so i gathered what's needed to make it happen : from tensor manipulation (views, operations , allocations) the autograd ( a DAG that retains the previous operations and inputs in order… 7 Google DeepMind official-blog 1mo ago Gemini Robotics 2 brings whole body intelligence to robots July 30, 2026 Models Gemini Robotics 2 brings whole body intelligence to robots Carolina Parada Share From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks For decades, we’ve… 20 Hugging Face Daily Papers research 1mo ago Data Pyramid for Embodied Manipulation Abstract Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states and actions. These signals can be provided, to varying degrees, by… 18 Hugging Face Daily Papers research 1mo ago Progress Reward Modeling for Robotic Learning: A Comprehensive Survey Abstract Robotic learning takes place in dynamic environments with large behavior spaces. A terminal success signal only tells the robot whether the task is completed. It does not explain whether the current behavior is making progress, remaining unchanged, or undoing earlier… 31 Import AI (Jack Clark) community 1mo ago Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI's accidental AI hacker The warning shots will continue until civilization wakes up 16 TechCrunch — AI news-outlet 1mo ago Enigma raises $70M to make controlling a robot as easy as adjusting the volume The massive seed round was led by Index Ventures and Ribbit Capital, with participation from Sarah Guo's Conviction Partners. 32 Hugging Face official-blog 1mo ago NVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics Back to Articles a]:hidden"> NVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics Enterprise + Article Published July 27, 2026 Upvote 1 Lukas Zbinden lzbinden nvidia Javier Gamazo javirk1 nvidia Mostafa Toloui mtoloui nvidia Sean Huver shuver… 12 arXiv — Machine Learning research 1mo ago Ordered Action Tokens for Visuomotor Policy Learning arXiv:2607.21670v1 Announce Type: cross Abstract: Action tokenization maps continuous robot action chunks to discrete tokens and has become an important interface for modern visuomotor policies. Existing approaches either rely on analytical discretization methods that produce… 11 arXiv — NLP / Computation & Language research 1mo ago Progress Reward Modeling for Robotic Learning: A Comprehensive Survey arXiv:2607.21655v1 Announce Type: cross Abstract: Robotic learning takes place in dynamic environments with large behavior spaces. A terminal success signal only tells the robot whether the task is completed. It does not explain whether the current behavior is making progress,… 8 r/LocalLLaMA community 1mo ago Unexpected use of local llm I was refreshing my youtube and found out my favourite reviewer uploaded a battery test of 78 smartphones: https://youtu.be/MpgUFrsIWSQ the author said they started using robotic arm to simulate a person using the phone but they wanted to further enhance it by using agentic ai.… 19 TechCrunch — AI news-outlet 1mo ago Are brain waves the next unlock for physical AI? Forget YouTube videos—frontier physical AI models need multiple camera angles, dense annotation, and soon, brain wave readings. 28 Hacker News — AI on Front Page community 1mo ago London Gatwick has launched a robotic airport parking service Article URL: https://aerospaceglobalnews.com/news/gatwick-airport-robotic-parking-stanley-robotics/ Comments URL: https://news.ycombinator.com/item?id=49058669 Points: 211 # Comments: 146 33 r/MachineLearning community 1mo ago Why first person video may matter for robot learning[D] I can see why first-person video might help a robot model, but not because the robot can copy a human hand. The joints, reach, timing, and control space are all different. What may transfer is the sequence of visual attention: which object enters view, what changes before… 20 Hugging Face Daily Papers research 1mo ago Robostral Navigate Abstract Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently. Yet, today's best systems depend on depth sensors, multi-camera rigs, or pre-built maps, limiting the hardware they… 33 Hugging Face Daily Papers research 1mo ago TableVerse: A Large-scale Tabletop Dataset with Real-world Grounded Layouts for Generalizable Manipulation Abstract The development of generalizable robotic manipulation policies is inherently bounded by the availability of large-scale, high-fidelity scene data. While recent automated synthesis methods attempt to bridge this gap via text-to-layout hallucination or simplified… 5 Latent.Space news-outlet 1mo ago [AINews] Black Forest Labs FLUX 3 - Multimodal Flow Models that beat Seedance 2.0, Gemini Omni and Grok Imagine, and FLUX-mimic video-action robotics model A HUGE win for BFL! 22 Hugging Face Daily Papers research 1mo ago SLAM in Low-Light Environments: Project Report Abstract Simultaneous localization and mapping (SLAM) is one of the fundamental problems in robotics, as it enables autonomous operations in real-world scenarios. Under low illumination, reduced contrast, sensor noise, and motion blur degrade both feature extraction and feature… 15 Hugging Face Daily Papers research 1mo ago SeededGrasp: Language-Guided Grasping in Complex Scenes with Multiple Embodiments Abstract Practical robotic grasping in complex scenes requires both 3D spatial reasoning and alignment with task-specific requirements. Vision-language models (VLMs) offer a natural way to specify these requirements using language, but existing approaches either use a VLM to… 37 arXiv — Machine Learning research 1mo ago Towards Torque-Driven Reinforcement Learning for Quadruped Locomotion arXiv:2607.18365v1 Announce Type: cross Abstract: Reinforcement learning (RL) for legged robots is advancing locomotion, demonstrating its ability to adapt to new and challenging terrain. Traditionally, these RL locomotion frameworks are position-based, making the policy less… 25 Hugging Face Daily Papers research 1mo ago Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment Abstract Finetuning a pretrained vision-language model (VLM) on robot demonstrations via behavior cloning (BC) has become the standard recipe for vision-language-action (VLA) policies. However, BC finetuning progressively overwrites the pretrained representations that support… 31 TechCrunch — AI news-outlet 1mo ago Travis Kalanick’s robotics company raises $1.7B, led by a16z Uber is also investing in Travis Kalanick's company Atoms, which has made gauzy claims about using industrial AI to modernize the world. 6 Ars Technica — AI news-outlet 1mo ago Hyundai claims humanoid robot plan is not part of talks with striking workers Union previously warned automaker that any robot deployment must be negotiated. 21 Hugging Face Daily Papers research 1mo ago Masked Visual Actions for Unified World Modeling Abstract Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to communicate action to such models in a form aligned with the visual space in… 27 arXiv — Machine Learning research 1mo ago Deep learning-based prediction of time-resolved adhesive forces in viscoelastic Hertzian contacts arXiv:2607.19060v1 Announce Type: new Abstract: Fast prediction of the response of adhesive soft viscoelastic contacts represents a current challenge in soft robotics and for gripping and manipulation tasks. Determining the complete time-resolved force trajectory requires full… 34 Hugging Face official-blog 1mo ago The State of Simulation for Physical AI: An Overview Back to Articles a]:hidden"> The State of Simulation for Physical AI: An Overview Enterprise + Article Published July 21, 2026 Upvote - Johnny Nuñez Cano johnnynv nvidia Mitesh Patel mitp nvidia Asier Arranz asiernvidia nvidia lior ben horin liorbenhorin-nv nvidia Raymond Lo… 23 TechCrunch — AI news-outlet 1mo ago Gritt exits stealth with $34 million for robots to build solar plants—then, everything else Gritt is coming out of stealth with $34 million and plan to automate the hardest tasks on construction sites. 31 arXiv — NLP / Computation & Language research 1mo ago Cross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image Editing arXiv:2607.16898v1 Announce Type: cross Abstract: Unified multimodal models (UMMs) have recently demonstrated powerful instruction-based image editing capabilities, but they also raise serious concerns about unauthorized manipulation of personal portraits. Existing adversarial… 9 Hugging Face Daily Papers research 1mo ago JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models Abstract The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robot embodiments, and task objectives. Existing compute services, whether offered as direct accelerator rental or batch-workload submission, typically allocate an… 13 Hugging Face official-blog 1mo ago Grabette: an open system to record robot-manipulation data Back to Articles a]:hidden"> Grabette: an open system to record robot-manipulation data. And build a shared dataset, together. Published July 21, 2026 Update on GitHub Upvote 5 Steve Nguyen SteveNguyen pollen-robotics Claire Houziel chouziel pollen-robotics Gaelle Lannuzel… 17 Hugging Face official-blog 1mo ago Introducing Cosmos 3 Edge Back to Articles a]:hidden"> Introducing Cosmos 3 Edge Enterprise + Article Published July 20, 2026 Upvote - Pranjali Joshi PranjaliJoshi nvidia Saeed Babamohamadi SaeedBabamohamadi nvidia The real world is vast and to operate in it physical AI systems need to understand how a… 32 NVIDIA Developer Blog official-blog 1mo ago Integrate NVIDIA Omniverse RTX Sensor Simulation Into Existing Apps Developers building 3D, design, simulation, robotics, and industrial digital twin applications need ways to bring physical AI capabilities into the tools and... 12 Hugging Face Daily Papers research 1mo ago See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models Abstract Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe the scene in the camera frame, creating a frame mismatch between where… 16 r/LocalLLaMA community 1mo ago MiniCPM-Robot model series - MiniCPM-RobotManip & MiniCPM-RobotTrack 🚀 MiniCPM enters the physical world — enabling robots to understand, remember, and act. We open-source MiniCPM-Robot, our first embodied AI model series, including: 🤖 MiniCPM-RobotManip — a 1.5B general-purpose Vision-Language-Action (VLA) model for robotic manipulation. 🐕… 23 Page 3 of 9 · 419 articles ← Newer Older →