AllenAI has been iterating on their MolmoAct2 models for robotics
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
r/AllenAI is cooking with MolmoAct2, a 5B vision-language-action model for robot control. They keep releasing new fine-tunes on different kinds of robotics datasets, including (but not limited to, and they keep releasing new ones):
https://huggingface.co/allenai/MolmoAct2-LIBERO - general robotics tasks
https://huggingface.co/allenai/MolmoAct2-DROID - interactive robotics tasks
https://huggingface.co/allenai/MolmoAct2-BimanualYAM - absolute joint-pose control
https://huggingface.co/allenai/MolmoAct2-SO100_101 - also absolute joint-pose control
AllenAI has released these as fully open source models, publishing not only their weights but also their complete training datasets (including pretraining), their training software source code, and technical papers describing the theory, training, and assessments of these models.
If anyone is fiddling with robots controlled via LLM inference, you should give MolmoAct2 models a look.
[link] [comments]
More from r/LocalLLaMA
-
How bad do you think models like Qwen3.8-27B or GLM-5.3-Flash would be with H-Neurons disabled?
Aug 31
-
vote for the Qwen 3.8
Aug 31
-
CUDA: extend MOE fusion to specdec, earlier MOE glu fusion and topk-router fusion were restricted to 1 token by ynankani · Pull Request #27621 · ggml-org/llama.cpp
Aug 31
-
deepseek-ai/DeepSeek-V4-Flash-Vision-Exp · Hugging Face
Aug 31
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.