Zhi Jing, Jinbin Qiao, Ouyang Lu, Jicong Ao, Shuang Qiu, Yu-Gang Jiang, Chenjia Bai
AssemLM is a multimodal large language model enabling 3D spatial reasoning for robotic assembly, predicting 6D poses using assembly manuals and point clouds.
Existing vision-language models rely on 2D perception, lacking 3D geometric reasoning and struggling with 6D pose inference required for precise tasks like robotic assembly.
AssemLM extracts fine-grained geometric and rotational features via a point cloud encoder and integrates them into a multimodal language model for 3D spatial reasoning. It also constructs the AssemBench dataset with over 900K samples for training and evaluation.
Achieves state-of-the-art performance in 6D pose reasoning across diverse assembly scenarios, and successfully performs precise multi-step assembly in real-robot experiments, demonstrating potential for robotic assembly applications.