상세 보기
MulPlanLM: multimodal robotic task planning with vision-language models and physical feedback
- Son, Young-Chae;
- Lee, Dong-Han;
- Lim, Soo-Chul
WEB OF SCIENCE
0SCOPUS
0초록
This study proposes a multimodal robotic task planning framework based on large language models (LLMs) that utilizes both visual and force-derived physical feedback. The system integrates multimodal inputs, including camera images and force/torque (F/T) sensor data, to interpret the visual and physical properties of objects and generate a task sequence to execute given commands. The integration of vision and force data allows the system to compensate for the limitations of each modality, leading to more reliable decision-making. The framework consists of three components: Extractor, Planner, and Sub-planner. According to their assigned roles, each agent automatically translates high-level commands given in natural language into executable robot motion plans, enabling the robot to perform the required sequence of actions. The system combines visual and force data to handle tasks that are difficult or infeasible with a single modality. Furthermore, it supports sequential and condition-based task planning through the collaboration of multiple LLM agents. Experimental results in various task scenarios show that the proposed framework consistently improves overall task success rates compared with unimodal settings with different LLMs and achieves a higher success rate compared to using only visual or force data.
키워드
- 제목
- MulPlanLM: multimodal robotic task planning with vision-language models and physical feedback
- 저자
- Son, Young-Chae; Lee, Dong-Han; Lim, Soo-Chul
- 발행일
- 2026-08
- 유형
- Article
- 권
- 19
- 호
- 5