Pixel2Catch: Multi-Agent Sim-to-Real Transfer for Agile Manipulation With a Single RGB Camera

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

To catch a thrown object, a robot must be able to perceive the object’s motion and generate control actions in a timely manner. Rather than explicitly estimating the object’s 3D position, this work focuses on a novel approach that recognizes object motion using pixel-level visual information extracted from consecutive RGB frames. Such visual cues capture changes in the object’s position and scale, allowing the policy to reason about the object’s motion. Furthermore, to achieve stable learning in a high-DoF system composed of a robot arm equipped with a multi-fingered hand, we design a heterogeneous multi-agent reinforcement learning framework that defines the arm and hand as independent agents with distinct roles. Each agent is trained cooperatively using role-specific observations and rewards, and the learned policies are successfully transferred from simulation to the real world. © 2026 IEEE.

키워드

AI-enabled roboticsdeep learning in graspingmanipulationreinforcement learning
제목
Pixel2Catch: Multi-Agent Sim-to-Real Transfer for Agile Manipulation With a Single RGB Camera
저자
Kim, SeongyongCho, JunhyeonLee, Kang-WonLim, Soo-Chul
DOI
10.1109/LRA.2026.3703582
발행일
2026-08
유형
Article
저널명
IEEE Robotics and Automation Letters
11
8
페이지
9287 ~ 9294