Skip to main navigation Skip to search Skip to main content

A vision-language conditioned physics-aware imitation learning approach for bimanual robotic dexterous assembly

  • Tian Wang
  • , Benhua Gao
  • , Haofei Ma
  • , Guoquan Zhang
  • , Duidi Wu
  • , Pai Zheng (Corresponding Author)

Research output: Journal article publicationJournal articleAcademic researchpeer-review

Abstract

Increasing complexity and precision in Human–Robot Collaborative Assembly (HRCA) require robots to understand language, vision, and tactile information for contact-rich manipulation. To address the challenge, this paper proposes a vision-language conditioned physics-aware imitation learning approach for bimanual dexterous assembly. Firstly, a Mixed Reality (MR)-based bilateral teleoperation system is designed for multimodal human demo collection. Then, a Vision-Language Model (VLM)-augmented physics-aware diffusion policy is developed for manipulation skill learning. Furthermore, a coarse-to-fine visual Chain-of-Thought (CoT) strategy is integrated for task planning. Finally, the proposed method has been demonstrated on a dual-arm dexterous hand-based robotic platform by performing various HRCA tasks.

Original languageEnglish
Number of pages5
JournalCIRP Annals
DOIs
Publication statusAccepted/In press - 2026

Keywords

  • Human–robot collaboration
  • Imitation learning
  • Manufacturing system

ASJC Scopus subject areas

  • Mechanical Engineering
  • Industrial and Manufacturing Engineering

Fingerprint

Dive into the research topics of 'A vision-language conditioned physics-aware imitation learning approach for bimanual robotic dexterous assembly'. Together they form a unique fingerprint.

Cite this