Abstract
Increasing complexity and precision in Human–Robot Collaborative Assembly (HRCA) require robots to understand language, vision, and tactile information for contact-rich manipulation. To address the challenge, this paper proposes a vision-language conditioned physics-aware imitation learning approach for bimanual dexterous assembly. Firstly, a Mixed Reality (MR)-based bilateral teleoperation system is designed for multimodal human demo collection. Then, a Vision-Language Model (VLM)-augmented physics-aware diffusion policy is developed for manipulation skill learning. Furthermore, a coarse-to-fine visual Chain-of-Thought (CoT) strategy is integrated for task planning. Finally, the proposed method has been demonstrated on a dual-arm dexterous hand-based robotic platform by performing various HRCA tasks.
| Original language | English |
|---|---|
| Number of pages | 5 |
| Journal | CIRP Annals |
| DOIs | |
| Publication status | Accepted/In press - 2026 |
Keywords
- Human–robot collaboration
- Imitation learning
- Manufacturing system
ASJC Scopus subject areas
- Mechanical Engineering
- Industrial and Manufacturing Engineering
Fingerprint
Dive into the research topics of 'A vision-language conditioned physics-aware imitation learning approach for bimanual robotic dexterous assembly'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver