Collaborative robots are increasingly popular for assisting humans at work and daily tasks. However, designing and setting up interfaces for human–robot collaboration is challenging, requiring the integration of multiple components, from perception and robot task control to the hardware itself. Frequently, this leads to highly customized solutions that rely on large amounts of costly training data, diverging from the ideal of flexible and general interfaces that empower robots to perceive and adapt to unstructured environments where they can naturally collaborate with humans. To overcome these challenges, this paper presents the Detection-Robot Management GPT (D-RMGPT), a robot-assisted planner based on Large Multimodal Models (LMM) to assist inexperienced operators in tasks without requiring any markers or previous training. D-RMGPT is composed of DetGPT-V and R-ManGPT. DetGPT-V, based on GPT-4V(vision),now replaced by GPT-4o, perceives the surrounding environment through one-shot analysis of prompted images from the current scenario, identifying the surrounding objects, human actions and interpreting speech commands. R-ManGPT, based on GPT-4 and aided by a 3D CNN-LSTM network, acts as operations and robot task manager, planning the next actions, and generating discrete robot actions. Planning precedence dependencies between tasks is solved by giving the system a single image with the full list of precedence relationships. Experimental tests on assembling a toy aircraft demonstrated that D-RMGPT is flexible and intuitive to use, achieving an assembly success rate of 83% while reducing the assembly time for inexperienced operators by 33% compared to the manual process. The cooking and food preparation assistance scenario demonstrated that D-RMGPT successfully anticipated operator needs and assisted in producing a pizza from selected ingredients, achieving a success rate of 87%. http://robotics-and-ai.github.io/LMMmodels/

D-RMGPT: Robot-assisted collaborative tasks driven by large multimodal models / Forlini, M., Babcinschi, M., Palmieri, G., Neto, P.. - In: BIOMIMETIC INTELLIGENCE AND ROBOTICS. - ISSN 2097-0242. - (2026). [Epub ahead of print] [10.1016/j.birob.2026.100334]

D-RMGPT: Robot-assisted collaborative tasks driven by large multimodal models

Forlini, Matteo
Primo
;
Palmieri, Giacomo
Penultimo
;
2026-01-01

Abstract

Collaborative robots are increasingly popular for assisting humans at work and daily tasks. However, designing and setting up interfaces for human–robot collaboration is challenging, requiring the integration of multiple components, from perception and robot task control to the hardware itself. Frequently, this leads to highly customized solutions that rely on large amounts of costly training data, diverging from the ideal of flexible and general interfaces that empower robots to perceive and adapt to unstructured environments where they can naturally collaborate with humans. To overcome these challenges, this paper presents the Detection-Robot Management GPT (D-RMGPT), a robot-assisted planner based on Large Multimodal Models (LMM) to assist inexperienced operators in tasks without requiring any markers or previous training. D-RMGPT is composed of DetGPT-V and R-ManGPT. DetGPT-V, based on GPT-4V(vision),now replaced by GPT-4o, perceives the surrounding environment through one-shot analysis of prompted images from the current scenario, identifying the surrounding objects, human actions and interpreting speech commands. R-ManGPT, based on GPT-4 and aided by a 3D CNN-LSTM network, acts as operations and robot task manager, planning the next actions, and generating discrete robot actions. Planning precedence dependencies between tasks is solved by giving the system a single image with the full list of precedence relationships. Experimental tests on assembling a toy aircraft demonstrated that D-RMGPT is flexible and intuitive to use, achieving an assembly success rate of 83% while reducing the assembly time for inexperienced operators by 33% compared to the manual process. The cooking and food preparation assistance scenario demonstrated that D-RMGPT successfully anticipated operator needs and assisted in producing a pizza from selected ingredients, achieving a success rate of 87%. http://robotics-and-ai.github.io/LMMmodels/
2026
Collaborative roboticsLarge multimodal modelsHuman–robot interaction
File in questo prodotto:
File Dimensione Formato  
1-s2.0-S2667379726000628-main.pdf

accesso aperto

Tipologia: Versione editoriale (versione pubblicata con il layout dell'editore)
Licenza d'uso: Creative commons
Dimensione 8.89 MB
Formato Adobe PDF
8.89 MB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11566/357917
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact