Augmented Reality and Vision-Language Models to Guide Humans Across Manual Tasks
Abstract
We present a novel mixed-reality pipeline for assisted assembly verification that enhances productivity, skill development, reliability and operational efficiency in manual tasks such as manufacturing, maintenance and DIY end-user assembly. Our system incorporates modern augmented reality (AR) headsets, vision-language models (VLM) and object detection models to construct an immersive assembly guidance platform. In real time, our system parses the state of the assembly and location of components from the user's point of view, providing labels in three-dimensional space, as well as visual and textual feedback, through an AR interface. Each assembly step is validated using the VLM and a state-space encoding of all sub-assemblies, such that user mistakes are detected and clear feedback is given on how to revert to a valid state. In this work, we present a proof-of-concept instance of this platform for the test case of guiding users through building a LEGO set. Our proof-of-concept uses an Apple VisionPro AR headset for spatial computing and user interfacing, a CNN-based foundational object detection model to label and track individual LEGO pieces, and a foundational vision-language model to validate assemblies and provide textual feedback. Users with no prior knowledge of the LEGO set's assembly are guided step-by-step throughout the process, detecting mistakes and giving textual and visual guidance on corrective measures, until the final assembly is validated. We present this test case as proof of a new paradigm in human-machine integration, where human judgment, perception and skill are augmented by artificial intelligence, seamlessly, in the human's domain.
Copyright and License
Files
3721245.3734043.pdf
Additional details
Related works
- Is published in
- Conference Proceeding: 10.1145/3721245 (DOI)
Caltech Custom Metadata
- Caltech groups
- Division of Engineering and Applied Science (EAS)
- Publication Status
- Published