Published August 9, 2025 | Version Published
Conference Paper Open

Augmented Reality and Vision-Language Models to Guide Humans Across Manual Tasks

  • 1. ROR icon California Institute of Technology

Abstract

We present a novel mixed-reality pipeline for assisted assembly verification that enhances productivity, skill development, reliability and operational efficiency in manual tasks such as manufacturing, maintenance and DIY end-user assembly. Our system incorporates modern augmented reality (AR) headsets, vision-language models (VLM) and object detection models to construct an immersive assembly guidance platform. In real time, our system parses the state of the assembly and location of components from the user's point of view, providing labels in three-dimensional space, as well as visual and textual feedback, through an AR interface. Each assembly step is validated using the VLM and a state-space encoding of all sub-assemblies, such that user mistakes are detected and clear feedback is given on how to revert to a valid state. In this work, we present a proof-of-concept instance of this platform for the test case of guiding users through building a LEGO set. Our proof-of-concept uses an Apple VisionPro AR headset for spatial computing and user interfacing, a CNN-based foundational object detection model to label and track individual LEGO pieces, and a foundational vision-language model to validate assemblies and provide textual feedback. Users with no prior knowledge of the LEGO set's assembly are guided step-by-step throughout the process, detecting mistakes and giving textual and visual guidance on corrective measures, until the final assembly is validated. We present this test case as proof of a new paradigm in human-machine integration, where human judgment, perception and skill are augmented by artificial intelligence, seamlessly, in the human's domain.

Copyright and License

© 2025 Copyright held by the owner/author(s). Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s).

Files

3721245.3734043.pdf

Files (28.2 MB)

Name Size
md5:2d3dee91cce3a2cf9c9545defae83f6d
2.1 MB Preview Download
md5:14b1f6516fef8a1b72f62819241a0632
26.1 MB Preview Download

Additional details

Related works

Is published in
Conference Proceeding: 10.1145/3721245 (DOI)

Caltech Custom Metadata