Published January 10, 2023 | Version Submitted
Discussion Paper Open

Vision Transformers Are Good Mask Auto-Labelers

Abstract

We propose Mask Auto-Labeler (MAL), a high-quality Transformer-based mask auto-labeling framework for instance segmentation using only box annotations. MAL takes box-cropped images as inputs and conditionally generates their mask pseudo-labels.We show that Vision Transformers are good mask auto-labelers. Our method significantly reduces the gap between auto-labeling and human annotation regarding mask quality. Instance segmentation models trained using the MAL-generated masks can nearly match the performance of their fully-supervised counterparts, retaining up to 97.4% performance of fully supervised models. The best model achieves 44.1% mAP on COCO instance segmentation (test-dev 2017), outperforming state-of-the-art box-supervised methods by significant margins. Qualitative results indicate that masks produced by MAL are, in some cases, even better than human annotations.

Additional Information

Attribution 4.0 International (CC BY 4.0).

Attached Files

Submitted - 2301.03992.pdf

Files

2301.03992.pdf

Files (6.3 MB)

Name Size
md5:b40db0f22ef93222e5a3ec51ceca3fc5
6.3 MB Preview Download

Additional details

Identifiers

Eprint ID
120089
Resolver ID
CaltechAUTHORS:20230316-153757695

Related works

Dates

Created
2023-03-16
Created from EPrint's datestamp field
Updated
2023-03-16
Created from EPrint's last_modified field