WILDS: A Benchmark of in-the-Wild Distribution Shifts
Creators
- Koh, Pang Wei
- Sagawa, Shiori
- Marklund, Henrik
- Xie, Sang Michael
- Zhang, Marvin
- Balsubramani, Akshay
- Hu, Weihua
- Yasunaga, Michihiro
- Phillips, Richard Lanas
- Gao, Irena
- Lee, Tony
- David, Etienne
-
Stavness, Ian
-
Guo, Wei
- Earnshaw, Berton A.
- Haque, Imran S.
-
Beery, Sara
-
Leskovec, Jure
-
Kundaje, Anshul
- Pierson, Emma
- Levine, Sergey
- Finn, Chelsea
- Liang, Percy
Abstract
Distribution shifts—where the training distribution differs from the test distribution—can substantially degrade the accuracy of machine learning (ML) systems deployed in the wild. Despite their ubiquity in the real-world deployments, these distribution shifts are under-represented in the datasets widely used in the ML community today. To address this gap, we present WILDS, a curated benchmark of 10 datasets reflecting a diverse range of distribution shifts that naturally arise in real-world applications, such as shifts across hospitals for tumor identification; across camera traps for wildlife monitoring; and across time and location in satellite imaging and poverty mapping. On each dataset, we show that standard training yields substantially lower out-of-distribution than in-distribution performance. This gap remains even with models trained by existing methods for tackling distribution shifts, underscoring the need for new methods for training models that are more robust to the types of distribution shifts that arise in practice. To facilitate method development, we provide an open-source package that automates dataset loading, contains default model architectures and hyperparameters, and standardizes evaluations. The full paper, code, and leaderboards are available at https://wilds.stanford.edu.
Additional Information
© 2021 by the author(s). Reproducibility: An executable version of our paper, hosted on CodaLab, can be found at https://wilds.stanford.edu/codalab. This contains the exact commands, code, environment, and data used for the experiments reported in our paper, as well as all trained model weights. The WILDS package is open-source and can be found at https://github.com/p-lambda/wilds. Many people generously volunteered their time and expertise to advise us on WILDS. We are grateful for all of the helpful suggestions and constructive feedback from: Aditya Khosla, Andreas Schlueter, Annie Chen, Aleksander Madry, Alexander D'Amour, Allison Koenecke, Alyssa Lees, Ananya Kumar, Andrew Beck, Behzad Haghgoo, Charles Sutton, Christopher Yeh, Cody Coleman, Dan Hendrycks, Dan Jurafsky, Daniel Levy, Daphne Koller, David Tellez, Erik Jones, Evan Liu, Fisher Yu, Georgi Marinov, Hongseok Namkoong, Irene Chen, Jacky Kang, Jacob Schreiber, Jacob Steinhardt, Jared Dunnmon, Jean Feng, Jeffrey Sorensen, Jianmo Ni, John Hewitt, John Miller, Kate Saenko, Kelly Cochran, Kensen Shi, Kyle Loh, Li Jiang, Lucy Vasserman, Ludwig Schmidt, Luke Oakden-Rayner, Marco Tulio Ribeiro, Matthew Lungren, Megha Srivastava, Nelson Liu, Nimit Sohoni, Pranav Rajpurkar, Robin Jia, Rohan Taori, Sarah Bird, Sharad Goel, Sherrie Wang, Shyamal Buch, Stefano Ermon, Steve Yadlowsky, Tatsunori Hashimoto, Tengyu Ma, Vincent Hellendoorn, Yair Carmon, Zachary Lipton, and Zhenghao Chen. The design of the WILDS benchmark was inspired by the Open Graph Benchmark (Hu et al., 2020b), and we are grateful to the Open Graph Benchmark team for their advice and help in setting up our benchmark. This project was funded by an Open Philanthropy Project Award and NSF Award Grant No. 1805310. Shiori Sagawa was supported by the Herbert Kunzel Stanford Graduate Fellowship. Henrik Marklund was supported by the Dr. Tech. Marcus Wallenberg Foundation for Education in International Industrial Entrepreneurship, CIFAR, and Google. Sang Michael Xie and Marvin Zhang were supported by NDSEG Graduate Fellowships. Weihua Hu was supported by the Funai Overseas Scholarship and the Masason Foundation Fellowship. Sara Beery was supported by an NSF Graduate Research Fellowship and is a PIMCO Fellow in Data Science. Jure Leskovec is a Chan Zuckerberg Biohub investigator. Chelsea Finn is a CIFAR Fellow in the Learning in Machines and Brains Program. We also gratefully acknowledge the support of DARPA under Nos. N660011924033 (MCS); ARO under Nos. W911NF-16-1-0342 (MURI), W911NF-16-1-0171 (DURIP); NSF under Nos. OAC-1835598 (CINES), OAC-1934578 (HDR), CCF-1918940 (Expeditions), IIS-2030477 (RAPID); Stanford Data Science Initiative, Wu Tsai Neurosciences Institute, Chan Zuckerberg Biohub, Amazon, JPMorgan Chase, Docomo, Hitachi, JD.com, KDDI, NVIDIA, Dell, Toshiba, and UnitedHealth Group.Attached Files
Published - koh21a.pdf
Submitted - 2012.07421.pdf
Supplemental Material - koh21a-supp.pdf
Files
2012.07421.pdf
Additional details
Identifiers
- Eprint ID
- 111571
- Resolver ID
- CaltechAUTHORS:20211021-160959922
Related works
- Describes
- https://arxiv.org/abs/2012.07421 (URL)
- https://wilds.stanford.edu (URL)
- https://github.com/p-lambda/wilds (URL)
Funding
- Open Philanthropy
- NSF
- CNS-1805310
- Stanford University
- Knut and Alice Wallenberg Foundation
- Canadian Institute for Advanced Research (CIFAR)
- Natural Sciences and Engineering Research Council of Canada (NSERC)
- Funai Overseas Scholarship
- Masason Foundation
- NSF Graduate Research Fellowship
- PIMCO
- Chan Zuckerberg Initiative
- Defense Advanced Research Projects Agency (DARPA)
- N660011924033
- Army Research Office (ARO)
- W911NF-16-1-0342
- Army Research Office (ARO)
- W911NF-16-1-0171
- NSF
- OAC-1835598
- NSF
- OAC-1934578
- NSF
- CCF-1918940
- NSF
- IIS-2030477
- Stanford Data Science Initiative
- Wu Tsai Neurosciences Institute
- Amazon
- JPMorgan Chase
- Docomo
- Hitachi
- JD.com
- KDDI
- NVIDIA Corporation
- Dell
- Toshiba
- UnitedHealth Group
Dates
- Created
-
2021-10-26Created from EPrint's datestamp field
- Updated
-
2023-06-02Created from EPrint's last_modified field