Published December 2017 | Version public
Book Section - Chapter

Massively-parallel best subset selection for ordinary least-squares regression

  • 1. ROR icon University of Copenhagen
  • 2. ROR icon Heidelberg Institute for Theoretical Studies
  • 3. ROR icon California Institute of Technology
  • 4. ROR icon Radboud University Nijmegen

Abstract

Selecting an optimal subset of k out of d features for linear regression models given n training instances is often considered intractable for feature spaces with hundreds or thousands of dimensions. We propose an efficient massively-parallel implementation for selecting such optimal feature subsets in a brute-force fashion for small k. By exploiting the enormous compute power provided by modern parallel devices such as graphics processing units, it can deal with thousands of input dimensions even using standard commodity hardware only. We evaluate the practical runtime using artificial datasets and sketch the applicability of our framework in the context of astronomy.

Additional Information

© 2017 IEEE. Fabian Gieseke acknowledges support from the Danish Industry Foundation through the Industrial Data Analysis Service (IDAS) and Christian Igel acknowledges support from the Innovation Fund Denmark through the Danish Center for Big Data Analytics Driven Innovation (DABAI).

Additional details

Identifiers

Eprint ID
84875
DOI
10.1109/SSCI.2017.8285225
Resolver ID
CaltechAUTHORS:20180216-161024980

Related works

Funding

Danish Industry Foundation
Danish Center for Big Data Analytics Driven Innovation (DABAI)

Dates

Created
2018-02-22
Created from EPrint's datestamp field
Updated
2021-11-15
Created from EPrint's last_modified field