On the optimal size of candidate feature set in random forest

Sunwoo Han, Hyunjoong Kim

Research output: Contribution to journalArticlepeer-review

28 Citations (Scopus)

Abstract

Random forest is an ensemble method that combines many decision trees. Each level of trees is determined by an optimal rule among a candidate feature set. The candidate feature set is a random subset of all features, and is different at each level of trees. In this article, we investigated whether the accuracy of Random forest is affected by the size of the candidate feature set. We found that the optimal size differs from data to data without any specific pattern. To estimate the optimal size of feature set, we proposed a novel algorithm which uses the out-of-bag error and the 'SearchSize' exploration. The proposed method is significantly faster than the standard grid search method while giving almost the same accuracy. Finally, we demonstrated that the accuracy of Random forest using the proposed algorithm has increased significantly compared to using a typical size of feature set.

Original languageEnglish
Article number898
JournalApplied Sciences (Switzerland)
Volume9
Issue number5
DOIs
Publication statusPublished - 2019

Bibliographical note

Publisher Copyright:
© 2019 by the authors.

All Science Journal Classification (ASJC) codes

  • Materials Science(all)
  • Instrumentation
  • Engineering(all)
  • Process Chemistry and Technology
  • Computer Science Applications
  • Fluid Flow and Transfer Processes

Fingerprint

Dive into the research topics of 'On the optimal size of candidate feature set in random forest'. Together they form a unique fingerprint.

Cite this