Feature selection
Features are measured variables supplied to a predictive model. Feature selection chooses a subset intended to retain useful information while removing unnecessary inputs.
Why select features?
Removing irrelevant or highly redundant variables can simplify a model, reduce computational cost and sometimes improve generalisation.
Methods
Filter methods rank features using properties measured before fitting the final model. Wrapper methods evaluate subsets through predictive performance. Embedded methods perform selection as part of training, as in some regularised models.
Avoiding leakage
Feature selection must be performed using training data only. Selecting features using the full dataset can contaminate later evaluation.
Key idea. Feature selection is part of model training and must be evaluated within the same safeguards used for all other model choices.