Imputace mediánem

Machine Learning s balíčkem caret v R

Max Kuhn

Software Engineer at RStudio and creator of caret

Práce s chybějícími hodnotami

  • Většina modelů vyžaduje čísla, chybějící data neumí zpracovat
  • Běžný přístup: odstranění řádků s chybějícími hodnotami
    • Může vést ke zkreslení dat
    • Generuje příliš sebejisté modely
  • Lepší strategie: imputace mediánem!
    • Nahradí chybějící hodnoty mediány
    • Funguje dobře, pokud data chybí náhodně (MAR)
Machine Learning s balíčkem caret v R

Příklad: mtcars

# Generate some data with missing values
data(mtcars)
set.seed(42)
mtcars[sample(1:nrow(mtcars), 10), "hp"] <- NA
# Split target from predictors
Y <- mtcars$mpg
X <- mtcars[, 2:4]
# Try to fit a caret model
library(caret)
model <- train(X, Y)
Error in train.default(X, Y) : Stopping 
Machine Learning s balíčkem caret v R

Jednoduché řešení

# Now fit with median imputation
model <- train(X, Y, preProcess = "medianImpute")
print(model)
Random Forest 

32 samples
 3 predictor

Pre-processing: median imputation (3) 
Resampling: Bootstrapped (25 reps) 
Summary of sample sizes: 32, 32, 32, 32, 32, 32, ... 
Resampling results across tuning parameters:

  mtry  RMSE      Rsquared 
  2     2.617096  0.8234652
  3     2.670550  0.8164535

RMSE was used to select the optimal model using the smallest value.
The final value used for the model was mtry = 2. 
Machine Learning s balíčkem caret v R

Pojďme si procvičit!

Machine Learning s balíčkem caret v R

Preparing Video For Download...