R 的樹狀模型機器學習
Sandro Raabe
Data Scientist
parsnip 決策樹的超參數:
min_n:節點分裂所需的最少樣本數tree_depth:樹的最大深度cost_complexity:對樹複雜度的懲罰parsnip 的預設值:
decision_tree(min_n = 20, tree_depth = 30, cost_complexity = 0.01)
調參的目標是找到一組最佳的超參數值。




spec_untuned <- decision_tree(min_n = tune(), tree_depth = tune()) %>% set_engine("rpart") %>% set_mode("classification")
Decision Tree Model Specification (classification)Main Arguments: tree_depth = tune() min_n = tune()
tune() 標記要調參的參數tree_grid <- grid_regular(parameters(spec_untuned),levels = 3 )
# A tibble: 9 x 2
min_n tree_depth
1 2 1
2 21 1
3 40 1
4 2 8
5 21 8
6 40 8
7 2 15
8 21 15
9 40 15
parameters()levels:每個超參數的網格點數
用法與參數:
metric_set() 包成的指標清單tune_results <- tune_grid(spec_untuned,outcome ~ .,resamples = my_folds,grid = tree_grid,metrics = metric_set(accuracy))
autoplot(tune_results)

# 選出表現最佳的參數 final_params <- select_best(tune_results)final_params
# A tibble: 1 x 3
min_n tree_depth .config
<int> <int> <chr>
1 2 8 Model4
# 套用到模型規格 best_spec <- finalize_model(spec_untuned, final_params)best_spec
Decision Tree Model Specification
(classification)
Main Arguments:
tree_depth = 8
min_n = 2
Computational engine: rpart
R 的樹狀模型機器學習