R での tidymodels によるモデリング
David Svancer
Data Scientist

列の役割を定義
変数の型を判定
recipe() 関数で実施

必要な前処理ステップを追加
各ステップは固有の step_*() 関数で追加

recipe オブジェクトはデータ(通常は学習データ)で訓練されます
prep() 関数でレシピを訓練します

学習済みのデータ前処理を適用
bake() 関数でレシピを適用します

リードスコアの total_time を対数変換
leads_training
# A tibble: 996 x 7
purchased total_visits total_time pages_per_visit total_clicks lead_source us_location
<fct> <dbl> <dbl> <dbl> <dbl> <fct> <fct>
1 yes 7 1148 7 59 direct_traffic west
2 no 5 228 2.5 25 email southeast
3 no 7 481 2.33 21 organic_search west
4 no 4 177 4 37 direct_traffic west
5 no 2 1273 2 26 email midwest
# ... with 991 more rows
recipe() 関数
data 引数
step_log() に recipe を渡して対数変換ステップを追加
total_time と底を指定leads_log_rec <- recipe(purchased ~ ., data = leads_training) %>%step_log(total_time, base = 10)
leads_log_rec
Data Recipe
Inputs:
role #variables
outcome 1
predictor 6
Operations:
Log transformation on total_time
summary() に recipe オブジェクトを渡す
type 列role 列leads_log_rec %>%
summary()
# A tibble: 7 x 4
variable type role source
<chr> <chr> <chr> <chr>
1 total_visits numeric predictor original
2 total_time numeric predictor original
3 pages_per_visit numeric predictor original
4 total_clicks numeric predictor original
5 lead_source nominal predictor original
6 us_location nominal predictor original
7 purchased nominal outcome original
prep() 関数
recipe オブジェクトtraining 引数
学習済み recipe の表示
[trained] と表示leads_log_rec_prep <- leads_log_rec %>%
prep(training = leads_training)
leads_log_rec_prep
Data Recipe
Inputs:
role #variables
outcome 1
predictor 6
Training data contained 996 data points and
no missing data.
Operations:
Log transformation on total_time [trained]
bake() 関数
recipenew_data 引数leads_training でレシピを学習prep() が保持new_data に NULL を渡すleads_log_rec_prep %>%
bake(new_data = NULL)
# A tibble: 996 x 7
total_visits total_time ... us_location purchased
<dbl> <dbl> ... <fct> <fct>
1 7 3.06 ... west yes
2 5 2.36 ... southeast no
3 7 2.68 ... west no
4 4 2.25 ... west no
5 2 3.10 ... midwest no
# ... with 991 more rows
レシピ学習に未使用のデータを変換
new_data に渡すleads_log_rec_prep %>%
bake(new_data = leads_test)
# A tibble: 332 x 7
total_visits total_time ... us_location purchased
<dbl> <dbl> ... <fct> <fct>
1 8 2 ... west no
2 4 3.13 ... northeast yes
3 3 2.25 ... west no
4 2 1.20 ... midwest no
5 9 3.01 ... west yes
# ... with 327 more rows
R での tidymodels によるモデリング