検証の活用

Pythonで挑むKaggleコンペティション

Yauhen Babakhin

Kaggle Grandmaster

データ漏洩

 

漏洩のロゴ

 

  • 特徴量でのリーク – 実運用では利用不可のデータを使用

  • 検証戦略でのリーク – 実運用と異なる検証戦略

Pythonで挑むKaggleコンペティション

時系列データ

 

時系列データに不適切な検証戦略

Pythonで挑むKaggleコンペティション

時系列K-fold交差検証

 

 

時系列K-fold交差検証の方式

Pythonで挑むKaggleコンペティション

時系列K-fold交差検証

# Import TimeSeriesSplit
from sklearn.model_selection import TimeSeriesSplit

# Create a TimeSeriesSplit object
time_kfold = TimeSeriesSplit(n_splits=5)
# Sort train by date
train = train.sort_values('date')

# Loop through each cross-validation split
for train_index, test_index in time_kfold.split(train):
    cv_train, cv_test = train.iloc[train_index], train.iloc[test_index]
Pythonで挑むKaggleコンペティション

検証パイプライン

# List for the results
fold_metrics = []

for train_index, test_index in CV_STRATEGY.split(train): cv_train, cv_test = train.iloc[train_index], train.iloc[test_index]
# Train a model model.fit(cv_train)
# Make predictions predictions = model.predict(cv_test)
# Calculate the metric metric = evaluate(cv_test, predictions) fold_metrics.append(metric)
Pythonで挑むKaggleコンペティション

モデル比較

 

フォールド モデルA MSE モデルB MSE
フォールド1 2.95 2.97
フォールド2 2.84 2.45
フォールド3 2.62 2.73
フォールド4 2.79 2.83
Pythonで挑むKaggleコンペティション

総合検証スコア

import numpy as np

# Simple mean over the folds
mean_score = np.mean(fold_metrics)
# Overall validation score
overall_score_minimizing = np.mean(fold_metrics) + np.std(fold_metrics)
# Or
overall_score_maximizing = np.mean(fold_metrics) - np.std(fold_metrics)
Pythonで挑むKaggleコンペティション

モデル比較

フォールド モデルA MSE モデルB MSE
フォールド1 2.95 2.97
フォールド2 2.84 2.45
フォールド3 2.62 2.73
フォールド4 2.79 2.83
平均 2.80 2.75
総合 2.919 2.935
Pythonで挑むKaggleコンペティション

練習しましょう!

Pythonで挑むKaggleコンペティション

Preparing Video For Download...