Python으로 배우는 금융 분야 Machine Learning
Nathan George
Data Science Professor




랜덤 포레스트
from sklearn.ensemble import RandomForestRegressor
random_forest = RandomForestRegressor()
random_forest.fit(train_features, train_targets)
print(random_forest.score(train_features, train_targets))
random_forest = RandomForestRegressor(n_estimators=200,
max_depth=5,
max_features=4,
random_state=42)
from sklearn.model_selection import ParameterGrid grid = {'n_estimators': [200], 'max_depth':[3, 5], 'max_features': [4, 8]}from pprint import pprint pprint(list(ParameterGrid(grid)))
[{'max_depth': 3, 'max_features': 4, 'n_estimators': 200},
{'max_depth': 3, 'max_features': 8, 'n_estimators': 200},
{'max_depth': 5, 'max_features': 4, 'n_estimators': 200},
{'max_depth': 5, 'max_features': 8, 'n_estimators': 200}]
test_scores = [] # 파라미터 그리드를 순회하며 하이퍼파라미터를 설정하고 점수를 저장합니다 for g in ParameterGrid(grid): rfr.set_params(**g) # **는 딕셔너리 "언패킹"을 의미합니다 rfr.fit(train_features, train_targets) test_scores.append(rfr.score(test_features, test_targets))# 테스트 점수에서 최적 하이퍼파라미터를 찾아 출력합니다 best_idx = np.argmax(test_scores) print(test_scores[best_idx]) print(ParameterGrid(grid)[best_idx])
0.05594252725411142
{'max_depth': 5, 'max_features': 8, 'n_estimators': 200}
Python으로 배우는 금융 분야 Machine Learning