集成学习

Python 树模型机器学习

Elie Kawerk

Data Scientist

CART 的优点

  • 易于理解。

  • 易于解释。

  • 易于使用。

  • 灵活性:可表征非线性依赖。

  • 预处理:无需标准化或归一化特征等。

Python 树模型机器学习

CART 的局限性

  • 分类:只能产生正交决策边界。

  • 对训练集的微小变化敏感。

  • 高方差:不加约束的 CART 可能过拟合。

  • 解决方案:集成学习。

Python 树模型机器学习

集成学习

  • 在同一数据集上训练不同模型。

  • 让每个模型各自预测。

  • 元模型:聚合各模型预测。

  • 最终预测:更稳健、错误更少。

  • 最佳效果:模型擅长的方面互补。

Python 树模型机器学习

集成学习:可视化说明

ensemble-visual

Python 树模型机器学习

实践中的集成学习:投票分类器

  • 二分类任务。

  • 有 N 个分类器预测:P1、P2、…、PN,且 Pi = 0 或 1。

  • 元模型预测:硬投票。

Python 树模型机器学习

硬投票

hard-voting

Python 树模型机器学习

sklearn 中的投票分类器(乳腺癌数据集)

# Import functions to compute accuracy and split data
from sklearn.metrics import accuracy_score
from sklearn.model_selection import train_test_split

# Import models, including VotingClassifier meta-model
from sklearn.linear_model import LogisticRegression
from sklearn.tree import DecisionTreeClassifier
from sklearn.neighbors import KNeighborsClassifier as KNN
from sklearn.ensemble import VotingClassifier

# Set seed for reproducibility
SEED = 1
Python 树模型机器学习

sklearn 中的投票分类器(乳腺癌数据集)

# Split data into 70% train and 30% test
X_train, X_test, y_train, y_test = train_test_split(X, y,
                                                    test_size= 0.3,
                                                    random_state= SEED)
# Instantiate individual classifiers
lr = LogisticRegression(random_state=SEED)
knn = KNN()
dt = DecisionTreeClassifier(random_state=SEED)

# Define a list called classifier that contains the tuples (classifier_name, classifier) classifiers = [('Logistic Regression', lr), ('K Nearest Neighbours', knn), ('Classification Tree', dt)]
Python 树模型机器学习
# Iterate over the defined list of tuples containing the classifiers
for clf_name, clf in classifiers:
    #fit clf to the training set
    clf.fit(X_train, y_train)

    # Predict the labels of the test set
    y_pred = clf.predict(X_test)

    # Evaluate the accuracy of clf on the test set
    print('{:s} : {:.3f}'.format(clf_name, accuracy_score(y_test, y_pred)))
Logistic Regression: 0.947
K Nearest Neighbours: 0.930
Classification Tree: 0.930
Python 树模型机器学习

sklearn 中的投票分类器(乳腺癌数据集)

# Instantiate a VotingClassifier 'vc'
vc = VotingClassifier(estimators=classifiers) 

# Fit 'vc' to the traing set and predict test set labels
vc.fit(X_train, y_train)   
y_pred = vc.predict(X_test)

# Evaluate the test-set accuracy of 'vc'
print('Voting Classifier: {.3f}'.format(accuracy_score(y_test, y_pred)))
Voting Classifier: 0.953
Python 树模型机器学习

Passons à la pratique !

Python 树模型机器学习

Preparing Video For Download...