Biến đổi đặc trưng để phân cụm tốt hơn

Unsupervised Learning bằng Python

Benjamin Wilson

Director of Research at lateral.io

Bộ dữ liệu rượu vang Piedmont

  • 178 mẫu từ 3 giống rượu vang đỏ: Barolo, Grignolino, Barbera

  • Đặc trưng đo thành phần hóa học, ví dụ độ cồn

  • Thuộc tính thị giác như “độ đậm màu”

1 Source: https://archive.ics.uci.edu/ml/datasets/Wine
Unsupervised Learning bằng Python

Phân cụm các mẫu rượu

from sklearn.cluster import KMeans
model = KMeans(n_clusters=3)
labels = model.fit_predict(samples)
Unsupervised Learning bằng Python

Cụm so với giống rượu

df = pd.DataFrame({'labels': labels, 
                       'varieties': varieties})
ct = pd.crosstab(df['labels'], df['varieties'])

print(ct)
varieties  Barbera  Barolo  Grignolino
labels                                
0               29      13          20
1                0      46           1
2               19       0          50
Unsupervised Learning bằng Python

Phương sai đặc trưng

  • Các đặc trưng của rượu có phương sai rất khác nhau!

  • Phương sai đo độ phân tán giá trị của đặc trưng

feature     variance
alcohol         0.65
malic_acid      1.24
...
od280           0.50
proline     99166.71

Biểu đồ phân tán od280 so với malic_acid

Unsupervised Learning bằng Python

Phương sai đặc trưng

  • Các đặc trưng của rượu có phương sai rất khác nhau!

  • Phương sai đo độ phân tán của giá trị đặc trưng

feature     variance
alcohol         0.65
malic_acid      1.24
...
od280           0.50
proline     99166.71

Biểu đồ phân tán od280 theo số quan sát

Unsupervised Learning bằng Python

StandardScaler

  • Trong k-means: phương sai đặc trưng = mức ảnh hưởng của đặc trưng

  • StandardScaler biến đổi mỗi đặc trưng có trung bình 0, phương sai 1

  • Khi đó các đặc trưng được “chuẩn hóa”

Biểu đồ phân tán od280 và proline sau chuẩn hóa

Unsupervised Learning bằng Python

sklearn StandardScaler

from sklearn.preprocessing import StandardScaler

scaler = StandardScaler()
scaler.fit(samples) StandardScaler(copy=True, with_mean=True, with_std=True)
samples_scaled = scaler.transform(samples)
Unsupervised Learning bằng Python

Phương thức tương tự

  • StandardScalerKMeans có phương thức tương tự

  • Dùng fit() / transform() với StandardScaler

  • Dùng fit() / predict() với KMeans

Unsupervised Learning bằng Python

StandardScaler, rồi KMeans

  • Cần 2 bước: StandardScaler, rồi KMeans

  • Dùng pipeline của sklearn để ghép nhiều bước

  • Dữ liệu chảy từ bước này sang bước kế tiếp

Unsupervised Learning bằng Python

Pipeline ghép nhiều bước

from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
scaler = StandardScaler()
kmeans = KMeans(n_clusters=3)

from sklearn.pipeline import make_pipeline
pipeline = make_pipeline(scaler, kmeans)
pipeline.fit(samples)
Pipeline(steps=...)
labels = pipeline.predict(samples)
Unsupervised Learning bằng Python

Chuẩn hóa đặc trưng cải thiện phân cụm

Khi chuẩn hóa đặc trưng:

varieties  Barbera  Barolo  Grignolino
labels                                
0                0      59           3
1               48       0           3
2                0       0          65

Không chuẩn hóa đặc trưng thì rất tệ:

varieties  Barbera  Barolo  Grignolino
labels                                
0               29      13          20
1                0      46           1
2               19       0          50
Unsupervised Learning bằng Python

Các bước tiền xử lý trong sklearn

  • StandardScaler là bước “tiền xử lý”

  • MaxAbsScalerNormalizer là ví dụ khác

Unsupervised Learning bằng Python

Ayo berlatih!

Unsupervised Learning bằng Python

Preparing Video For Download...