특성 공학과 과적합

Python으로 설계하는 Machine Learning 워크플로

Dr. Chris Anagnostopoulos

Honorary Associate Professor

비정형 데이터의 특성 추출

단일 환자의 ECG는 보통 종이에 시간열로 기록됨.

arrhythmias.head()
   age  sex  height  weight  ...    chV6_TwaveAmp  chV6_QRSA  chV6_QRSTA  class
0   75    0     190      80  ...              2.9       23.3        49.4      0
1   56    1     165      64  ...              2.1       20.4        38.8      0
2   54    0     172      95  ...              3.4       12.3        49.0      0
3   55    0     175      94  ...              2.6       34.6        61.6      1
4   75    0     190      80  ...              3.9       25.4        62.8      0
Python으로 설계하는 Machine Learning 워크플로

범주형 변수의 레이블 인코딩

numpy.unique(credit_scoring['purpose'])
array(['business', 'buy_domestic_appliance', 'buy_furniture_equipment',
       'buy_new_car', 'buy_radio_tv', 'buy_used_car', 'education',
       'other', 'repairs', 'retraining'], dtype=object)
numpy.unique(LabelEncoder().fit_transform(credit_scoring['purpose']))
array([0, 1, 2, 3, 4, 5, 6, 7, 8, 9])
Python으로 설계하는 Machine Learning 워크플로

범주형 변수의 레이블 인코딩

범주를 숫자로 인코딩하고, 의사결정나무가 이를 임의의 연속 구간으로 분할하는 예시.

Python으로 설계하는 Machine Learning 워크플로

범주형 변수의 원-핫 인코딩

pd.get_dummies(credit_scoring['purpose']).iloc[1]
purpose_business                   0
purpose_buy_domestic_appliance     0
purpose_buy_furniture_equipment    0
purpose_buy_new_car                0
purpose_buy_radio_tv               1
purpose_buy_used_car               0
purpose_education                  0
purpose_other                      0
purpose_repairs                    0
purpose_retraining                 0
Python으로 설계하는 Machine Learning 워크플로

범주형 변수의 키워드 인코딩

from sklearn.feature_extraction.text import CountVectorizer
vec = CountVectorizer()

credit_scoring['purpose'] = credit_scoring['purpose'].apply( lambda s: ' '.join(s.split('_')), 0)
dummy_matrix = vec.fit_transform(credit_scoring['purpose']).toarray()
pd.DataFrame(dummy_matrix, columns=vec.get_feature_names()).head()
      appliance  business  buy  car  ...   repairs  retraining  tv  used
0          0         0    1    0  ...         0           0   1     0
1          0         0    1    0  ...         0           0   1     0
2          0         0    0    0  ...         0           0   0     0
3          0         0    1    0  ...         0           0   0     0
Python으로 설계하는 Machine Learning 워크플로

차원과 특성 공학

credit의 범주형 변수:

  • 레이블 인코딩: 1열
  • 원-핫 인코딩: 10열
  • 키워드 인코딩: 15열

arrhythmias의 ECG 특성:

  • 250개 이상
Python으로 설계하는 Machine Learning 워크플로

추가 열이 0개일 때는 학습용 정확도가 0.8이고, 가짜 열을 더할수록 거의 1.0까지 상승. 검증용 정확도는 가짜 열을 추가하는 즉시 하락.

Python으로 설계하는 Machine Learning 워크플로

특성 선택

from np.random import uniform
fakes = pd.DataFrame(
    uniform(low=0.0, high=1.0, size=n * 100).reshape(X.shape[0], 100),
    columns=['fake_' + str(j) for j in range(100)]
)
X_with_fakes = pd.concat([X, fakes], 1)
Python으로 설계하는 Machine Learning 워크플로

특성 선택

from sklearn.feature_selection import chi2, SelectKBest
sk = SelectKBest(chi2, k=20)
which_selected = sk.fit(X_with_fakes, y).get_support()
X_with_fakes.columns[which_selected]
['checking_status', 'duration', 'credit_history', 'purpose',
    'credit_amount', 'savings_status', 'installment_commitment',
    'personal_status', 'property_magnitude', 'age', 'other_payment_plans',
    'job', 'own_telephone', 'fake_28', 'fake_46', 'fake_51', 'fake_60',
    'fake_83', 'fake_95', 'fake_99']
Python으로 설계하는 Machine Learning 워크플로

어디에나 존재하는 절충

Python으로 설계하는 Machine Learning 워크플로

Preparing Video For Download...