探索的データ分析

PythonでMachine Learningを使ってCTRを予測する

Kevin Huo

Instructor

特徴量を詳しく見る

print(df.columns)
['id', 'click', 'hour', 'C1', ... ]
print(df.dtypes)
id                  object
click                int64
...
  • int: 整数: 12 など
  • float: 小数: 3.024.56 など
  • object: 文字列: "hello""world" など
  • datetime: 日時: 2018-01-01 など
df.select_dtypes(
  include=['int', 'float'])
click                int64
...
PythonでMachine Learningを使ってCTRを予測する

欠損値

df.info()
Data columns (total 24 columns):
id            50000 non-null object
df['id'].isnull()
[False, False, False, False, ... ]
df.isnull().sum(axis = 0)
dtype: object
id                  0
...
df.isnull().sum(axis = 0).sum()
0
PythonでMachine Learningを使ってCTRを予測する

分布を見る

df.groupby(['search_engine_type', 
            'click']).size()
search_engine_type    click
1002          0          940
              1          240
                   ...
df.groupby(['search_engine_type',
            'click']).size().unstack()
click                  0     1
search_engine_type               
1002                 940   240
               ...
PythonでMachine Learningを使ってCTRを予測する

CTR別の内訳

df.reset_index()
click  search_engine_type      0     1
                     1002    940   240
df = df.rename(columns = {0: 'non_clicks'})
click  search_engine_type  non_clicks  clicks
                     1002         940     240
PythonでMachine Learningを使ってCTRを予測する

Passons à la pratique !

PythonでMachine Learningを使ってCTRを予測する

Preparing Video For Download...