PySpark로 하는 Feature Engineering
John Hogue
Lead Data Scientist, General Mills

df.select(['NO', 'UNITNUMBER', 'CLASS']).show()
+----+----------+-----+
| NO|UNITNUMBER|CLASS|
+----+----------+-----+
| 1| null| SF|
| 156| A8| SF|
| 157| 207| SF|
| 158| 701G| SF|
| 159| 36| SF|
분석에 불필요한 여러 필드
'NO' 자동 생성된 레코드 번호'UNITNUMBER' 관련 없는 데이터'CLASS' 모두 동일한 값 drop(*cols)
*cols – 삭제할 열 이름 또는 열 이름 목록.# List of columns to drop cols_to_drop = ['NO', 'UNITNUMBER', 'CLASS']# Drop the columns df = df.drop(*cols_to_drop)
where(condition)types.BooleanType 열 또는 SQL 표현식 문자열.like(other)~df = df.where(~df['POTENTIALSHORTSALE'].like('Not Disclosed'))
평균(μ)의 3 표준편차(3σ) 이내로 데이터 필터링

# Calculate values used for filtering std_val = df.agg({'SALESCLOSEPRICE': 'stddev'}).collect()[0][0] mean_val = df.agg({'SALESCLOSEPRICE': 'mean'}).collect()[0][0]# Create three standard deviation (μ ± 3σ) upper and lower bounds for data hi_bound = mean_val + (3 * std_val) low_bound = mean_val - (3 * std_val)# Use where() to filter the DataFrame between values df = df.where((df['LISTPRICE'] < hi_bound) & (df['LISTPRICE'] > low_bound))
DataFrame.dropna()
how: 'any' 또는 'all'. 'any'이면 null이 하나라도 있는 레코드를 삭제하고, 'all'이면 모든 값이 null인 경우에만 삭제합니다.thresh: 정수, 기본값 None. 지정 시, 비-null 값이 thresh 미만인 레코드를 삭제합니다. how 매개변수보다 우선 적용됩니다.subset: 고려할 열 이름의 선택적 목록.# Drop any records with NULL values df = df.dropna()# drop records if both LISTPRICE and SALESCLOSEPRICE are NULL df = df.dropna(how='all', subset['LISTPRICE', 'SALESCLOSEPRICE '])# Drop records where at least two columns have NULL values df = df.dropna(thresh=2)
중복이란 무엇인가?
dropDuplicates()
# Entire DataFrame df.dropDuplicates()# Check only a column list df.dropDuplicates(['streetaddress'])
PySpark로 하는 Feature Engineering