Python을 활용한 고객 세분화
Karolis Urbonas
Head of Data Science, Amazon
행:
열:

행:
열:

영국 기반 온라인 소매점의 50만 건 이상 거래.
본 강의에서는 무작위 20% 표본을 사용합니다.
online.head()

def get_month(x): return dt.datetime(x.year, x.month, 1)online['InvoiceMonth'] = online['InvoiceDate'].apply(get_month)grouping = online.groupby('CustomerID')['InvoiceMonth']online['CohortMonth'] = grouping.transform('min')online.head()
year, month, day 정수 값을 추출하는 함수를 정의합니다.
본 강의 전체에서 사용합니다.
def get_date_int(df, column):
year = df[column].dt.year
month = df[column].dt.month
day = df[column].dt.day
return year, month, day
invoice_year, invoice_month, _ = get_date_int(online, 'InvoiceMonth') cohort_year, cohort_month, _ = get_date_int(online, 'CohortMonth')years_diff = invoice_year - cohort_year months_diff = invoice_month - cohort_monthonline['CohortIndex'] = years_diff * 12 + months_diff + 1 online.head()
grouping = online.groupby(['CohortMonth', 'CohortIndex'])cohort_data = grouping['CustomerID'].apply(pd.Series.nunique)cohort_data = cohort_data.reset_index()cohort_counts = cohort_data.pivot(index='CohortMonth', columns='CohortIndex', values='CustomerID')print(cohort_counts)

Python을 활용한 고객 세분화