Python으로 배우는 데이터 프라이버시와 익명화
Rebeca Gonzalez
Data engineer



"사회적·법적 규범을 충족하는 정보 흐름을 보장하는 능력."

단독으로 또는 다른 관련 데이터와 결합하여 개인을 식별할 수 있는 데이터.



단독으로는 개인을 추적할 수 없는 데이터
성별, 직업, 우편번호, 출생 도시처럼 단독으로는 개인을 추적할 수 없는 데이터.


개인 정보 보호를 위해 선택 정보 제거.
# Attribute suppression on Sensitive PII "name" suppressed_salaries = salaries.drop('name', axis="columns")# Explore obtained dataset suppressed_salaries.head()
gender status salary pay_basis position_title
0 Male Employee 64400.0 Per Annum DEPUTY DIRECTOR
1 Male Employee 43600.0 Per Annum ASSOCIATE DIRECTOR
2 Male Employee 120000.0 Per Annum SPECIAL ASSISTANT TO THE PRESIDENT AND DEPUTY ...
3 Male Employee 86200.0 Per Annum LEAD ADVANCE REPRESENTATIVE
4 Male Employee 106000.0 Per Annum SPECIAL ASSISTANT TO THE PRESIDENT AND DIRECTO...
# Explore the DataFrame
salaries.head()
hours performance salary
0 72 51 $80,500.00
1 20 99 $2,805,000.00
3 75 62 $75,800.00
4 74 58 $60,000.00
5 70 54 $79,000.00
# Drop rows with salaries higher than 2,000,000
salaries = salaries.drop(salaries[salaries.Salary > 2000000].index)
# See reasulting DataFrame
salaries.head()
hours performance salary
0 72 51 80500
2 75 62 75800
3 74 58 60000
4 70 54 79000
5 68 53 62000

Python으로 배우는 데이터 프라이버시와 익명화