Ẩn danh dữ liệu phân loại

Bảo mật dữ liệu và Ẩn danh trong Python

Rebeca Gonzalez

Instructor

Khái quát hóa

    Age    Gender    Department    Condition
0    30    F         Finance       Anxiety disorders
1    42    M         Production    Bronchitis
2    35    F         Marketing     Dysthymia
3    39    F         Production    Dysthymia
4    40    M         Marketing     Flu
    Age    Gender    Department    Condition
0    <40   F         Finance       Anxiety disorders
1    >=40  M         Production    Bronquitis
2    <40   F         Finance       Dysthymia
3    <40   F         Production    Dysthymia
4    >=40  M         Marketing     Flu
Bảo mật dữ liệu và Ẩn danh trong Python

Khái quát hóa dữ liệu phân loại

# Xem bộ dữ liệu
hr.head()
    Age    BusinessTravel        Department                EducationField    EmployeeNumber
0    41    Travel_Rarely         Sales                     Life Sciences     1
1    49    Travel_Frequently     Research & Development    Life Sciences     2
2    37    Travel_Rarely         Research & Development    Other             4
3    33    Travel_Frequently     Research & Development    Life Sciences     5
4    27    Travel_Rarely         Research & Development    Medical           7
Bảo mật dữ liệu và Ẩn danh trong Python

Dữ liệu phân loại

Số lượng giá trị có thể có hạn chế hoặc cố định.

  • chủng tộc
  • giới tính
  • quê quán
  • nhóm tuổi
  • trình độ học vấn
  • phim yêu thích và sở thích
Bảo mật dữ liệu và Ẩn danh trong Python

Ẩn danh dữ liệu phân loại

Biểu đồ tần suất các giá trị phân loại trong Education Field

Bảo mật dữ liệu và Ẩn danh trong Python

Ẩn danh dữ liệu phân loại

     Department                EducationField
0    Sales                     Life Sciences
1    Research & Development    Life Sciences
2    Research & Development    Other
3    Research & Development    Life Sciences
4    Research & Development    Medical

Bộ dữ liệu gốc

     Department                EducationField
0    Sales                     Medical
1    Research & Development    Marketing
2    Research & Development    Life Sciences
3    Research & Development    Other
4    Research & Development    Life Sciences

Bộ dữ liệu sau khi lấy mẫu theo phân phối xác suất của cột educationField trong bộ dữ liệu gốc.

Bảo mật dữ liệu và Ẩn danh trong Python

Lấy mẫu từ dữ liệu

Cục Điều tra Dân số Hoa Kỳ công khai các mẫu dữ liệu thu thập về công dân.

Giúp tính các mẫu thống kê quy mô lớn:

  • trung bình
  • phương sai
  • cụm
Bảo mật dữ liệu và Ẩn danh trong Python

Khám phá phân phối

# Hiển thị tần suất tuyệt đối của từng giá trị duy nhất
hr['EducationField'].value_counts()
Life Sciences       606
Medical             464
Marketing           159
Technical Degree    132
Other                82
Human Resources      27
Name: EducationField, dtype: int64
Bảo mật dữ liệu và Ẩn danh trong Python

Khám phá phân phối

# Vẽ biểu đồ cột cho các hạng mục
df['BusinessTravel'].value_counts().plot(kind='bar')

Biểu đồ cột các giá trị phân loại trong cột Business Travel

Bảo mật dữ liệu và Ẩn danh trong Python

Khám phá phân phối

# Lấy tần suất tuyệt đối của từng giá trị duy nhất
counts = hr['EducationField'].value_counts()

# In danh sách chỉ mục
print(counts.index)
Index(['Life Sciences', 'Medical', 'Marketing', 
      'Technical Degree', 'Other', 'Human Resources'],
       dtype='object')
Bảo mật dữ liệu và Ẩn danh trong Python

Khám phá phân phối

# Phân phối xác suất của từng giá trị duy nhất
counts = df['EducationField'].value_counts(normalize=True)
Life Sciences       0.412245
Medical             0.315646
Marketing           0.108163
Technical Degree    0.089796
Other               0.055782
Human Resources     0.018367
Name: EducationField, dtype: float64
Bảo mật dữ liệu và Ẩn danh trong Python

Khám phá phân phối

# Các giá trị tần suất của từng giá trị duy nhất
df['EducationField'].value_counts(normalize=True).values
array([0.4122449 , 0.31564626, 0.10816327, 0.08979592, 0.05578231,
       0.01836735])
Bảo mật dữ liệu và Ẩn danh trong Python

Lấy mẫu cùng phân phối

# Lấy mẫu theo phân phối xác suất
hr_sample['EducationField']= np.random.choice(counts.index, 
                                              p=counts.values, 
                                              size=len(hr))


# Xem bộ dữ liệu kết quả hr.head()
    Age    BusinessTravel        Department                EducationField    EmployeeNumber
0    41    Travel_Rarely         Sales                     Life Sciences     1
1    49    Travel_Frequently     Research & Development    Medical           2
2    37    Travel_Rarely         Research & Development    Marketing         4
3    33    Travel_Frequently     Research & Development    Technical Degree  5
4    27    Travel_Rarely         Research & Development    Medical           7
Bảo mật dữ liệu và Ẩn danh trong Python

Lấy mẫu cùng phân phối

# Hiển thị tần suất tuyệt đối của từng hạng mục
hr['EducationField'].value_counts()
Life Sciences       606
Medical             464
Marketing           159
Technical Degree    132
Other                82
Human Resources      27
Name: EducationField, dtype: int64
# Hiển thị tần suất của cột kết quả
hr_sample['EducationField'].value_counts()
Life Sciences       604
Medical             493
Marketing           158
Technical Degree    120
Other                61
Human Resources      34
Name: EducationField, dtype: int64
Bảo mật dữ liệu và Ẩn danh trong Python

Ayo berlatih!

Bảo mật dữ liệu và Ẩn danh trong Python

Preparing Video For Download...