Tạo tập dữ liệu thực tế với Faker

Bảo mật dữ liệu và Ẩn danh trong Python

Rebeca Gonzalez

Data engineer

Tạo dữ liệu với Faker

fake_data.name()
'Kelly Clark'
fake_data.name_male()
'Antonio Henderson'
fake_data.name_female()
'Jennifer Ortega'
Bảo mật dữ liệu và Ẩn danh trong Python

DataFrame khách hàng

clients_df
gender    active
0    Female    No
1    Male      Yes
2    Male      No
3    Female    Yes
4    Male      Yes
...    ...    ...
1465    Male    Yes
1466    Male    Yes
1467    Male    Yes
1468    Male    Yes
1469    Male    Yes
1470 rows × 2 columns
Bảo mật dữ liệu và Ẩn danh trong Python

Tạo tập dữ liệu với Faker

  • Tạo tên duy nhất, phù hợp giới tính
  • Tạo thành phố ngẫu nhiên
  • Tạo thành phố xác định theo phân phối xác suất
  • Tạo email
  • Tạo ngày trong một khoảng thời gian
Bảo mật dữ liệu và Ẩn danh trong Python

Làm cho tên khớp với giới tính

Tạo tên duy nhất trong tập dữ liệu

Tránh trùng lặp
# Import lớp Faker
from faker import Faker

# Khởi tạo Faker fake_data = Faker()
# Tạo tên theo giới tính, đảm bảo duy nhất trong tập dữ liệu clients_df['name'] = [fake_data.unique.name_female() if x == "Female"
else fake_data.unique.name_male()
for x in clients_df['gender']]
Bảo mật dữ liệu và Ẩn danh trong Python

Làm cho tên khớp với giới tính

# Khám phá dữ liệu
clients_df
    gender    active    name
0    Female    No       Michelle Lang
1    Male      Yes      Robert Norton
2    Male      No       Matthew Brown
3    Female    Yes      Sherry Jones
4    Male      Yes      Steven Vega
...    ...    ...    ...
1465    Male    Yes    Bradley Smith
1466    Male    Yes    Tyler Yu
1467    Male    Yes    Mr. Joshua Gallegos
1468    Male    Yes    Brian Aguilar
1469    Male    Yes    David Johnson
1470 rows × 3 columns
Bảo mật dữ liệu và Ẩn danh trong Python

Tạo thành phố ngẫu nhiên

# Tạo thành phố ngẫu nhiên
clients_df['city'] = [fake_data.city() 
                      for x in range(len(clients_df))]


clients_df.head()
    Gender    Active    Name              City
0    Female   No        Stacy Hooper      Reedland
1    Male     Yes       Michael Rogers    North Michellestad
2    Male     No        James Sanchez     West Josephburgh
3    Female   Yes       Taylor Berger     Hermanton
4    Male     Yes       Joshua Coleman    South Amandaland
Bảo mật dữ liệu và Ẩn danh trong Python

Tạo email

# Tạo email với các tên miền khác nhau
clients_df['contact email'] = [fake_data.company_email() 
                               for x in range(len(clients_df))]


# Khám phá dữ liệu clients_df.head()
    gender     active    name            city                  contact email
0    Female    No        Stacy Hooper    Reedland              [email protected]
1    Male      Yes       Michael Rogers  North Michellestad    [email protected]
2    Male      No        James Sanchez   West Josephburgh      [email protected]
3    Female    Yes       Taylor Berger   Hermanton             [email protected]
4    Male      Yes       Joshua Coleman  South Amandaland      [email protected]
Bảo mật dữ liệu và Ẩn danh trong Python

Tạo email

# Tạo email với tên miền giống công ty và username giống tên
clients_df['Contact email'] = [x.replace(" ", "") + "@" +
                               fake_data.domain_name() 
                               for x in clients_df['Name']]

# Khám phá DataFrame kết quả clients_df.head()
    gender     active    name            city                  contact email
0    Female    No        Stacy Hooper    Reedland              [email protected]
1    Male      Yes       Michael Rogers  North Michellestad    [email protected]
2    Male      No        James Sanchez   West Josephburgh      [email protected]
3    Female    Yes       Taylor Berger   Hermanton             [email protected]
4    Male      Yes       Joshua Coleman  South Amandaland      [email protected]
Bảo mật dữ liệu và Ẩn danh trong Python

Tạo ngày

Ngày trong khoảng thời gian

# Tạo ngày trong khoảng thời gian
clients_df['date'] = [fake_data.date_between(start_date="-10y", end_date="now") 
                               for x in range(len(clients_df))]

# Khám phá DataFrame kết quả clients_df.head()
    gender     active    name            city                  contact email                  date
0    Female    No        Stacy Hooper    Reedland              [email protected]       2019-11-20
1    Male      Yes       Michael Rogers  North Michellestad    [email protected]       2015-02-22
2    Male      No        James Sanchez   West Josephburgh      [email protected]  2015-12-11
3    Female    Yes       Taylor Berger   Hermanton             [email protected]       2012-12-13
4    Male      Yes       Joshua Coleman  South Amandaland      [email protected]      2014-05-22
Bảo mật dữ liệu và Ẩn danh trong Python

Tạo thành phố theo phân phối xác suất

Khi mô phỏng dữ liệu thực, ta có thể tránh lộ tên giá trị thật.

# Import numpy
import numpy as np


# Xác định hoặc chỉ định xác suất p = (0.58, 0.23, 0.16, 0.03) cities = ["New York", "Chicago", "Seattle", "Dallas"]
# Tạo các thành phố đã chọn theo một phân phối clients_df['city'] = np.random.choice(cities, size=len(clients_df), p=p)
Bảo mật dữ liệu và Ẩn danh trong Python

Tạo thành phố theo phân phối xác suất

# Xem dữ liệu kết quả
clients_df.head()
    gender     active    name            city        contact email                  date
0    Female    No        Stacy Hooper    Chicago     [email protected]       2019-11-20
1    Male      Yes       Michael Rogers  New York    [email protected]       2015-02-22
2    Male      No        James Sanchez   New York    [email protected]  2015-12-11
3    Female    Yes       Taylor Berger   Chicago     [email protected]       2012-12-13
4    Male      Yes       Joshua Coleman  New York    [email protected]      2014-05-22
Bảo mật dữ liệu và Ẩn danh trong Python

Hãy tạo dữ liệu!

Bảo mật dữ liệu và Ẩn danh trong Python

Preparing Video For Download...