使用 Faker 生成逼真的数据集

Python 中的数据隐私与匿名化

Rebeca Gonzalez

Data engineer

用 Faker 生成数据

fake_data.name()
'Kelly Clark'
fake_data.name_male()
'Antonio Henderson'
fake_data.name_female()
'Jennifer Ortega'
Python 中的数据隐私与匿名化

Clients DataFrame

clients_df
gender    active
0    Female    No
1    Male      Yes
2    Male      No
3    Female    Yes
4    Male      Yes
...    ...    ...
1465    Male    Yes
1466    Male    Yes
1467    Male    Yes
1468    Male    Yes
1469    Male    Yes
1470 rows × 2 columns
Python 中的数据隐私与匿名化

使用 Faker 生成数据集

  • 生成与性别一致的唯一姓名
  • 生成随机城市
  • 生成按概率分布的指定城市
  • 生成邮箱
  • 生成指定时间范围内的日期
Python 中的数据隐私与匿名化

使姓名与性别匹配

在数据集中生成唯一姓名

避免重复
# 导入 Faker 类
from faker import Faker

# 初始化 Faker fake_data = Faker()
# 按性别生成姓名,使其在数据集中唯一 clients_df['name'] = [fake_data.unique.name_female() if x == "Female"
else fake_data.unique.name_male()
for x in clients_df['gender']]
Python 中的数据隐私与匿名化

使姓名与性别匹配

# 查看数据集
clients_df
    gender    active    name
0    Female    No       Michelle Lang
1    Male      Yes      Robert Norton
2    Male      No       Matthew Brown
3    Female    Yes      Sherry Jones
4    Male      Yes      Steven Vega
...    ...    ...    ...
1465    Male    Yes    Bradley Smith
1466    Male    Yes    Tyler Yu
1467    Male    Yes    Mr. Joshua Gallegos
1468    Male    Yes    Brian Aguilar
1469    Male    Yes    David Johnson
1470 rows × 3 columns
Python 中的数据隐私与匿名化

生成随机城市

# 生成随机城市
clients_df['city'] = [fake_data.city() 
                      for x in range(len(clients_df))]


clients_df.head()
    Gender    Active    Name              City
0    Female   No        Stacy Hooper      Reedland
1    Male     Yes       Michael Rogers    North Michellestad
2    Male     No        James Sanchez     West Josephburgh
3    Female   Yes       Taylor Berger     Hermanton
4    Male     Yes       Joshua Coleman    South Amandaland
Python 中的数据隐私与匿名化

生成邮箱

# 生成不同域名的邮箱
clients_df['contact email'] = [fake_data.company_email() 
                               for x in range(len(clients_df))]


# 查看数据集 clients_df.head()
    gender     active    name            city                  contact email
0    Female    No        Stacy Hooper    Reedland              [email protected]
1    Male      Yes       Michael Rogers  North Michellestad    [email protected]
2    Male      No        James Sanchez   West Josephburgh      [email protected]
3    Female    Yes       Taylor Berger   Hermanton             [email protected]
4    Male      Yes       Joshua Coleman  South Amandaland      [email protected]
Python 中的数据隐私与匿名化

生成邮箱

# 生成类似公司域名且用户名接近姓名的邮箱
clients_df['Contact email'] = [x.replace(" ", "") + "@" +
                               fake_data.domain_name() 
                               for x in clients_df['Name']]

# 查看结果 DataFrame clients_df.head()
    gender     active    name            city                  contact email
0    Female    No        Stacy Hooper    Reedland              [email protected]
1    Male      Yes       Michael Rogers  North Michellestad    [email protected]
2    Male      No        James Sanchez   West Josephburgh      [email protected]
3    Female    Yes       Taylor Berger   Hermanton             [email protected]
4    Male      Yes       Joshua Coleman  South Amandaland      [email protected]
Python 中的数据隐私与匿名化

生成日期

两个时间点之间的日期

# 生成指定时间范围内的日期
clients_df['date'] = [fake_data.date_between(start_date="-10y", end_date="now") 
                               for x in range(len(clients_df))]

# 查看结果 DataFrame clients_df.head()
    gender     active    name            city                  contact email                  date
0    Female    No        Stacy Hooper    Reedland              [email protected]       2019-11-20
1    Male      Yes       Michael Rogers  North Michellestad    [email protected]       2015-02-22
2    Male      No        James Sanchez   West Josephburgh      [email protected]  2015-12-11
3    Female    Yes       Taylor Berger   Hermanton             [email protected]       2012-12-13
4    Male      Yes       Joshua Coleman  South Amandaland      [email protected]      2014-05-22
Python 中的数据隐私与匿名化

按概率分布生成城市

在仿真真实数据集时,可避免泄露真实取值的名称。

# 导入 numpy
import numpy as np


# 获取或指定概率 p = (0.58, 0.23, 0.16, 0.03) cities = ["New York", "Chicago", "Seattle", "Dallas"]
# 按给定分布从所选城市中生成值 clients_df['city'] = np.random.choice(cities, size=len(clients_df), p=p)
Python 中的数据隐私与匿名化

按概率分布生成城市

# 查看结果数据集
clients_df.head()
    gender     active    name            city        contact email                  date
0    Female    No        Stacy Hooper    Chicago     [email protected]       2019-11-20
1    Male      Yes       Michael Rogers  New York    [email protected]       2015-02-22
2    Male      No        James Sanchez   New York    [email protected]  2015-12-11
3    Female    Yes       Taylor Berger   Chicago     [email protected]       2012-12-13
4    Male      Yes       Joshua Coleman  New York    [email protected]      2014-05-22
Python 中的数据隐私与匿名化

让我们生成数据集!

Python 中的数据隐私与匿名化

Preparing Video For Download...