Fakerで現実的なデータセットを生成する

Pythonで学ぶデータプライバシーと匿名化

Rebeca Gonzalez

Data engineer

Fakerでデータを生成

fake_data.name()
'Kelly Clark'
fake_data.name_male()
'Antonio Henderson'
fake_data.name_female()
'Jennifer Ortega'
Pythonで学ぶデータプライバシーと匿名化

Clients DataFrame

clients_df
gender    active
0    Female    No
1    Male      Yes
2    Male      No
3    Female    Yes
4    Male      Yes
...    ...    ...
1465    Male    Yes
1466    Male    Yes
1467    Male    Yes
1468    Male    Yes
1469    Male    Yes
1470 rows × 2 columns
Pythonで学ぶデータプライバシーと匿名化

Fakerでデータセットを生成

  • 性別に合う一意の名前を生成
  • ランダムな都市を生成
  • 確率分布に従う指定都市を生成
  • メールを生成
  • 期間内の日付を生成
Pythonで学ぶデータプライバシーと匿名化

性別に合わせて名前を生成

データセットで一意の名前を生成

重複を避ける
# Fakerクラスをインポート
from faker import Faker

# Fakerを初期化 fake_data = Faker()
# 性別に応じた一意の名前を生成 clients_df['name'] = [fake_data.unique.name_female() if x == "Female"
else fake_data.unique.name_male()
for x in clients_df['gender']]
Pythonで学ぶデータプライバシーと匿名化

性別に合わせて名前を生成

# データセットを確認
clients_df
    gender    active    name
0    Female    No       Michelle Lang
1    Male      Yes      Robert Norton
2    Male      No       Matthew Brown
3    Female    Yes      Sherry Jones
4    Male      Yes      Steven Vega
...    ...    ...    ...
1465    Male    Yes    Bradley Smith
1466    Male    Yes    Tyler Yu
1467    Male    Yes    Mr. Joshua Gallegos
1468    Male    Yes    Brian Aguilar
1469    Male    Yes    David Johnson
1470 rows × 3 columns
Pythonで学ぶデータプライバシーと匿名化

ランダムな都市の生成

# ランダムな都市を生成
clients_df['city'] = [fake_data.city() 
                      for x in range(len(clients_df))]


clients_df.head()
    Gender    Active    Name              City
0    Female   No        Stacy Hooper      Reedland
1    Male     Yes       Michael Rogers    North Michellestad
2    Male     No        James Sanchez     West Josephburgh
3    Female   Yes       Taylor Berger     Hermanton
4    Male     Yes       Joshua Coleman    South Amandaland
Pythonで学ぶデータプライバシーと匿名化

メールの生成

# ドメインが異なるメールを生成
clients_df['contact email'] = [fake_data.company_email() 
                               for x in range(len(clients_df))]


# データセットを確認 clients_df.head()
    gender     active    name            city                  contact email
0    Female    No        Stacy Hooper    Reedland              [email protected]
1    Male      Yes       Michael Rogers  North Michellestad    [email protected]
2    Male      No        James Sanchez   West Josephburgh      [email protected]
3    Female    Yes       Taylor Berger   Hermanton             [email protected]
4    Male      Yes       Joshua Coleman  South Amandaland      [email protected]
Pythonで学ぶデータプライバシーと匿名化

メールの生成

# 会社風ドメイン+氏名に近いユーザー名のメールを生成
clients_df['Contact email'] = [x.replace(" ", "") + "@" +
                               fake_data.domain_name() 
                               for x in clients_df['Name']]

# 結果のDataFrameを確認 clients_df.head()
    gender     active    name            city                  contact email
0    Female    No        Stacy Hooper    Reedland              [email protected]
1    Male      Yes       Michael Rogers  North Michellestad    [email protected]
2    Male      No        James Sanchez   West Josephburgh      [email protected]
3    Female    Yes       Taylor Berger   Hermanton             [email protected]
4    Male      Yes       Joshua Coleman  South Amandaland      [email protected]
Pythonで学ぶデータプライバシーと匿名化

日付の生成

2つの時点の間の日付

# 期間内の日付を生成
clients_df['date'] = [fake_data.date_between(start_date="-10y", end_date="now") 
                               for x in range(len(clients_df))]

# 結果のDataFrameを確認 clients_df.head()
    gender     active    name            city                  contact email                  date
0    Female    No        Stacy Hooper    Reedland              [email protected]       2019-11-20
1    Male      Yes       Michael Rogers  North Michellestad    [email protected]       2015-02-22
2    Male      No        James Sanchez   West Josephburgh      [email protected]  2015-12-11
3    Female    Yes       Taylor Berger   Hermanton             [email protected]       2012-12-13
4    Male      Yes       Joshua Coleman  South Amandaland      [email protected]      2014-05-22
Pythonで学ぶデータプライバシーと匿名化

確率分布に従って都市を生成

実データを模倣する際、本物の値の名前を漏らさないようにできます。

# numpyをインポート
import numpy as np


# 確率を取得または指定 p = (0.58, 0.23, 0.16, 0.03) cities = ["New York", "Chicago", "Seattle", "Dallas"]
# 指定都市から分布に従って生成 clients_df['city'] = np.random.choice(cities, size=len(clients_df), p=p)
Pythonで学ぶデータプライバシーと匿名化

確率分布に従って都市を生成

# 結果のデータセットを確認
clients_df.head()
    gender     active    name            city        contact email                  date
0    Female    No        Stacy Hooper    Chicago     [email protected]       2019-11-20
1    Male      Yes       Michael Rogers  New York    [email protected]       2015-02-22
2    Male      No        James Sanchez   New York    [email protected]  2015-12-11
3    Female    Yes       Taylor Berger   Chicago     [email protected]       2012-12-13
4    Male      Yes       Joshua Coleman  New York    [email protected]      2014-05-22
Pythonで学ぶデータプライバシーと匿名化

データセットを生成しましょう!

Pythonで学ぶデータプライバシーと匿名化

Preparing Video For Download...