การเปรียบเทียบสตริง

การทำความสะอาดข้อมูลใน Python

Adel Nehme

VP of AI Curriculum, DataCamp

ในบทนี้

 

 

 

 

 

 

บทที่ 4 - Record linkage

การทำความสะอาดข้อมูลใน Python

Minimum edit distance

จำนวนขั้นตอนน้อยที่สุดที่ต้องใช้ในการเปลี่ยนสตริงหนึ่งไปเป็นอีกสตริงหนึ่ง

การทำความสะอาดข้อมูลใน Python

Minimum edit distance

จำนวนขั้นตอนน้อยที่สุดที่ต้องใช้ในการเปลี่ยนสตริงหนึ่งไปเป็นอีกสตริงหนึ่ง

การทำความสะอาดข้อมูลใน Python

Minimum edit distance

การทำความสะอาดข้อมูลใน Python

Minimum edit distance

Minimum edit distance ณ ตอนนี้: 2

การทำความสะอาดข้อมูลใน Python

Minimum edit distance

Minimum edit distance: 5

การทำความสะอาดข้อมูลใน Python

Minimum edit distance

 

การทำความสะอาดข้อมูลใน Python

อัลกอริทึม Minimum edit distance

อัลกอริทึม การดำเนินการ
Damerau-Levenshtein การแทรก, การแทนที่, การลบ, การสลับตำแหน่ง
Levenshtein การแทรก, การแทนที่, การลบ
Hamming การแทนที่เท่านั้น
Jaro distance การสลับตำแหน่งเท่านั้น
... ...

 

แพ็กเกจที่ใช้ได้: nltk, thefuzz, textdistance ..

การทำความสะอาดข้อมูลใน Python

อัลกอริทึม Minimum edit distance

อัลกอริทึม การดำเนินการ
Damerau-Levenshtein การแทรก, การแทนที่, การลบ, การสลับตำแหน่ง
Levenshtein การแทรก, การแทนที่, การลบ
Hamming การแทนที่เท่านั้น
Jaro distance การสลับตำแหน่งเท่านั้น
... ...

 

แพ็กเกจที่ใช้ได้: thefuzz

การทำความสะอาดข้อมูลใน Python

การเปรียบเทียบสตริงอย่างง่าย

# Lets us compare between two strings
from thefuzz import fuzz

# Compare reeding vs reading fuzz.WRatio('Reeding', 'Reading')
86
การทำความสะอาดข้อมูลใน Python

สตริงบางส่วนและลำดับที่แตกต่างกัน

# Partial string comparison
fuzz.WRatio('Houston Rockets', 'Rockets')
90
# Partial string comparison with different order
fuzz.WRatio('Houston Rockets vs Los Angeles Lakers', 'Lakers vs Rockets')
86
การทำความสะอาดข้อมูลใน Python

การเปรียบเทียบกับอาร์เรย์

# Import process
from thefuzz import process

# Define string and array of possible matches
string = "Houston Rockets vs Los Angeles Lakers"
choices = pd.Series(['Rockets vs Lakers', 'Lakers vs Rockets', 
                     'Houson vs Los Angeles', 'Heat vs Bulls'])

process.extract(string, choices, limit = 2)
[('Rockets vs Lakers', 86, 0), ('Lakers vs Rockets', 86, 1)]
การทำความสะอาดข้อมูลใน Python

การรวมหมวดหมู่ด้วยความคล้ายคลึงของสตริง

บทที่ 2

ใช้ .replace() เพื่อรวม "eur" เป็น "Europe"

 

แล้วถ้ามีรูปแบบที่หลากหลายเกินไปล่ะ?

"EU", "eur", "Europ", "Europa", "Erope", "Evropa"...

 

                                                                                                ความคล้ายคลึงของสตริง!

การทำความสะอาดข้อมูลใน Python

การรวมหมวดหมู่ด้วยการจับคู่สตริง

print(survey['state'].unique())
id          state
0      California
1            Cali
2      Calefornia
3      Calefornie
4      Californie
5       Calfornia
6      Calefernia
7        New York
8   New York City
...
categories
  state
0 California
1 New York
การทำความสะอาดข้อมูลใน Python

การรวมข้อมูลรัฐทั้งหมด

# For each correct category
for state in categories['state']:

# Find potential matches in states with typoes matches = process.extract(state, survey['state'], limit = survey.shape[0])
# For each potential match match for potential_match in matches: # If high similarity score if potential_match[1] >= 80:
# Replace typo with correct category survey.loc[survey['state'] == potential_match[0], 'state'] = state
การทำความสะอาดข้อมูลใน Python

Record linkage

record linkage

การทำความสะอาดข้อมูลใน Python

มาฝึกกันเถอะ!

การทำความสะอาดข้อมูลใน Python

Preparing Video For Download...