文字編碼入門

Feature Engineering for Machine Learning in Python

Robert O'Callaghan

Director of Data Science, Ordergroove

文字標準化

自由文字範例:

Fellow-Citizens of the Senate and of the House of Representatives: AMONG the vicissitudes incident to life no event could have filled me with greater anxieties than that of which the notification was transmitted by your order, and received on the th day of the present month.

Feature Engineering for Machine Learning in Python

資料集

print(speech_df.head())
                  Name           Inaugural Address    \ 
0    George Washington     First Inaugural Address
1    George Washington    Second Inaugural Address
2    John Adams                  Inaugural Address    
3    Thomas Jefferson      First Inaugural Address    
4    Thomas Jefferson     Second Inaugural Address

                        Date                               text
0    Thursday, April 30, 1789    Fellow-Citizens of the Sena...
1       Monday, March 4, 1793    Fellow Citizens: I AM again...
2     Saturday, March 4, 1797    WHEN it was first perceived...
3    Wednesday, March 4, 1801    Friends and Fellow-Citizens...
4       Monday, March 4, 1805    PROCEEDING, fellow-citizens...
Feature Engineering for Machine Learning in Python

移除不需要的字元

  • [a-zA-Z]:所有字母字元
  • [^a-zA-Z]:所有非字母字元
speech_df['text'] = speech_df['text']\
                   .str.replace('[^a-zA-Z]', ' ')
Feature Engineering for Machine Learning in Python

移除不需要的字元

處理前:

"Fellow-Citizens of the Senate and of the House of  
Representatives: AMONG the vicissitudes incident to   
life no event could have filled me with greater" ...

處理後:

"Fellow Citizens of the Senate and of the House of  
Representatives AMONG the vicissitudes incident to   
life no event could have filled me with greater" ...
Feature Engineering for Machine Learning in Python

統一大小寫

speech_df['text'] = speech_df['text'].str.lower()
print(speech_df['text'][0])
"fellow citizens of the senate and of the house of  
representatives among the vicissitudes incident to   
life no event could have filled me with greater"...
Feature Engineering for Machine Learning in Python

文字長度

speech_df['char_cnt'] = speech_df['text'].str.len()
print(speech_df['char_cnt'].head())
0    1889  
1     806  
2    2408  
3    1495  
4    2465
Name: char_cnt, dtype: int64
Feature Engineering for Machine Learning in Python

詞數

speech_df['word_cnt'] = 
    speech_df['text'].str.split()
speech_df['word_cnt'].head(1)
['fellow', 'citizens', 'of', 'the', 'senate', 'and',...
Feature Engineering for Machine Learning in Python

詞數

speech_df['word_counts'] = 
    speech_df['text'].str.split().str.len()
print(speech_df['word_splits'].head())
0    1432
1     135
2    2323
3    1736
4    2169
Name: word_cnt, dtype: int64
Feature Engineering for Machine Learning in Python

平均詞長

speech_df['avg_word_len'] = 
         speech_df['char_cnt'] / speech_df['word_cnt']
Feature Engineering for Machine Learning in Python

一起來練習吧!

Feature Engineering for Machine Learning in Python

Preparing Video For Download...