TFIDF

R 自然语言处理入门

Kasey Jones

Research Data Scientist

词袋法的陷阱

t1 <- "My name is John. My best friend is Joe. We like tacos."
t2 <- "Two common best friend names are John and Joe."
t3 <- "Tacos are my favorite food. I eat them with my buddy Joe."
clean_t1 <- "john friend joe tacos"
clean_t2 <- "common friend john joe names"
clean_t3 <- "tacos favorite food eat buddy joe"
R 自然语言处理入门

共享常见词

clean_t1 <- "john friend joe tacos"
clean_t2 <- "common friend john joe names"
clean_t3 <- "tacos favorite food eat buddy joe"

比较 t1 与 t2

  • t1 中有 3/4 的词出现在 t2 中
  • t2 中有 3/5 的词出现在 t1 中

比较 t1 与 t3

  • t1 中有 2/4 的词出现在 t3 中
  • t3 中有 2/6 的词出现在 t1 中
R 自然语言处理入门

Tacos 很关键

t1 <- "My name is John. My best friend is Joe. We like tacos."
t2 <- "Two common best friend names are John and Joe."
t3 <- "Tacos are my favorite food. I eat them with my friend Joe."

各文本中的词:

  • John:t1, t2
  • Joe:t1, t2, t3
  • Tacos:t1, t3
R 自然语言处理入门

TFIDF

clean_t1 <- "john friend joe tacos"
clean_t2 <- "common friend john joe names"
clean_t3 <- "tacos favorite food eat buddy joe"
  • TF:词频
    • 某词在文本中所占比例
    • john 占 clean_t1 的 1/4,tf = .25
  • IDF:逆文档频率
    • 反映该词在所有文档中的普遍程度
    • john 出现在 3/3 个文档中,IDF = 0
R 自然语言处理入门

IDF 公式

 

$ IDF = log \frac{N}{n_{t}} $

  • N:语料库中的文档总数
  • $n_{t}$:包含该词的文档数

示例:

  • Taco 的 IDF:$log (\frac{3}{2}) = .405$
  • Buddy 的 IDF:$log (\frac{3}{1}) = 1.10$
  • John 的 IDF:$log (\frac{3}{3}) = 0$
R 自然语言处理入门

TF + IDF

clean_t1 <- "john friend joe tacos"
clean_t2 <- "common friend john joe names"
clean_t3 <- "tacos favorite food eat buddy joe"

"tacos"的TFIDF:

  • clean_t1:TF * IDF = (1/4) * (.405) = 0.101
  • clean_t2:TF * IDF = (0/4) * (.405) = 0
  • clean_t3:TF * IDF = (1/6) * (.405) = 0.068
R 自然语言处理入门

计算TFIDF矩阵

# Create a data.frame
df <- data.frame('text' = c(t1, t2, t3), 'ID' = c(1, 2, 3))
df %>%
  unnest_tokens(output = "word", token = "words", input = text) %>%
  anti_join(stop_words) %>%
  count(ID, word, sort = TRUE) %>%
  bind_tf_idf(word, ID, n)
  • word:包含术语的列
  • ID:包含文档ID的列
  • n:count() 生成的词频
R 自然语言处理入门

bind_tf_idf 输出

# A tibble: 15 x 6
       X word         n    tf   idf tf_idf
   <dbl> <chr>    <int> <dbl> <dbl>  <dbl>
 1     1 friend       1 0.25  0.405 0.101 
 2     1 joe          1 0.25  0     0     
 3     1 john         1 0.25  0.405 0.101 
 4     1 tacos        1 0.25  0.405 0.101 
 5     2 common       1 0.2   1.10  0.220 
 6     2 friend       1 0.2   0.405 0.0811
 ...
R 自然语言处理入门

TFIDF 练习

R 自然语言处理入门

Preparing Video For Download...