超越单个词语

使用 R 的 Bag-of-Words 进行文本挖掘

Ted Kwartler

Instructor

Unigram、Bigram、Trigram,全都要!

# 仅使用前 2 条关于 coffee 的推文
tweets$text[1:2]
[1] @ayyytylerb that is so true drink lots of coffee
[2] RT @bryzy_brib: Senior March tmw morning at 7:25 A.M. in the SENIOR lot. Get up early, make yo coffee/breakfast, cus this will only happen…
# 对前 2 条推文创建 unigram DTM
unigram_dtm <- DocumentTermMatrix(text_corp)
unigram_dtm
<<DocumentTermMatrix (documents: 2, terms: 18)>>
Non-/sparse entries: 18/18
Sparsity           : 50%
Maximal term length: 15
Weighting          : term frequency (tf)
使用 R 的 Bag-of-Words 进行文本挖掘

Unigram、Bigram、Trigram,全都要!

# 加载 RWeka 包
library(RWeka)
# 定义 bigram 分词器
tokenizer <- function(x) NGramTokenizer(x, Weka_control(min = 2, max = 2))

# 创建 bigram TDM bigram_tdm <- TermDocumentMatrix(clean_corpus(text_corp), control = list(tokenize = tokenizer)) bigram_tdm
<<DocumentTermMatrix (documents: 2, terms: 21)>>
Non-/sparse entries: 21/21
Sparsity           : 50%
Maximal term length: 19
Weighting          : term frequency (tf)
使用 R 的 Bag-of-Words 进行文本挖掘

让我们练习吧!

使用 R 的 Bag-of-Words 进行文本挖掘

Preparing Video For Download...