Chuẩn bị văn bản cho mô hình hóa

Nhập môn Xử lý Ngôn ngữ Tự nhiên với R

Kasey Jones

Research Data Scientist

Học có giám sát trong R: phân loại

Nhập môn Xử lý Ngôn ngữ Tự nhiên với R

Mô hình phân loại

  • phương pháp học có giám sát
  • phân loại quan sát vào các nhóm
    • thắng/thua
    • nguy hiểm, thân thiện, hay thờ ơ
  • có thể dùng nhiều kỹ thuật:
    • hồi quy logistic
    • cây quyết định/rừng ngẫu nhiên/xgboost
    • mạng nơ-ron
    • v.v.
Nhập môn Xử lý Ngôn ngữ Tự nhiên với R

Các bước cơ bản để mô hình hóa

  1. Làm sạch/chuẩn bị dữ liệu
  2. Tạo tập huấn luyện và kiểm tra
  3. Huấn luyện mô hình trên tập huấn luyện
  4. Báo cáo độ chính xác trên tập kiểm tra
Nhập môn Xử lý Ngôn ngữ Tự nhiên với R

Nhận diện nhân vật

Napoloeon Napoleon

Boxer Boxer

1 https://comicvine.gamespot.com/napoleon/4005-141035/ 2 https://hero.fandom.com/wiki/Boxer_(Animal_Farm)
Nhập môn Xử lý Ngôn ngữ Tự nhiên với R

Câu có động vật

# Make sentences
sentences <- animal_farm %>%
  unnest_tokens(output = "sentence", token = "sentences", input = text_column)
# Label sentences by animal
sentences$boxer <- grepl('boxer', sentences$sentence)
sentences$napoleon <- grepl('napoleon', sentences$sentence)
# Replace the animal name
sentences$sentence <- gsub("boxer", "animal X", sentences$sentence)
sentences$sentence <- gsub("napoleon", "animal X", sentences$sentence)
animal_sentences <- sentences[sentences$boxer + sentences$napoleon == 1, ]
Nhập môn Xử lý Ngôn ngữ Tự nhiên với R

Tiếp tục về câu

animal_sentences$Name <-
    as.factor(ifelse(animal_sentences$boxer, "boxer", "napoleon"))
# 75 of each
animal_sentences <- 
  rbind(animal_sentences[animal_sentences$Name == "boxer", ][c(1:75), ],
        animal_sentences[animal_sentences$Name == "napoleon", ][c(1:75), ])
animal_sentences$sentence_id <- c(1:dim(animal_sentences)[1])
Nhập môn Xử lý Ngôn ngữ Tự nhiên với R

Chuẩn bị dữ liệu

library(tm); library(tidytext)
library(dplyr); library(SnowballC)
animal_tokens <- animal_sentences %>%
  unnest_tokens(output = "word", token = "words", input = sentence) %>%
  anti_join(stop_words) %>%
  mutate(word = wordStem(word))
Nhập môn Xử lý Ngôn ngữ Tự nhiên với R

Chuẩn bị (tiếp)

animal_matrix <- animal_tokens %>%
  count(sentence_id, word) %>%
  cast_dtm(document = sentence_id, term = word,
           value = n, weighting = tm::weightTfIdf)
animal_matrix
<<DocumentTermMatrix (documents: 150, terms: 694)>>
Non-/sparse entries: 1235/102865
Sparsity           : 99%
Maximal term length: 17
Weighting          : term frequency - inverse document frequency
Nhập môn Xử lý Ngôn ngữ Tự nhiên với R

Loại bỏ thuật ngữ thưa

  • Có giá trị (1.235) + rỗng (102.865)
  • Kích thước ma trận 150 × 694
  • Độ thưa: 102.865 / 104.100 (99%)

Giải pháp: removeSparseTerms()

Nhập môn Xử lý Ngôn ngữ Tự nhiên với R

Bao nhiêu là quá thưa?

removeSparseTerms(animal_matrix, sparse = .90)
<<DocumentTermMatrix (documents: 150, terms: 4)>>
Non-/sparse entries: 207/393
Sparsity           : 66%
removeSparseTerms(animal_matrix, sparse = .99)
removeSparseTerms(animal_matrix, sparse = .99)
<<DocumentTermMatrix (documents: 150, terms: 172)>>
Non-/sparse entries: 713/25087
Sparsity           : 97%
Nhập môn Xử lý Ngôn ngữ Tự nhiên với R

Ayo berlatih!

Nhập môn Xử lý Ngôn ngữ Tự nhiên với R

Preparing Video For Download...