Mô hình phân loại

Nhập môn Xử lý Ngôn ngữ Tự nhiên với R

Kasey Jones

Research Data Scientist

Tóm tắt các bước

  1. Làm sạch/chuẩn bị dữ liệu

    • Lọc câu của Boxer/Napoleon
    • Tạo token đã làm sạch
    • Tạo ma trận tài liệu–thuật ngữ với trọng số TFIDF
  2. Tạo tập huấn luyện và kiểm tra

  3. Huấn luyện mô hình trên tập huấn luyện
  4. Báo cáo độ chính xác trên tập kiểm tra
Nhập môn Xử lý Ngôn ngữ Tự nhiên với R

Bước 2: chia dữ liệu

set.seed(1111)
sample_size <- floor(0.80 * nrow(animal_matrix))
train_ind <- sample(nrow(animal_matrix), size = sample_size)
train <- animal_matrix[train_ind, ]
test <- animal_matrix[-train_ind, ]
Nhập môn Xử lý Ngôn ngữ Tự nhiên với R

Mô hình rừng ngẫu nhiên

Nhập môn Xử lý Ngôn ngữ Tự nhiên với R

Ví dụ phân loại

library(randomForest)
rfc <- randomForest(x = as.data.frame(as.matrix(train)), 
                    y = animal_sentences$Name[train_ind], nTree = 50)
rfc
Call:
 randomForest(...
        OOB estimate of  error rate: 23.33%
Confusion matrix:
         boxer napoleon class.error
boxer       37       20   0.3508772
napoleon     8       55   0.1269841
Nhập môn Xử lý Ngôn ngữ Tự nhiên với R

Ma trận nhầm lẫn

Call:
 randomForest(...
        OOB estimate of  error rate: 23.33%
Confusion matrix:
         boxer napoleon class.error
boxer       37       20   0.3508772
napoleon     8       55   0.1269841

Độ chính xác: (37 + 55) / (37 + 20 + 8 + 55) = 76%

Nhập môn Xử lý Ngôn ngữ Tự nhiên với R

Dự đoán trên tập kiểm tra

y_pred <- predict(rfc, newdata = as.data.frame(as.matrix(test)))
table(animal_sentences[-train_ind, ]$Name, y_pred)
          y_pred
           boxer napoleon
  boxer       14        4
  napoleon     2       10
  • Độ chính xác cho boxer: 14/18
  • Độ chính xác cho napoleon: 10/12
  • Độ chính xác tổng thể: 24/30 = 80%
Nhập môn Xử lý Ngôn ngữ Tự nhiên với R

Luyện tập phân loại

Nhập môn Xử lý Ngôn ngữ Tự nhiên với R

Preparing Video For Download...