Giới thiệu tiền xử lý văn bản

Deep Learning cho Văn bản với PyTorch

Shubham Jain

Data Scientist

Nội dung học

  • Phân loại văn bản
  • Sinh văn bản
  • Mã hóa
  • Mô hình học sâu cho văn bản
  • Kiến trúc Transformer
  • Bảo vệ mô hình

Trường hợp sử dụng:

  • Phân tích cảm xúc
  • Tóm tắt văn bản
  • Dịch máy

Phân tích cảm xúc

Deep Learning cho Văn bản với PyTorch

Bạn cần biết gì

Học phần tiền đề: Intermediate Deep Learning with PyTorch

  • Mô hình học sâu với PyTorch
  • Vòng lặp huấn luyện và đánh giá
  • Mạng nơ-ron tích chập (CNN) và mạng nơ-ron hồi quy (RNN)
Deep Learning cho Văn bản với PyTorch

Pipeline xử lý văn bản

 

 

Quy trình xử lý PyTorch

Deep Learning cho Văn bản với PyTorch

Pipeline xử lý văn bản

 

 

Quy trình xử lý PyTorch

 

  • Làm sạch và chuẩn bị văn bản
Deep Learning cho Văn bản với PyTorch

PyTorch và NLTK

Logo PyTorch

Logo NLTK

  • Bộ công cụ xử lý ngôn ngữ tự nhiên
    • Biến văn bản thô thành văn bản đã xử lý
Deep Learning cho Văn bản với PyTorch

Kỹ thuật tiền xử lý

  • Tách từ (tokenization)
  • Loại bỏ stop word
  • Stemming
  • Loại bỏ từ hiếm
Deep Learning cho Văn bản với PyTorch

Tách từ (Tokenization)

  • Trích xuất các token hoặc từ từ văn bản
  • Tách từ bằng torchtext
from torchtext.data.utils import get_tokenizer

tokenizer = get_tokenizer("basic_english")
tokens = tokenizer("I am reading a book now. I love to read books!") print(tokens)
["I", "am", "reading", "a", "book", "now", ".", "I", "love", "to", "read", 
"books", "!"]
Deep Learning cho Văn bản với PyTorch

Loại bỏ stop word

  • Loại bỏ các từ phổ biến không mang nhiều nghĩa
  • Stop word: "a", "the", "and", "or", ...
import nltk
nltk.download('stopwords')
from nltk.corpus import stopwords

stop_words = set(stopwords.words('english'))
tokens = ["I", "am", "reading", "a", "book", "now", ".", "I", "love", "to", "read", "books", "!"] filtered_tokens = [token for token in tokens if token.lower() not in stop_words]
print(filtered_tokens)
["reading", "book", ".", "love", "read", "books", "!"]
Deep Learning cho Văn bản với PyTorch

Stemming

  • Rút gọn từ về dạng gốc
  • Ví dụ: "running", "runs", "ran" thành run
import nltk
from nltk.stem import PorterStemmer

stemmer = PorterStemmer()
filtered_tokens = ["reading", "book", ".", "love", "read", "books", "!"]
stemmed_tokens = [stemmer.stem(token) for token in filtered_tokens]
print(stemmed_tokens)
["read", "book", ".", "love", "read", "book", "!"]
Deep Learning cho Văn bản với PyTorch

Loại bỏ từ hiếm

  • Loại bỏ từ ít xuất hiện, ít giá trị
from nltk.probability import FreqDist
stemmed_tokens= ["read", "book", ".", "love", "read", "book", "!"]  
freq_dist = FreqDist(stemmed_tokens)

threshold = 2
common_tokens = [token for token in stemmed_tokens if freq_dist[token] > threshold] print(common_tokens)
["read", "book", "read", "book"]
Deep Learning cho Văn bản với PyTorch

Kỹ thuật tiền xử lý

Tách từ, loại bỏ stop word, stemming và loại bỏ từ hiếm

  • Giảm số lượng đặc trưng
  • Dữ liệu sạch và đại diện hơn
  • Còn nhiều kỹ thuật khác
Deep Learning cho Văn bản với PyTorch

Ayo berlatih!

Deep Learning cho Văn bản với PyTorch

Preparing Video For Download...