Giới thiệu về mô hình ngôn ngữ

Mạng nơ-ron hồi quy (RNN) cho Mô hình ngôn ngữ với Keras

David Cecchini

Data Scientist

Xác suất câu

Nhiều mô hình sẵn có

  • Xác suất của "I loved this movie".
  • Unigram
    • $$P(\text{sentence}) = P(\text{I})P(\text{loved})P(\text{this})P(\text{movie})$$
  • N-gram
    • N = 2 (bigram): $$P(\text{sentence}) = P(\text{I})P(\text{loved} | \text{I})P(\text{this} | \text{loved})P(\text{movie} | \text{this})$$
    • N = 3 (trigram): $$P(\text{sentence}) = P(\text{I})P(\text{loved} | \text{I})P(\text{this} | \text{I loved})P(\text{movie} | \text{loved this})$$
Mạng nơ-ron hồi quy (RNN) cho Mô hình ngôn ngữ với Keras

Xác suất câu (tiếp)

  • Skip-gram
    • $$P(\text{sentence}) = P(\text{context of I} | \text{I})P(\text{context of loved} | \text{loved}) \ $$ $$P(\text{context of this} | \text{this})P(\text{context of movie} | \text{movie})$$
  • Mạng nơ-ron
    • Xác suất câu được cho bởi hàm softmax ở lớp đầu ra của mạng
Mạng nơ-ron hồi quy (RNN) cho Mô hình ngôn ngữ với Keras

Liên hệ với RNN

Mô hình ngôn ngữ hiện diện khắp nơi trong RNN!

  • Bản thân mạng

Các mô hình RNN có thể coi là mô hình ngôn ngữ vì chúng dự đoán từ kế tiếp.

Mạng nơ-ron hồi quy (RNN) cho Mô hình ngôn ngữ với Keras

Liên hệ với RNN (tiếp)

  • Lớp embedding

Hiển thị sơ đồ các lớp của mô hình. Lớp embedding phải là lớp đầu sau lớp đầu vào và tạo biểu diễn dày đặc cho từ.

Mạng nơ-ron hồi quy (RNN) cho Mô hình ngôn ngữ với Keras

Xây dựng từ điển từ vựng

# Get unique words
unique_words = list(set(text.split(' ')))
# Create dictionary: word is key, index is value
word_to_index = {k:v for (v,k) in enumerate(unique_words)}
# Create dictionary: index is key, word is value
index_to_word = {k:v for (k,v) in enumerate(unique_words)}
Mạng nơ-ron hồi quy (RNN) cho Mô hình ngôn ngữ với Keras

Tiền xử lý đầu vào

# Initialize variables X and y
X = []
y = []

# Loop over the text: length `sentence_size` per time with step equal to `step` for i in range(0, len(text) - sentence_size, step):
X.append(text[i:i + sentence_size]) y.append(text[i + sentence_size])
# Example (numbers are numerical indexes of vocabulary):
# Sentence is: "i loved this movie" -> (["i", "loved", "this"], "movie")
X[0],y[0] = ([10, 444, 11], 17)
Mạng nơ-ron hồi quy (RNN) cho Mô hình ngôn ngữ với Keras

Chuyển đổi văn bản mới

# Create list to keep the sentences of indexes
new_text_split = []

# Loop and get the indexes from dictionary for sentence in new_text:
sent_split = []
for wd in sentence.split(' '):
ix = wd_to_index[wd]
sent_split.append(ix)
new_text_split.append(sent_split)
Mạng nơ-ron hồi quy (RNN) cho Mô hình ngôn ngữ với Keras

Ayo berlatih!

Mạng nơ-ron hồi quy (RNN) cho Mô hình ngôn ngữ với Keras

Preparing Video For Download...