Phần 1: Tiền xử lý dữ liệu

Machine Translation với Keras

Thushan Ganegedara

Data Scientist and Author

Giới thiệu dữ liệu

  • Dữ liệu

    • en_text: Danh sách câu tiếng Anh; mỗi câu là chuỗi từ cách nhau bằng dấu cách.
    • fr_text: Danh sách câu tiếng Pháp; mỗi câu là chuỗi từ cách nhau bằng dấu cách.
  • In một số mẫu trong tập dữ liệu

for en_sent, fr_sent in zip(en_text[:3], fr_text[:3]):
  print("English: ", en_sent)
  print("\tFrench: ", fr_sent)
English:  new jersey is sometimes quiet during autumn , and it is snowy in april .
    French:  new jersey est parfois calme pendant l' automne , et il est neigeux en avril .
English:  the united states is usually chilly during july , and it is usually freezing in november .
    French:  les états-unis est généralement froid en juillet , et il gèle habituellement en novembre .
...
Machine Translation với Keras

Tách từ

  • Tách từ (Tokenization)

    • Quá trình tách câu/cụm thành các từ/ký tự riêng lẻ
    • Ví dụ: "I watched a movie last night, it was okay." thành
    • [I, watched, a, movie, last, night, it, was, okay]
  • Tách từ với Keras

    • Học ánh xạ từ → ID từ bằng một corpus cho trước.
    • Dùng để chuyển chuỗi văn bản thành dãy ID
from tensorflow.keras.preprocessing.text import Tokenizer
en_tok = Tokenizer()
Machine Translation với Keras

Fit Tokenizer

  • Fit Tokenizer trên dữ liệu
    • Cần fit trên một số câu để học ánh xạ từ ↔ ID.
en_tok = Tokenizer()
en_tok.fit_on_texts(en_text)
  • Lấy ánh xạ từ → ID
    • Dùng thuộc tính word_index của Tokenizer.
id = en_tok.word_index["january"] # => returns 51
  • Lấy ánh xạ ID → từ
w = en_tok.index_word[51] # => returns 'january'
Machine Translation với Keras

Chuyển câu thành chuỗi ID

seq = en_tok.texts_to_sequences(['she likes grapefruit , peaches , and lemons .'])
[[26, 70, 27, 73, 7, 74]]
Machine Translation với Keras

Giới hạn kích thước từ vựng

  • Có thể giới hạn kích thước từ vựng trong Keras Tokenizer.
tok = Tokenizer(num_words=50)
  • Từ ngoài từ vựng (OOV)

    • Từ hiếm trong corpus huấn luyện.
    • Từ không xuất hiện trong tập huấn luyện.
  • Ví dụ

    • tok.fit_on_texts(["I drank milk"])
    • tok.texts_to_sequences(["I drank water"])
    • Từ water là OOV và sẽ bị bỏ qua.
Machine Translation với Keras

Xử lý từ ngoài từ vựng (OOV)

  • Định nghĩa token OOV
tok = Tokenizer(num_words=50, oov_token='UNK')
  • Ví dụ
    • tok.fit_on_texts(["I drank milk"])
    • tok.texts_to_sequences(["I drank water"])
    • Từ water là OOV và sẽ được thay bằng UNK.
      • tức là Keras sẽ thấy "I drank UNK"
Machine Translation với Keras

Luyện tập nào!

Machine Translation với Keras

Preparing Video For Download...