การประมวลผลภาษาธรรมชาติ (NLP) ด้วย Python
Fouad Trad
Machine Learning Engineer

การทำความเข้าใจหัวข้อของข้อความ

การทำความเข้าใจหัวข้อของข้อความ

งานที่ต้องใช้ทุกคำในข้อความ

NLTK มีรายการ stop words สำหรับหลายภาษา
from nltk.corpus import stopwords nltk.download('stopwords')stop_words = stopwords.words('english')print(stop_words[:10])
['a', 'about', 'above', 'after', 'again', 'against', 'ain', 'all', 'am', 'an']
from nltk.tokenize import word_tokenizetext = "This is an example to demonstrate removing stop words."tokens = word_tokenize(text)# The .lower() method helps with case sensitivity filtered_tokens = [word for word in tokens if word.lower() not in stop_words]print(filtered_tokens)
['example', 'demonstrate', 'removing', 'stop', 'words', '.']

งานที่ต้องการค้นหาคำที่พบบ่อยหรือคำสำคัญในเอกสาร

งานที่ต้องการค้นหาคำที่พบบ่อยหรือคำสำคัญในเอกสาร

งานที่ต้องรักษาโครงสร้างประโยคเพื่อความชัดเจน

import string
print(string.punctuation)
!"#$%&'()*+,-./:;<=>?@[\]^_`{|}~
text = "This is an example to demonstrate removing stop words." tokens = word_tokenize(text) filtered_tokens = [word for word in tokens if word.lower() not in stop_words]clean_tokens = [word for word in filtered_tokens if word not in string.punctuation]print(clean_tokens)
['example', 'demonstrate', 'removing', 'stop', 'words']
การประมวลผลภาษาธรรมชาติ (NLP) ด้วย Python