Procesarea Limbajului Natural (NLP) în Python
Fouad Trad
Machine Learning Engineer

Înțelegerea subiectului unui text

Înțelegerea subiectului unui text

Sarcini care necesită fiecare cuvânt din text

NLTK oferă o listă de cuvinte de oprire pentru mai multe limbi
from nltk.corpus import stopwords nltk.download('stopwords')stop_words = stopwords.words('english')print(stop_words[:10])
['a', 'about', 'above', 'after', 'again', 'against', 'ain', 'all', 'am', 'an']
from nltk.tokenize import word_tokenizetext = "This is an example to demonstrate removing stop words."tokens = word_tokenize(text)# The .lower() method helps with case sensitivity filtered_tokens = [word for word in tokens if word.lower() not in stop_words]print(filtered_tokens)
['example', 'demonstrate', 'removing', 'stop', 'words', '.']

Sarcini care necesită identificarea cuvintelor frecvente sau importante din documente

Sarcini care necesită identificarea cuvintelor frecvente sau importante din documente

Sarcini care necesită menținerea structurii propoziției pentru claritate

import string
print(string.punctuation)
!"#$%&'()*+,-./:;<=>?@[\]^_`{|}~
text = "This is an example to demonstrate removing stop words." tokens = word_tokenize(text) filtered_tokens = [word for word in tokens if word.lower() not in stop_words]clean_tokens = [word for word in filtered_tokens if word not in string.punctuation]print(clean_tokens)
['example', 'demonstrate', 'removing', 'stop', 'words']
Procesarea Limbajului Natural (NLP) în Python