Xử lý ngôn ngữ tự nhiên với spaCy
Azadeh Mobasher
Principal Data Scientist
spaCy trước tiên tách từ (tokenize) để tạo đối tượng DocDoc được xử lý qua nhiều bước của pipeline xử lý
import spacy
nlp = spacy.load("en_core_web_sm")
doc = nlp(example_text)
spaCy cho NER:print([ent.text for ent in doc.ents])
sentencizer: thành phần pipeline của spaCy để tách câu.text = " ".join(["This is a test sentence."]*10000)en_core_sm_nlp = spacy.load("en_core_web_sm") start_time = time.time() doc = en_core_sm_nlp(text)print(f"Finished processing with en_core_web_sm model in {round((time.time() - start_time)/60.0 , 5)} minutes")
>>> Finished processing with en_core_web_sm model in 0.09332 minutes
sentencizer:blank_nlp = spacy.blank("en")blank_nlp.add_pipe("sentencizer")start_time = time.time() doc = blank_nlp(text) print(f"Finished processing with blank model in {round((time.time() - start_time)/60.0 , 5)} minutes")
>>> Finished processing with blank model in 0.00091 minutes
nlp.analyze_pipes() phân tích pipeline spaCy để xác định:
pretty=True sẽ in bảng thay vì chỉ trả về dữ liệu có cấu trúc.import spacy
nlp = spacy.load("en_core_web_sm")
analysis = nlp.analyze_pipes(pretty=True)
Xử lý ngôn ngữ tự nhiên với spaCy