使用 spaCy 的自然語言處理
Azadeh Mobasher
Principal Data Scientist
EntityRuler 會把命名實體加入 Doc 容器EntityRecognizer 搭配{"label": "ORG", "pattern": "Microsoft"}
{"label": "GPE", "pattern": [{"LOWER": "san"}, {"LOWER": "francisco"}]}
.add_pipe() 方法.add_patterns() 方法加入樣式清單
nlp = spacy.blank("en")
entity_ruler = nlp.add_pipe("entity_ruler")
patterns = [{"label": "ORG", "pattern": "Microsoft"},
{"label": "GPE", "pattern": [{"LOWER": "san"}, {"LOWER": "francisco"}]}]
entity_ruler.add_patterns(patterns)
.ents 會儲存 EntityLinker 元件的結果
doc = nlp("Microsoft is hiring software developer in San Francisco.")
print([(ent.text, ent.label_) for ent in doc.ents])
[("Microsoft", "ORG"), ("San Francisco", "GPE")]
spaCy 的 pipeline 元件整合強化命名實體辨識器
沒有 EntityRuler 的 spaCy 模型:
nlp = spacy.load("en_core_web_sm")
doc = nlp("Manhattan associates is a company in the U.S.")
print([(ent.text, ent.label_) for ent in doc.ents])
>>> [("Manhattan", "GPE"), ("U.S.", "GPE")]
EntityRuler 加在現有 ner 元件之後:nlp = spacy.load("en_core_web_sm")
ruler = nlp.add_pipe("entity_ruler", after='ner')
patterns = [{"label": "ORG", "pattern": [{"lower": "manhattan"}, {"lower": "associates"}]}]
ruler.add_patterns(patterns)
doc = nlp("Manhattan associates is a company in the U.S.")
print([(ent.text, ent.label_) for ent in doc.ents])
>>> [("Manhattan", "GPE"), ("U.S.", "GPE")]
EntityRuler 加在現有 ner 元件之前:nlp = spacy.load("en_core_web_sm")
ruler = nlp.add_pipe("entity_ruler", before='ner')
patterns = [{"label": "ORG", "pattern": [{"lower": "manhattan"}, {"lower": "associates"}]}]
ruler.add_patterns(patterns)
doc = nlp("Manhattan associates is a company in the U.S.")
print([(ent.text, ent.label_) for ent in doc.ents])
>>> [("Manhattan associates", "ORG"), ("U.S.", "GPE")]
使用 spaCy 的自然語言處理