Nhận dạng thực thể có tên trên văn bản chuyển âm

Xử lý Ngôn ngữ Nói bằng Python

Daniel Bourke

Machine Learning Engineer/YouTube Creator

Cài đặt spaCy

# Install spaCy
$ pip install spacy
# Download spaCy language model
$ python -m spacy download en_core_web_sm
Xử lý Ngôn ngữ Nói bằng Python

Sử dụng spaCy

import spacy

# Load spaCy language model nlp = spacy.load("en_core_web_sm")
# Create a spaCy doc
doc = nlp("I'd like to talk about a smartphone I ordered on July 31st from your 
Sydney store, my order number is 40939440. I spoke to Georgia about it last week.")
Xử lý Ngôn ngữ Nói bằng Python

Token trong spaCy

# Show different tokens and positions
for token in doc:
  print(token.text, token.idx)
I 0
'd 1
like 4
to 9
talk 12
about 17
a 23
smartphone 25...
Xử lý Ngôn ngữ Nói bằng Python

Câu trong spaCy

# Show sentences in doc
for sentences in doc.sents:
  print(sentence)
I'd like to talk about a smartphone I ordered on July 31st from your Sydney store, 
my order number is 4093829.
I spoke to one of your customer service team, Georgia, yesterday.
Xử lý Ngôn ngữ Nói bằng Python

Thực thể có tên trong spaCy

Một số loại thực thể có tên tích hợp của spaCy:

  • PERSON Con người, gồm cả hư cấu.
  • ORG Công ty, cơ quan, tổ chức, v.v.
  • GPE Quốc gia, thành phố, bang/tỉnh.
  • PRODUCT Vật thể, phương tiện, thực phẩm, v.v. (Không gồm dịch vụ.)
  • DATE Ngày hoặc khoảng thời gian tuyệt đối/tương đối.
  • TIME Mốc thời gian nhỏ hơn một ngày.
  • MONEY Giá trị tiền tệ, gồm đơn vị.
  • CARDINAL Số không thuộc loại khác.
Xử lý Ngôn ngữ Nói bằng Python

Thực thể có tên trong spaCy

# Find named entities in doc
for entity in doc.ents:
  print(entity.text, entity.label_)
July 31st DATE
Sydney GPE
4093829 CARDINAL
one CARDINAL
Georgia GPE
yesterday DATE
Xử lý Ngôn ngữ Nói bằng Python

Tùy chỉnh thực thể có tên

# Import EntityRuler class
from spacy.pipeline import EntityRuler
# Check spaCy pipeline
print(nlp.pipeline)
[('tagger', <spacy.pipeline.pipes.Tagger at 0x1c3aa8a470>),
 ('parser', <spacy.pipeline.pipes.DependencyParser at 0x1c3bb60588>),
 ('ner', <spacy.pipeline.pipes.EntityRecognizer at 0x1c3bb605e8>)]
Xử lý Ngôn ngữ Nói bằng Python

Thay đổi pipeline

# Create EntityRuler instance
ruler = EntityRuler(nlp)
# Add token pattern to ruler
ruler.add_patterns([{"label":"PRODUCT", "pattern": "smartphone"}])
# Add new rule to pipeline before ner
nlp.add_pipe(ruler, before="ner")
# Check updated pipeline
nlp.pipeline
Xử lý Ngôn ngữ Nói bằng Python

Thay đổi pipeline

[('tagger', <spacy.pipeline.pipes.Tagger at 0x1c1f9c9b38>),
 ('parser', <spacy.pipeline.pipes.DependencyParser at 0x1c3c9cba08>),
 ('entity_ruler', <spacy.pipeline.entityruler.EntityRuler at 0x1c1d834b70>),
 ('ner', <spacy.pipeline.pipes.EntityRecognizer at 0x1c3c9cba68>)]
Xử lý Ngôn ngữ Nói bằng Python

Kiểm tra pipeline mới

# Test new entity rule
for entity in doc.ents:
    print(entity.text, entity.label_)
smartphone PRODUCT
July 31st DATE
Sydney GPE
4093829 CARDINAL
one CARDINAL
Georgia GPE
yesterday DATE
Xử lý Ngôn ngữ Nói bằng Python

Hãy bắt tay vào luyện tập spaCy!

Xử lý Ngôn ngữ Nói bằng Python

Preparing Video For Download...