Sémantické vyhledávání a obohacená vnoření

Introduction to Embeddings with the OpenAI API

Emmanuel Pire

Senior Software Engineer, DataCamp

Sémantické vyhledávání

  • Vnoření slouží k nalezení nejpodobnějších výsledků pro vyhledávací dotaz
  • Příklad: Sémantické vyhledávání pro zpravodajský web

Jak funguje sémantické vyhledávání: vyhledávaný text je vložen do modelu vnoření, poté je vyhodnocena vzdálenost mezi vložením hledaného textu a vložením nadpisů. Nejbližší nadpis je vrácen.

Introduction to Embeddings with the OpenAI API

Sémantické vyhledávání

Jak funguje sémantické vyhledávání: vyhledávaný text je vložen do modelu vnoření, poté je vyhodnocena vzdálenost mezi vložením hledaného textu a vložením nadpisů. Nejbližší nadpis je vrácen.

  1. Vložení vyhledávacího dotazu a ostatních textů
  2. Výpočet kosinových vzdáleností
  3. Extrakce textů s nejmenší kosinovou vzdáleností
Introduction to Embeddings with the OpenAI API

Obohacená vnoření

articles = [
    {"headline": "Economic Growth Continues Amid Global Uncertainty",
     "topic": "Business",
     "keywords": ["economy", "business", "finance"]},
    ...
    {"headline": "1.5 Billion Tune-in to the World Cup Final",
     "topic": "Sport",
     "keywords": ["soccer", "world cup", "tv"]}
]
Headline: Economic Growth Continues Amid Global Uncertainty
Topic: Business
Keywords: economy, business, finance
Introduction to Embeddings with the OpenAI API

Kombinování atributů pomocí f-řetězců

articles = [..., {"headline": "1.5 Billion Tune-in to the World Cup ",
                  "topic": "Sport",
                  "keywords": ["soccer", "world cup", "tv"]}]


def create_article_text(article):
return f"""Headline: {article['headline']} Topic: {article['topic']} Keywords: {', '.join(article['keywords'])}"""
print(create_article_text(articles[-1]))
Headline: 1.5 Billion Tune-in to the World Cup Final
Topic: Sport
Keywords: soccer, world cup, tv
Introduction to Embeddings with the OpenAI API

Vytváření obohacených vnoření

article_texts = [create_article_text(article) for article in articles]

article_embeddings = create_embeddings(article_texts)
print(article_embeddings)
[[-0.019609929993748665, -0.03331860154867172, ...],
 ...,
 [..., -0.014373429119586945, -0.005235843360424042]]
Introduction to Embeddings with the OpenAI API

Výpočet vzdáleností

from scipy.spatial import distance

def find_n_closest(query_vector, embeddings, n=3):

distances = [] for index, embedding in enumerate(embeddings): dist = distance.cosine(query_vector, embedding) distances.append({"distance": dist, "index": index})
distances_sorted = sorted(distances, key=lambda x: x["distance"])
return distances_sorted[0:n]
Introduction to Embeddings with the OpenAI API

Vrácení výsledků vyhledávání

query_text = "AI"

query_vector = create_embeddings(query_text)[0]
hits = find_n_closest(query_vector, article_embeddings)
for hit in hits: article = articles[hit['index']] print(article['headline'])
Tech Giant Buys 49% Stake In AI Startup
Tech Company Launches Innovative Product to Improve Online Accessibility
India Successfully Lands Near Moon's South Pole
Introduction to Embeddings with the OpenAI API

Lass uns üben!

Introduction to Embeddings with the OpenAI API

Preparing Video For Download...