探索向量空間

Introduction to Embeddings with the OpenAI API

Emmanuel Pire

Senior Software Engineer, DataCamp

範例:為標題建立嵌入向量

articles = [
    {"headline": "Economic Growth Continues Amid Global Uncertainty", "topic": "Business"},
    {"headline": "Interest rates fall to historic lows", "topic": "Business"},
    {"headline": "Scientists Make Breakthrough Discovery in Renewable Energy", "topic": "Science"},
    {"headline": "India Successfully Lands Near Moon's South Pole", "topic": "Science"},
    {"headline": "New Particle Discovered at CERN", "topic": "Science"},
    {"headline": "Tech Company Launches Innovative Product to Improve Online Accessibility", "topic": "Tech"},
    {"headline": "Tech Giant Buys 49% Stake In AI Startup", "topic": "Tech"},
    {"headline": "New Social Media Platform Has Everyone Talking!", "topic": "Tech"},
    {"headline": "The Blues get promoted on the final day of the season!", "topic": "Sport"},
    {"headline": "1.5 Billion Tune-in to the World Cup Final", "topic": "Sport"}
]
Introduction to Embeddings with the OpenAI API

範例:為標題建立嵌入向量

包含標題及其嵌入向量的字典清單。

Introduction to Embeddings with the OpenAI API

同時為多筆輸入建立嵌入向量

headline_text = [article['headline'] for article in articles]
headline_text
["Economic Growth Continues Amid Global Uncertainty",
 ...,
 "1.5 Billion Tune-in to the World Cup Final"]
response = client.embeddings.create(
  model="text-embedding-3-small",
  input=headline_text
)
response_dict = response.model_dump()
  • 批次處理 比多次呼叫 API 更有效率
Introduction to Embeddings with the OpenAI API
[...]

'data': [
    {
      "embedding": [-0.017142612487077713, ..., -0.0012911480152979493],
      "index": 0,
      "object": "embedding"
    },
    {
      "embedding": [-0.032995883375406265, ..., -0.0028605300467461348],
      "index": 1,
      "object": "embedding"
    },
    ...
  ]

[...]
Introduction to Embeddings with the OpenAI API

同時為多筆輸入建立嵌入向量

articles = [
    {"headline": "Economic Growth Continues Amid Global Uncertainty", "topic": "Business"},
     ...
]
for i, article in enumerate(articles):

article['embedding'] = response_dict['data'][i]['embedding']
print(articles[:2])
[{'headline': 'Economic Growth Continues Amid Global Uncertainty',
  'topic': 'Business',
  'embedding': [-0.017142612487077713, ..., -0.0012911480152979493]}
 {'headline': 'Interest rates fall to historic lows',
  'topic': 'Business',
  'embedding': [-0.032995883375406265, ..., -0.0028605300467461348]}]
Introduction to Embeddings with the OpenAI API

嵌入向量的長度是多少?

  • 「Economic Growth Continues Amid Global Uncertainty」
len(articles[0]['embedding'])
1536
  • 「Tech Company Launches Innovative Product to Improve Accessibility」
len(articles[5]['embedding'])
1536
  • 永遠回傳 1536 個數值!
Introduction to Embeddings with the OpenAI API

降維與 t-SNE

 

  • 有多種技術可用來「降維」
  • t-SNEt-distributed Stochastic Neighbor Embedding)
1 https://www.datacamp.com/tutorial/introduction-t-sne
Introduction to Embeddings with the OpenAI API

實作 t-SNE

from sklearn.manifold import TSNE
import numpy as np


embeddings = [article['embedding'] for article in articles]
tsne = TSNE(n_components=2, perplexity=5)
embeddings_2d = tsne.fit_transform(np.array(embeddings))
  • n_components:結果維度數
  • perplexity:演算法用的參數,必須小於資料點數
  • 會造成部分資訊流失
1 https://www.datacamp.com/tutorial/introduction-t-sne
Introduction to Embeddings with the OpenAI API

視覺化嵌入向量

import matplotlib.pyplot as plt

plt.scatter(embeddings_2d[:, 0], embeddings_2d[:, 1])

topics = [article['topic'] for article in articles] for i, topic in enumerate(topics): plt.annotate(topic, (embeddings_2d[i, 0], embeddings_2d[i, 1])) plt.show()
Introduction to Embeddings with the OpenAI API

視覺化嵌入向量

 

  • 類似的文章會被聚在一起!
  • 模型捕捉到語意

 

  • 接下來:計算相似度

 

2D 向量空間圖,顯示相同情感與主題的評論在向量空間中更靠近。

Introduction to Embeddings with the OpenAI API

一起來練習吧!

Introduction to Embeddings with the OpenAI API

Preparing Video For Download...