Phân loại ảnh zero-shot

Mô hình đa phương thức với Hugging Face

James Chapman

Curriculum Manager, DataCamp

CLIP

  • Contrastive Language-Image Pre-training
  • Chấm điểm độ giống nhau giữa ảnh và văn bản
  • Huấn luyện trên 400M cặp ảnh–văn bản
  • Hai bộ mã hóa:
    • Bộ mã hóa ảnh
    • Bộ mã hóa văn bản
  • Cặp ảnh–văn bản khớp chặt cho mảng tương tự nhau

Sơ đồ mã hóa văn bản và ảnh của CLIP

1 https://openai.com/index/clip/
Mô hình đa phương thức với Hugging Face

Học zero-shot

  • Thực hiện tác vụ mà mô hình chưa được huấn luyện trực tiếp

Xếp hạng zero-shot của một máy bay

1 https://openai.com/index/clip/
Mô hình đa phương thức với Hugging Face

Tác vụ: phân loại sản phẩm

from datasets import load_dataset
import matplotlib.pyplot as plt

dset = "rajuptvs/ecommerce_products_clip"
dataset = load_dataset(dset)

print(dataset["train"][0]["Description"])
plt.imshow(dataset["train"][0]["image"]) plt.show()
Blive High quality premium Full sleeves printed 
Shirt direct from the manufacturers.Gives you 
a clean and classy look while also 
making you feel comfortable.Trusted 
brand online and no compromise on quality.

Ảnh một chiếc áo sơ mi từ bộ dữ liệu

Mô hình đa phương thức với Hugging Face

Học zero-shot với CLIP

model = CLIPModel.from_pretrained("openai/clip-vit-base-patch32")
processor = CLIPProcessor.from_pretrained("openai/clip-vit-base-patch32")


categories = ["shirt", "trousers", "shoes", "dress", "hat", "bag", "watch"]
inputs = processor(text=categories, images=dataset["train"][0]["image"], return_tensors="pt", padding=True) outputs = model(**inputs)
probs = outputs.logits_per_image.softmax(dim=1)
categories[probs.argmax().item()]
shirt
Mô hình đa phương thức với Hugging Face

Điểm CLIP

  • Độ tương đồng giữa ảnh đã mã hóa và mô tả đã mã hóa
  • Thang từ 100 (phù hợp hoàn hảo) đến 0 (không phù hợp)
from torchmetrics.functional.multimodal import clip_score


image = dataset["train"][0]["image"] description = dataset["train"][0]["Description"]
from torchvision.transforms import ToTensor image = ToTensor()(image)*255
score = clip_score(image, description, "openai/clip-vit-base-patch32")
print(f"CLIP score: {score}")
CLIP score: 28.495952606201172
Mô hình đa phương thức với Hugging Face

Ayo berlatih!

Mô hình đa phương thức với Hugging Face

Preparing Video For Download...