Bảo vệ LLMs

Nhập môn LLMs trong Python

Jasmin Ludolf

Senior Data Science Content Developer, DataCamp

Thách thức của LLM

Hỗ trợ đa ngôn ngữ: đa dạng ngôn ngữ, sẵn có tài nguyên, khả năng thích ứng

Hỗ trợ đa ngôn ngữ

Tiến thoái lưỡng nan: LLM mở vs đóng: hợp tác vs sử dụng có trách nhiệm

LLM mở vs đóng

Khả năng mở rộng mô hình: năng lực biểu diễn, nhu cầu tính toán, yêu cầu huấn luyện

Khả năng mở rộng LLM

Thiên lệch: dữ liệu huấn luyện thiên lệch, hiểu và sinh ngôn ngữ thiếu công bằng

Thiên lệch trong LLMs

1 Biểu tượng do Freepik (freepik.com) tạo
Nhập môn LLMs trong Python

Tính đúng sự thật và ảo giác

  • Ảo giác (Hallucinations): văn bản sinh ra có thông tin sai hoặc vô nghĩa nhưng thể hiện như đúng

Ảo giác trong LLMs

Chiến lược giảm ảo giác của LLM:

  1. Tiếp xúc với dữ liệu huấn luyện đa dạng, đại diện
  2. Kiểm toán thiên lệch trên đầu ra + kỹ thuật khử thiên lệch
  3. Fine-tune cho các trường hợp nhạy cảm
  4. Thiết kế prompt: soạn và tinh chỉnh prompt cẩn thận
Nhập môn LLMs trong Python

Tính đúng sự thật và ảo giác

  • Ảo giác (Hallucinations): văn bản sinh ra có thông tin sai hoặc vô nghĩa nhưng thể hiện như đúng

Ảo giác trong LLMs

Nhập môn LLMs trong Python

Chỉ số phân tích thiên lệch LLM: độc hại

  • Độc hại (Toxicity): định lượng mức độ độc hại bằng bộ phân loại hate speech đã huấn luyện sẵn
  • aggregation="maximum" trả về điểm độc hại tối đa trên các đầu vào
  • aggregation="ratio" trả về % dự đoán có độc hại > 0.5
toxicity_metric = load("toxicity")
texts_1 = ["Everyone in the team adores him", "He is a true genius, pure talent"]
texts_2 = ["Nobody in the team likes him", "He is a useless 'good-for-nothing'"]

toxicity_results_1 = toxicity_metric.compute(predictions=texts_1,aggregation="maximum") toxicity_results_2 = toxicity_metric.compute(predictions=texts_2, aggregation="maximum") print("Toxicity Sentences 1:", toxicity_results_1) print("Toxicity Sentences 2:", toxicity_results_2)
Toxicity Sentences 1: {'max_toxicity': 0.0001386617950629443}
Toxicity Sentences 2: {'max_toxicity': 0.012273495085537434}
Nhập môn LLMs trong Python

Chỉ số phân tích thiên lệch LLM: regard

  • Regard: mức độ thiện cảm/định kiến ngôn ngữ đối với một số nhóm nhân khẩu

  • Đánh giá điểm regard trên đầu ra LLM gắn với hai nhóm riêng biệt

regard = load("regard")

group1 = ['abc are described as loyal employees',
          'abc are ambitious in their career expectations']
group2 = ['abc are known for causing lots of team conflicts',
          'abc are verbally violent']


polarity_results_1 = regard.compute(data=group1) polarity_results_2 = regard.compute(data=group2)
Nhập môn LLMs trong Python

Chỉ số phân tích thiên lệch LLM: regard

for result in polarity_results_1['regard']:
  print(result)
[{'label': 'positive', 'score': 0.9098386764526367}, 
 {'label': 'neutral', 'score': 0.059396952390670776}, 
 {'label': 'other', 'score': 0.026468101888895035}, 
 {'label': 'negative', 'score': 0.004296252969652414}]
[{'label': 'positive', 'score': 0.7809812426567078}, 
 {'label': 'neutral', 'score': 0.18085983395576477}, 
 {'label': 'other', 'score': 0.030492952093482018}, 
 {'label': 'negative', 'score': 0.007666013203561306}]
for result in polarity_results_2['regard']:
  print(result)
[{'label': 'negative', 'score': 0.9658734202384949}, 
 {'label': 'other', 'score': 0.021555885672569275}, 
 {'label': 'neutral', 'score': 0.012026479467749596},
 {'label': 'positive', 'score': 0.0005441228277049959}]
[{'label': 'negative', 'score': 0.9774736166000366}, 
 {'label': 'other', 'score': 0.012994581833481789},  
 {'label': 'neutral', 'score': 0.008945506066083908}, 
 {'label': 'positive', 'score': 0.0005862844991497695}]
Nhập môn LLMs trong Python

Ayo berlatih!

Nhập môn LLMs trong Python

Preparing Video For Download...