LLMの安全性の確保

Pythonで学ぶ LLM 入門

Jasmin Ludolf

Senior Data Science Content Developer, DataCamp

LLMの課題

多言語対応: 言語多様性、リソース可用性、適応性

多言語対応

オープン vs クローズド LLMのジレンマ: 共同利用 vs 責任ある利用

オープンとクローズドのLLM

モデルのスケーラビリティ: 表現力、計算負荷、学習要件

LLMのスケーラビリティ

バイアス: 偏った学習データ、不公正な理解と生成

LLMのバイアス

1 Icon made by Freepik (freepik.com)
Pythonで学ぶ LLM 入門

真実性と幻覚

  • 幻覚: 生成文が正確なように見せかけて虚偽・無意味な情報を含む

LLMの幻覚

幻覚を抑える戦略:

  1. 多様で代表性のある学習データへの曝露
  2. 出力のバイアス監査 + バイアス除去手法
  3. センシティブ用途向けにファインチューニング
  4. プロンプト設計: プロンプトの精緻化
Pythonで学ぶ LLM 入門

真実性と幻覚

  • 幻覚: 生成文が正確なように見せかけて虚偽・無意味な情報を含む

LLMの幻覚

Pythonで学ぶ LLM 入門

LLMバイアス分析の指標: 毒性

  • 毒性: 事前学習のヘイトスピーチ分類器でテキストの毒性を定量化
  • aggregation="maximum" は入力全体での最大毒性スコアを返す
  • aggregation="ratio" は毒性>0.5の予測割合(%)を返す
toxicity_metric = load("toxicity")
texts_1 = ["Everyone in the team adores him", "He is a true genius, pure talent"]
texts_2 = ["Nobody in the team likes him", "He is a useless 'good-for-nothing'"]

toxicity_results_1 = toxicity_metric.compute(predictions=texts_1,aggregation="maximum") toxicity_results_2 = toxicity_metric.compute(predictions=texts_2, aggregation="maximum") print("Toxicity Sentences 1:", toxicity_results_1) print("Toxicity Sentences 2:", toxicity_results_2)
Toxicity Sentences 1: {'max_toxicity': 0.0001386617950629443}
Toxicity Sentences 2: {'max_toxicity': 0.012273495085537434}
Pythonで学ぶ LLM 入門

LLMバイアス分析の指標: regard

  • Regard: 特定属性集団に対する言語の極性と偏見

  • 2つのグループに対応するLLM出力で別々にregardスコアを評価

regard = load("regard")

group1 = ['abc are described as loyal employees',
          'abc are ambitious in their career expectations']
group2 = ['abc are known for causing lots of team conflicts',
          'abc are verbally violent']


polarity_results_1 = regard.compute(data=group1) polarity_results_2 = regard.compute(data=group2)
Pythonで学ぶ LLM 入門

LLMバイアス分析の指標: regard

for result in polarity_results_1['regard']:
  print(result)
[{'label': 'positive', 'score': 0.9098386764526367}, 
 {'label': 'neutral', 'score': 0.059396952390670776}, 
 {'label': 'other', 'score': 0.026468101888895035}, 
 {'label': 'negative', 'score': 0.004296252969652414}]
[{'label': 'positive', 'score': 0.7809812426567078}, 
 {'label': 'neutral', 'score': 0.18085983395576477}, 
 {'label': 'other', 'score': 0.030492952093482018}, 
 {'label': 'negative', 'score': 0.007666013203561306}]
for result in polarity_results_2['regard']:
  print(result)
[{'label': 'negative', 'score': 0.9658734202384949}, 
 {'label': 'other', 'score': 0.021555885672569275}, 
 {'label': 'neutral', 'score': 0.012026479467749596},
 {'label': 'positive', 'score': 0.0005441228277049959}]
[{'label': 'negative', 'score': 0.9774736166000366}, 
 {'label': 'other', 'score': 0.012994581833481789},  
 {'label': 'neutral', 'score': 0.008945506066083908}, 
 {'label': 'positive', 'score': 0.0005862844991497695}]
Pythonで学ぶ LLM 入門

Passons à la pratique !

Pythonで学ぶ LLM 入門

Preparing Video For Download...