การประเมินโมเดล RLHF

Reinforcement Learning from Human Feedback (RLHF)

Mina Parham

AI Engineer

เมตริกอัตโนมัติ

 

  • งาน Classification: Accuracy, F1 score
classification_results.head(3)
| ID | Feedback_Text                         | True_Category | Predicted_Category |
|----|---------------------------------------|---------------|--------------------|
| 1  | "Arrived on time and works great."    | Positive      | Positive           |
| 2  | "I had issues with customer service." | Negative      | Neutral            |
| 3  | "The website is easy to navigate."    | Positive      | Positive           |
Reinforcement Learning from Human Feedback (RLHF)

เมตริกอัตโนมัติ

 

  • การสร้างข้อความ, การสรุปความ: ROUGE, BLEU
text_generation.head(3)
| ID | Prompt               | True_Completion  | Pred_Completion   |
|----|----------------------|------------------|-------------------|
| 1  | "Customer service"   | "can help you."  | "will assist."    |
| 2  | "To get a refund,"   | "contact us."    | "reach out."      |
| 3  | "Support team is"    | "here 24/7."     | "available 24/7." |
Reinforcement Learning from Human Feedback (RLHF)

เมตริกอัตโนมัติ

 

 

ประโยคอ้างอิง:

  • RLHF ปรับปรุงการปรับแนว (alignment) ของโมเดล ให้สอดคล้องกับค่านิยมของมนุษย์

 

 

คะแนน ROUGE: 0.83

 

 

ประโยคที่นำมาเปรียบเทียบ:

  • RLHF aligns models with human values.
Reinforcement Learning from Human Feedback (RLHF)

กราฟ Artifact

config = PPOConfig(
    model_name="lvwerra/gpt2-imdb",learning_rate=1.41e-5, log_with="wandb")
import wandb
wandb.init()

ภาพหน้าจอผลลัพธ์ใน Weights and Biases

Reinforcement Learning from Human Feedback (RLHF)

กราฟ Artifact

  • รางวัลเพิ่มขึ้นเมื่อโมเดลเรียนรู้

กราฟแสดงแนวโน้มรางวัลที่เพิ่มขึ้น ซึ่งหมายความว่าโมเดลกำลังพัฒนาขึ้น

  • กราฟ KL ควรเพิ่มขึ้นอย่างค่อยเป็นค่อยไป

กราฟแสดงแนวโน้มที่เพิ่มขึ้นอย่างค่อยเป็นค่อยไปของ KL loss

Reinforcement Learning from Human Feedback (RLHF)

การประเมินที่เน้นมนุษย์เป็นศูนย์กลาง

  • การประเมินโดยมนุษย์: ใช้การตัดสินเชิงอัตนัยและความเข้าใจบริบทเชิงลึก

ผู้ประเมินมนุษย์กำลังทำงานหน้าแล็ปท็อป

  • การประเมินโดยโมเดล: รองรับการขยายขนาดและให้ผลสอดคล้องกัน

หุ่นยนต์กับฟองข้อความแทนการประเมินโดยโมเดล

Reinforcement Learning from Human Feedback (RLHF)

มาฝึกกันเถอะ!

Reinforcement Learning from Human Feedback (RLHF)

Preparing Video For Download...