บทนำสู่การ Bootstrapping

การสุ่มตัวอย่างใน Python

James Chapman

Curriculum Manager, DataCamp

แบบคืนค่าและไม่คืนค่า

การสุ่มตัวอย่างแบบไม่คืนค่า:

ไพ่บนโต๊ะคาสิโน

การสุ่มตัวอย่างแบบคืนค่า ("Resampling"):

ลูกเต๋าสี่ลูกกำลังกลิ้ง

การสุ่มตัวอย่างใน Python

การสุ่มตัวอย่างแบบง่ายโดยไม่คืนค่า

ประชากร:

เมล็ดกาแฟเรียงเป็นแถวและคอลัมน์

ตัวอย่าง:

เมล็ดกาแฟเรียงเป็นแถวและคอลัมน์ โดยส่วนใหญ่ถูกทำให้เป็นสีเทา

การสุ่มตัวอย่างใน Python

การสุ่มตัวอย่างแบบง่ายโดยคืนค่า

ประชากร:

เมล็ดกาแฟเรียงเป็นแถวและคอลัมน์

Resample:

ตัวอย่างเมล็ดกาแฟแบบสุ่ม บางส่วนเป็นข้อมูลซ้ำ

การสุ่มตัวอย่างใน Python

เหตุใดจึงสุ่มแบบคืนค่า?

  • coffee_ratings: ตัวอย่างจากประชากรกาแฟทั้งหมดที่มีขนาดใหญ่กว่า
  • กาแฟแต่ละรายการในตัวอย่างแทนกาแฟในประชากรสมมติอีกหลายรายการ
  • การสุ่มแบบคืนค่าทำหน้าที่เป็น proxy
การสุ่มตัวอย่างใน Python

การเตรียมข้อมูลกาแฟ

coffee_focus = coffee_ratings[["variety", "country_of_origin", "flavor"]]
coffee_focus = coffee_focus.reset_index()
      index  variety country_of_origin  flavor
0         0     None          Ethiopia    8.83
1         1    Other          Ethiopia    8.67
2         2  Bourbon         Guatemala    8.50
3         3     None          Ethiopia    8.58
4         4    Other          Ethiopia    8.50
...     ...      ...               ...     ...
1333   1333     None           Ecuador    7.58
1334   1334     None           Ecuador    7.67
1335   1335     None     United States    7.33
1336   1336     None             India    6.83
1337   1337     None           Vietnam    6.67

[1338 rows x 4 columns]
การสุ่มตัวอย่างใน Python

การ Resampling ด้วย .sample()

coffee_resamp = coffee_focus.sample(frac=1, replace=True)
      index  variety country_of_origin  flavor
1140   1140  Bourbon         Guatemala    7.25
57       57  Bourbon         Guatemala    8.00
1152   1152  Bourbon            Mexico    7.08
621     621  Caturra          Thailand    7.50
44       44     SL28             Kenya    8.08
...     ...      ...               ...     ...
996     996   Typica            Mexico    7.33
1090   1090  Bourbon         Guatemala    7.33
918     918    Other         Guatemala    7.42
249     249  Caturra          Colombia    7.67
467     467  Caturra          Colombia    7.50

[1338 rows x 4 columns]
การสุ่มตัวอย่างใน Python

กาแฟที่ซ้ำกัน

coffee_resamp["index"].value_counts()
658     5
167     4
363     4
357     4
1047    4
       ..
771     1
770     1
766     1
764     1
0       1
Name: index, Length: 868, dtype: int64
การสุ่มตัวอย่างใน Python

กาแฟที่หายไป

num_unique_coffees = len(coffee_resamp.drop_duplicates(subset="index"))
868
len(coffee_ratings) - num_unique_coffees
470
การสุ่มตัวอย่างใน Python

Bootstrapping

ตรงข้ามกับการสุ่มตัวอย่างจากประชากร

Sampling: ดึงข้อมูลจากประชากรมาเป็นตัวอย่างขนาดเล็ก

Bootstrapping: สร้างประชากรเชิงทฤษฎีจากตัวอย่าง

กรณีการใช้งาน Bootstrapping:

  • ศึกษาความแปรปรวนของการสุ่มตัวอย่างโดยใช้ตัวอย่างเพียงชุดเดียว

บูทคาวบอย

การสุ่มตัวอย่างใน Python

กระบวนการ Bootstrapping

  1. สร้าง resample ที่มีขนาดเท่ากับตัวอย่างเดิม
  2. คำนวณสถิติที่สนใจสำหรับ bootstrap sample นี้
  3. ทำซ้ำขั้นตอนที่ 1 และ 2 หลาย ๆ ครั้ง

สถิติที่ได้คือ bootstrap statistics และประกอบกันเป็น bootstrap distribution

การสุ่มตัวอย่างใน Python

Bootstrapping ค่าเฉลี่ยรสชาติกาแฟ

import numpy as np

mean_flavors_1000 = []
for i in range(1000):
mean_flavors_1000.append(
np.mean(coffee_sample.sample(frac=1, replace=True)['flavor'])
)
การสุ่มตัวอย่างใน Python

ฮิสโตแกรมของ Bootstrap Distribution

import matplotlib.pyplot as plt
plt.hist(mean_flavors_1000)
plt.show()

Bootstrap distribution ของค่าเฉลี่ยรสชาติ

การสุ่มตัวอย่างใน Python

มาฝึกกันเถอะ!

การสุ่มตัวอย่างใน Python

Preparing Video For Download...