Bootstrap-betrouwbaarheidsintervallen

Statistical Thinking in Python (deel 2)

Justin Bois

Lecturer at the California Institute of Technology

Bootstrap-replicatiefunctie

def bootstrap_replicate_1d(data, func):
    """Generate bootstrap replicate of 1D data."""
    bs_sample = np.random.choice(data, len(data))
    return func(bs_sample)

bootstrap_replicate_1d(michelson_speed_of_light, np.mean)

299859.20000000001

bootstrap_replicate_1d(michelson_speed_of_light, np.mean)

299855.70000000001

bootstrap_replicate_1d(michelson_speed_of_light, np.mean)

299850.29999999999

Veel bootstrap-replicaties

bs_replicates = np.empty(10000)

for i in range(10000):
    bs_replicates[i] = bootstrap_replicate_1d(
                  michelson_speed_of_light, np.mean)

Histogram van bootstrap-replicaties plotten

_ = plt.hist(bs_replicates, bins=30, normed=True)
_ = plt.xlabel('mean speed of light (km/s)')
_ = plt.ylabel('PDF')
plt.show()

Bootstrap-schatting van het gemiddelde

ch2-2.011.png

Betrouwbaarheidsinterval van een statistiek

Als we metingen keer op keer herhalen, valt p% van de waarnemingen binnen het p%-betrouwbaarheidsinterval.

Bootstrap-betrouwbaarheidsinterval

conf_int = np.percentile(bs_replicates, [2.5, 97.5])

array([ 299837.,  299868.])

ch2-2.016.png

Laten we oefenen!

Statistical Thinking in Python (deel 2)