Nền tảng Suy luận trong R
Jo Hardin
Instructor






Tạo phân phối của thống kê từ quần thể theo giả thuyết không cho biết liệu dữ liệu quan sát có không phù hợp với giả thuyết không hay không
Dữ liệu gốc
| Vị trí | Cola | Cam |
|---|---|---|
| Đông | 28 | 6 |
| Tây | 19 | 7 |
$\hat{p}_\text{east} = 28/(28 + 6) = 0.82$
$\hat{p}_\text{west} = 19/(19 + 7) = 0.73$
Lần tráo thứ nhất, giống dữ liệu gốc
| Vị trí | Cola | Cam |
|---|---|---|
| Đông | 28 | 6 |
| Tây | 19 | 7 |

Lần tráo thứ hai
| Vị trí | Cola | Cam |
|---|---|---|
| Đông | 27 | 7 |
| Tây | 20 | 6 |

Lần tráo thứ ba
| Vị trí | Cola | Cam |
|---|---|---|
| Đông | 28 | 8 |
| Tây | 21 | 5 |

Lần tráo thứ tư
| Vị trí | Cola | Cam |
|---|---|---|
| Đông | 25 | 9 |
| Tây | 22 | 4 |

Lần tráo thứ năm
| Vị trí | Cola | Cam |
|---|---|---|
| Đông | 29 | 5 |
| Tây | 18 | 8 |

Lần tráo thứ năm
| Vị trí | Cola | Cam |
|---|---|---|
| Đông | 29 | 5 |
| Tây | 18 | 8 |







soda %>%
group_by(location) %>%
summarize(prop_cola =
mean(drink == "cola")) %>%
summarize(diff(prop_cola))
# A tibble: 1 x 1
`diff(prop_cola)`
<dbl>
1 -0.09276018
library(infer)
soda %>% specify(drink ~ location,
success = "cola") %>%
hypothesize(null = "independence") %>%
generate(reps = 1, type = "permute") %>%
calculate(stat = "diff in props",
order = c("west","east"))
# A tibble: 1 x 2
replicate stat
<int> <dbl>
1 1 -0.02488688
soda %>%
specify(drink ~ location, success = "cola") %>%
hypothesize(null = "independence") %>%
generate(reps = 5, type = "permute") %>%
calculate(stat = "diff in props", order = c("west", "east"))
# A tibble: 5 x 2
replicate stat
<int> <dbl>
1 1 0.04298643
2 2 -0.09276018
3 3 0.11085973
4 4 0.17873303
5 5 -0.16063348

Nền tảng Suy luận trong R