R로 하는 가설 검정
Richie Cotton
Data Evangelist at DataCamp
$H_{0}$: 30세 미만과 30세 이상의 SO 사용자 중 취미 사용자 비율은 동일합니다.
$H_{0}$: $p_{\geq30} - p_{<30} = 0$
$H_{A}$: 30세 미만과 30세 이상의 SO 사용자 중 취미 사용자 비율은 다릅니다.
$H_{A}$: $p_{\geq30} - p_{<30} \neq 0$
alpha <- 0.05
$$ z = \frac{(\hat{p}_{\geq30} - \hat{p}_{<30}) - 0}{\text{SE}(\hat{p}_{\geq30} - \hat{p}_{<30})} $$
$$ \text{SE}(\hat{p}_{\geq30} - \hat{p}_{<30}) = \sqrt{\dfrac{\hat{p} \times (1 - \hat{p})}{n_{\geq30}} + \dfrac{\hat{p} \times (1 - \hat{p})}{n_{<30}}} $$
$\hat{p}$는 $p$(성공의 공통 미지 비율)에 대한 합동 추정치입니다.
$$ \hat{p} = \frac{n_{\geq30} \times \hat{p}_{\geq30} + n_{<30} \times \hat{p}_{<30}}{n_{\geq30} + n_{<30} } $$
필요한 수치는 4개입니다: $\hat{p}_{\geq30}$, $\hat{p}_{<30}$, $n_{\geq30}$, $n_{<30}$.
stack_overflow %>%
group_by(age_cat) %>%
summarize(
p_hat = mean(hobbyist == "Yes"),
n = n()
)
# A tibble: 2 x 3
age_cat p_hat n
<chr> <dbl> <int>
1 At least 30 0.773 1050
2 Under 30 0.843 1216
z_score
-4.217
library(infer) stack_overflow %>% prop_test(hobbyist ~ age_cat, # proportions ~ categoriesorder = c("At least 30", "Under 30"), # which p-hat to subtractsuccess = "Yes", # which response value to count proportions ofalternative = "two-sided", # type of alternative hypothesiscorrect = FALSE # should Yates' continuity correction be applied?)
# A tibble: 1 x 6
statistic chisq_df p_value alternative lower_ci upper_ci
<dbl> <dbl> <dbl> <chr> <dbl> <dbl>
1 17.8 1 0.0000248 two.sided 0.0605 0.165
R로 하는 가설 검정