R에서 대치(Imputation)로 결측치 다루기
Michal Oleszak
Machine Learning Engineer
이 과정을 마치면 다음을 수행할 수 있습니다:
이 과정은 다음 주제에 익숙하다고 가정합니다:
dplyr과 파이프 연산자(%>%)로 기본 데이터 조작lm(), glm())결측치를 다루는 최선의 방법은 처음부터 없게 하는 것입니다.
안타깝게도 결측치는 어디에나 있습니다:
결측치에 항상 유의해야 합니다.
head(nhanes, 3)
Age Gender Weight Height Diabetes TotChol Pulse PhysActive
1 16 male 73.2 172.0 FALSE 3.00 76 TRUE
2 17 male 72.3 176.0 FALSE 2.61 74 TRUE
3 12 male 57.7 158.9 FALSE 4.27 80 TRUE
nhanes %>% is.na() %>% colSums()
Age Gender Weight Height Diabetes TotChol Pulse PhysActive
0 0 9 8 1 85 32 26
model_1 <- lm(Diabetes ~ Age + Weight,
data = nhanes)
summary(model_1)의 일부:
Residual standard error: 0.08571 on 804
degrees of freedom (10 observations
deleted due to missingness)
Adjusted R-squared: 0.005706
F-statistic: 3.313 on 2 and 804 DF,
p-value: 0.03691
model_2 <- lm(Diabetes ~ Age + Weight +
TotChol, data = nhanes)
summary(model_2)의 일부:
Residual standard error: 0.08264 on 718
degrees of freedom (95 observations
deleted due to missingness)
Adjusted R-squared: 0.008422
F-statistic: 3.041 on 3 and 718 DF,
p-value: 0.02834
R에서 대치(Imputation)로 결측치 다루기