遺漏值:可能出什麼問題

在 R 中以插補處理遺漏值

Michal Oleszak

Machine Learning Engineer

你將學到什麼

完成本課後,你將能:

  • 了解為何遺漏值需要特別處理。
  • 用統計檢定與視覺化工具偵測遺漏值的模式。
  • 以多種統計與機器學習模型進行插補。
  • 把插補的不確定性納入分析與預測,讓結果更穩健。
在 R 中以插補處理遺漏值

先備知識

本課假設你熟悉以下主題:

  • 使用 dplyr 與管線運算子(%>%)進行基本資料操作。
  • 直線與羅吉斯迴歸模型(lm(), glm())。
  • 基本機率觀念:隨機變數、機率分配。
在 R 中以插補處理遺漏值

遺漏值入門

處理遺漏值的最佳方式,當然是不要有遺漏。

但遺漏值無所不在:

  • 問卷未作答。
  • 蒐集設備的技術問題。
  • 串接異質資料來源。
  • ……

你必須時時「留意遺漏值」。

1 Orchard, T., 和 M. A. Woodbury。1972 年。〈A Missing Information Principle: Theory and Applications〉。收錄於 Proceedings of the Sixth Berkeley Symposium on Mathematical Statistics and Probability,1:697–715。
在 R 中以插補處理遺漏值

NHANES 資料

head(nhanes, 3)
  Age Gender Weight Height Diabetes TotChol Pulse PhysActive
1  16   male   73.2  172.0    FALSE    3.00    76       TRUE
2  17   male   72.3  176.0    FALSE    2.61    74       TRUE
3  12   male   57.7  158.9    FALSE    4.27    80       TRUE
nhanes %>% is.na() %>% colSums()
Age     Gender     Weight     Height   Diabetes    TotChol    Pulse   PhysActive 
0       0          9          8        1            85        32      26
在 R 中以插補處理遺漏值

含遺漏值的線性迴歸

model_1 <- lm(Diabetes ~ Age + Weight, 
              data = nhanes)

summary(model_1) 的部分輸出:

Residual standard error: 0.08571 on 804 
degrees of freedom (10 observations 
deleted due to missingness)

Adjusted R-squared:  0.005706 
F-statistic: 3.313 on 2 and 804 DF,  
p-value: 0.03691
model_2 <- lm(Diabetes ~ Age + Weight +
              TotChol, data = nhanes)

summary(model_2) 的部分輸出:

Residual standard error: 0.08264 on 718 
degrees of freedom (95 observations
deleted due to missingness)

Adjusted R-squared:  0.008422 
F-statistic: 3.041 on 3 and 718 DF,
p-value: 0.02834
在 R 中以插補處理遺漏值

重點整理

  • 統計軟體有時會悄悄忽略遺漏值。
  • 因此,不同模型可能無法相互比較。
  • 直接刪除不完整觀測,可能導致偏誤。
  • 有遺漏值時,務必妥善處理
在 R 中以插補處理遺漏值

一起來練習吧!

在 R 中以插補處理遺漏值

Preparing Video For Download...