在建模前先轉換輸入

R 中的監督式學習:回歸

Nina Zumel and John Mount

Win-Vector LLC

為何要轉換輸入變數

  • 領域知識/合成變數
    • $Intelligence \sim \frac{mass.brain}{mass.body^{2/3}}$
R 中的監督式學習:回歸

為何要轉換輸入變數

  • 領域知識/合成變數
    • $Intelligence \sim \frac{mass.brain}{mass.body^{2/3}}$
  • 實務考量
    • 對數轉換可縮小動態範圍
    • 當變數的有意義變化是乘法性時用對數轉換
R 中的監督式學習:回歸

為何要轉換輸入變數

  • 領域知識/合成變數
    • $Intelligence \sim \frac{mass.brain}{mass.body^{2/3}}$
  • 實務考量
    • 對數轉換可縮小動態範圍
    • 當變數的有意義變化是乘法性時用對數轉換
    • $y$ 與 $f(x)$ 近似線性,而非與 $x$ 線性
R 中的監督式學習:回歸

範例:預測焦慮

R 中的監督式學習:回歸

轉換 hassles 變數

R 中的監督式學習:回歸

不同的擬合方式

哪一個最好?

  • anx ~ I(hassles^2)
  • anx ~ I(hassles^3)
  • anx ~ I(hassles^2) + I(hassles^3)
  • anx ~ exp(hassles)
  • ...

I():將運算式視為字面量(不是交互作用)

R 中的監督式學習:回歸

比較不同模型

線性、二次與三次模型

mod_lin <- lm(anx ~ hassles, hassleframe)
summary(mod_lin)$r.squared
0.5334847
mod_quad <- lm(anx ~ I(hassles^2), hassleframe)
summary(mod_quad)$r.squared
0.6241029
mod_tritic <- lm(anx ~ I(hassles^3), hassleframe)
summary(mod_tritic)$r.squared
0.6474421
R 中的監督式學習:回歸

比較不同模型

用交叉驗證評估模型

Model RMSE
Linear ($hassles$) 7.69
Quadratic ($hassles^2$) 6.89
Cubic ($hassles^3$) 6.70
R 中的監督式學習:回歸

一起來練習吧!

R 中的監督式學習:回歸

Preparing Video For Download...