貝葉斯線性迴歸

使用 rstanarm 的貝葉斯迴歸建模

Jake Thompson

Psychometrician, ATLAS, University of Kansas

為何使用貝葉斯方法?

  • P 值推論的是資料的機率,而非參數值
  • 後驗分佈:由概似與先驗組合而成
    • 對後驗分佈取樣
    • 彙總取樣結果
    • 用摘要對參數值進行推論
使用 rstanarm 的貝葉斯迴歸建模

rstanarm 套件

  • Stan 機率程式語言的介面
  • rstanarm 提供 Stan 的高階存取
  • 可自訂模型定義
使用 rstanarm 的貝葉斯迴歸建模
library(rstanarm)

stan_model <- stan_glm(kid_score ~ mom_iq, data = kidiq)
SAMPLING FOR MODEL 'continuous' NOW (CHAIN 1).
 Gradient evaluation took 0.000408 seconds
 1000 transitions using 10 leapfrog steps per transition would take
 4.08 seconds.
 Adjust your expectations accordingly!

 Iteration:    1 / 2000 [  0%]  (Warmup)
 Iteration:  200 / 2000 [ 10%]  (Warmup)
 Iteration:  400 / 2000 [ 20%]  (Warmup)
 Iteration:  600 / 2000 [ 30%]  (Warmup)
 Iteration:  800 / 2000 [ 40%]  (Warmup)
 Iteration: 1000 / 2000 [ 50%]  (Warmup)
 Iteration: 1001 / 2000 [ 50%]  (Sampling)
 Iteration: 1200 / 2000 [ 60%]  (Sampling)
 Iteration: 1400 / 2000 [ 70%]  (Sampling)
 Iteration: 1600 / 2000 [ 80%]  (Sampling)
 Iteration: 1800 / 2000 [ 90%]  (Sampling)
 Iteration: 2000 / 2000 [100%]  (Sampling)

  Elapsed Time: 0.37735 seconds (Warm-up)
                0.252244 seconds (Sampling)
                0.629594 seconds (Total)
使用 rstanarm 的貝葉斯迴歸建模
summary(stan_model)
 Model Info:
  function:     stan_glm
  family:       gaussian [identity]
  formula:      kid_score ~ mom_iq
  algorithm:    sampling
  priors:       see help('prior_summary')
  sample:       4000 (posterior sample size)
  observations: 434
  predictors:   2

 Estimates:
                 mean    sd      2.5%    25%     50%     75%     97.5%
 (Intercept)      25.7     6.0    13.8    21.6    25.7    30.0    37.0
 mom_iq            0.6     0.1     0.5     0.6     0.6     0.7     0.7
 sigma            18.3     0.6    17.1    17.9    18.3    18.7    19.5
 mean_PPD         86.8     1.2    84.3    85.9    86.8    87.6    89.2
 log-posterior -1885.4     1.2 -1888.5 -1886.0 -1885.1 -1884.5 -1884.0

 Diagnostics:
               mcse Rhat n_eff
 (Intercept)   0.1  1.0  4000 
 mom_iq        0.0  1.0  4000 
 sigma         0.0  1.0  3827 
 mean_PPD      0.0  1.0  4000 
 log-posterior 0.0  1.0  1896 

 For each parameter, mcse is Monte Carlo standard error, n_eff is a crude measure of effective sample size, and Rhat is the potential scale reduction factor
 on split chains (at convergence Rhat=1).
使用 rstanarm 的貝葉斯迴歸建模

rstanarm 摘要:估計值

 Estimates:
                 mean    sd      2.5%    25%     50%     75%     97.5%
 (Intercept)      25.7     6.0    13.8    21.6    25.7    30.0    37.0
 mom_iq            0.6     0.1     0.5     0.6     0.6     0.7     0.7
 sigma            18.3     0.6    17.1    17.9    18.3    18.7    19.5
 mean_PPD         86.8     1.2    84.3    85.9    86.8    87.6    89.2
 log-posterior -1885.4     1.2 -1888.5 -1886.0 -1885.1 -1884.5 -1884.0
  • sigma:殘差的標準差
  • mean_PPD:後驗預測樣本的平均
  • log-posterior:類似於概似
使用 rstanarm 的貝葉斯迴歸建模

rstanarm 摘要:診斷

Diagnostics:
             mcse Rhat n_eff
(Intercept)   0.1  1.0  4000 
mom_iq        0.0  1.0  4000 
sigma         0.0  1.0  3827 
mean_PPD      0.0  1.0  4000 
log-posterior 0.0  1.0  1896 

For each parameter, mcse is Monte Carlo standard error,
n_eff is a crude measure of effective sample size, and
Rhat is the potential scale reduction factor on split chains
 (at convergence Rhat=1).
  • Rhat:比較鏈內與鏈間變異的指標
  • 小於 1.1 表示已收斂
使用 rstanarm 的貝葉斯迴歸建模

一起來練習吧!

使用 rstanarm 的貝葉斯迴歸建模

Preparing Video For Download...