ベイズ線形回帰

rstanarm で学ぶベイズ回帰モデリング

Jake Thompson

Psychometrician, ATLAS, University of Kansas

ベイズ法を使う理由

  • p値はデータの確率を推測するものであり、パラメータ値ではない
  • 事後分布: 尤度と事前分布の組み合わせ
    • 事後分布からサンプリング
    • サンプルの要約
    • 要約を用いてパラメータ値を推測
rstanarm で学ぶベイズ回帰モデリング

rstanarmパッケージ

  • Stan確率的プログラミング言語へのインターフェース
  • rstanarmStanへの高レベルアクセスを提供
  • カスタムモデル定義が可能
rstanarm で学ぶベイズ回帰モデリング
library(rstanarm)

stan_model <- stan_glm(kid_score ~ mom_iq, data = kidiq)
SAMPLING FOR MODEL 'continuous' NOW (CHAIN 1).
 Gradient evaluation took 0.000408 seconds
 1000 transitions using 10 leapfrog steps per transition would take
 4.08 seconds.
 Adjust your expectations accordingly!

 Iteration:    1 / 2000 [  0%]  (Warmup)
 Iteration:  200 / 2000 [ 10%]  (Warmup)
 Iteration:  400 / 2000 [ 20%]  (Warmup)
 Iteration:  600 / 2000 [ 30%]  (Warmup)
 Iteration:  800 / 2000 [ 40%]  (Warmup)
 Iteration: 1000 / 2000 [ 50%]  (Warmup)
 Iteration: 1001 / 2000 [ 50%]  (Sampling)
 Iteration: 1200 / 2000 [ 60%]  (Sampling)
 Iteration: 1400 / 2000 [ 70%]  (Sampling)
 Iteration: 1600 / 2000 [ 80%]  (Sampling)
 Iteration: 1800 / 2000 [ 90%]  (Sampling)
 Iteration: 2000 / 2000 [100%]  (Sampling)

  Elapsed Time: 0.37735 seconds (Warm-up)
                0.252244 seconds (Sampling)
                0.629594 seconds (Total)
rstanarm で学ぶベイズ回帰モデリング
summary(stan_model)
 Model Info:
  function:     stan_glm
  family:       gaussian [identity]
  formula:      kid_score ~ mom_iq
  algorithm:    sampling
  priors:       see help('prior_summary')
  sample:       4000 (posterior sample size)
  observations: 434
  predictors:   2

 Estimates:
                 mean    sd      2.5%    25%     50%     75%     97.5%
 (Intercept)      25.7     6.0    13.8    21.6    25.7    30.0    37.0
 mom_iq            0.6     0.1     0.5     0.6     0.6     0.7     0.7
 sigma            18.3     0.6    17.1    17.9    18.3    18.7    19.5
 mean_PPD         86.8     1.2    84.3    85.9    86.8    87.6    89.2
 log-posterior -1885.4     1.2 -1888.5 -1886.0 -1885.1 -1884.5 -1884.0

 Diagnostics:
               mcse Rhat n_eff
 (Intercept)   0.1  1.0  4000 
 mom_iq        0.0  1.0  4000 
 sigma         0.0  1.0  3827 
 mean_PPD      0.0  1.0  4000 
 log-posterior 0.0  1.0  1896 

 For each parameter, mcse is Monte Carlo standard error, n_eff is a crude measure of effective sample size, and Rhat is the potential scale reduction factor
 on split chains (at convergence Rhat=1).
rstanarm で学ぶベイズ回帰モデリング

rstanarmの要約: 推定値

 Estimates:
                 mean    sd      2.5%    25%     50%     75%     97.5%
 (Intercept)      25.7     6.0    13.8    21.6    25.7    30.0    37.0
 mom_iq            0.6     0.1     0.5     0.6     0.6     0.7     0.7
 sigma            18.3     0.6    17.1    17.9    18.3    18.7    19.5
 mean_PPD         86.8     1.2    84.3    85.9    86.8    87.6    89.2
 log-posterior -1885.4     1.2 -1888.5 -1886.0 -1885.1 -1884.5 -1884.0
  • sigma: 誤差の標準偏差
  • mean_PPD: 事後予測サンプルの平均
  • log-posterior: 尤度に類似した指標
rstanarm で学ぶベイズ回帰モデリング

rstanarmの要約: 診断

Diagnostics:
             mcse Rhat n_eff
(Intercept)   0.1  1.0  4000 
mom_iq        0.0  1.0  4000 
sigma         0.0  1.0  3827 
mean_PPD      0.0  1.0  4000 
log-posterior 0.0  1.0  1896 

For each parameter, mcse is Monte Carlo standard error,
n_eff is a crude measure of effective sample size, and
Rhat is the potential scale reduction factor on split chains
 (at convergence Rhat=1).
  • Rhat: チェーン内分散とチェーン間分散の比
  • 1.1未満の値は収束を示す
rstanarm で学ぶベイズ回帰モデリング

練習しましょう!

rstanarm で学ぶベイズ回帰モデリング

Preparing Video For Download...