Faydaları ölçmek

R'da Paralel Programlama

Nabeel Imam

Data Scientist

Basit örnek

numbers <- 1:1000000


# Sequential sqroots <- lapply(numbers, sqrt)
# Parallel cl <- makeCluster(4) sqroots <- parLapply(cl, numbers, sqrt) stopCluster(my_cluster)

Hangisi daha iyi çalışır?

R'da Paralel Programlama

Performansı kıyaslama

Ortalama çalışma süresini görmek için kodu birkaç kez çalıştır

library(microbenchmark)


microbenchmark( "Sequential" = lapply(numbers, sqrt),
"Parallel" = { cl <- makeCluster(4) parLapply(cl, numbers, sqrt) stopCluster(my_cluster) },
times = 10 )

 

 

Unit: milliseconds
      expr     min    mean     max neval
Sequential  633.96  838.09  993.59    10
  Parallel 1136.95 1247.29 1557.58    10
  • Basit sayısal işlemler paralelleştirmeden nadiren fayda görür
  • Profiling satır satır rapor verir, benchmarking toplam çalışma süresini verir
R'da Paralel Programlama

Ortadaki fil

sqroots <- sqrt(numbers)

Bir oturma odasında kanepede oturan bir fil ve insanlar onun varlığını kabulleniyor.

R'da Paralel Programlama

Vektörizasyon

sqroots <- sqrt(numbers)
  • sqrt() gibi Base R fonksiyonları vektörizedir.
  • Tek bir fonksiyonu çok sayıda girdiye uygular
  • Çok hızlıdır ama sadece basit işlemler için geçerlidir
microbenchmark(
  "Vectorized" = sqrt(numbers),
  "Sequential" = lapply(numbers, sqrt),
  "Parallel" = {
    cl <- makeCluster(4)
    parLapply(cl, numbers, sqrt)
    stopCluster(my_cluster)
  },
  times = 10)
Unit: milliseconds
      expr       min      mean      max neval
Vectorized    2.3904    9.2071   66.303    10
Sequential  352.1166  771.7491 1004.753    10
  Parallel 1191.3176 1377.6926 1700.316    10
R'da Paralel Programlama

Bootstrap

Mevcut veriden geri koymalı örnekleme

print(ls_df)
$`2001`
   Country             Life_expectancy  Year
 1 Afghanistan                    56.3  2001
 2 Albania                        74.3  2001
 3 Algeria                        71.1  2001
...
$`2002`
   Country             Life_expectancy  Year
 1 Afghanistan                    56.8  2002
 2 Albania                        74.6  2002
 3 Algeria                        71.6  2002
...
R'da Paralel Programlama

Klasik sürüm

df <- ls_df$`2001`


estimates <- rep(0, 10000)
for (i in 1:10000) { b <- sample(df$Life_expectancy, replace = T)
estimates[i] <- mean(b) }

2001'de küresel ortalama yaşam beklentisinin bootstrap ile elde edilmiş tahminlerinin klasik çan eğrisi şeklindeki histogramı.

  • Kuantillerle güven aralığı: quantile(estimates, c(0.025, 0.975))
R'da Paralel Programlama

İyi haber

Bootstrap işlemleri paralelleştirilebilir

estimates <- rep(0, 10000)

for (i in 1:10000) {

  b <- sample(df$Life_expectancy,
              replace = T)

  estimates[i] <- mean(b)
  }
boot_dist <- function (df) {

  estimates <- rep(0, 10000)

  for (i in 1:10000) {
    b <- sample(df$Life_expectancy, replace = T)
    estimates[i] <- mean(b)
  }

  return(estimates)
}


cl <- makeCluster(4) ls_dists <- parLapply(cl, ls_df, boot_dist) stopCluster(cl)
R'da Paralel Programlama

Kazançlar

microbenchmark(
  "lapply" = lapply(ls_df, boot_dist),
  "parLapply" = {
    cl <- makeCluster(4)
    parLapply(cl, ls_df, boot_dist)
    stopCluster(cl)
  },
  times = 10
)
Unit: seconds
     expr    min   mean    max neval
   lapply 3.6938 4.2184 4.5267    10
parLapply 1.9006 2.5166 2.7292    10

Buna nasıl ulaşılır:

  • Mevcut kodu profile et, en yavaş kısmı bul
  • Bu adımı paralelleştir/iyileştir
  • Kıyasla ve karşılaştır
R'da Paralel Programlama

Hadi pratik yapalım!

R'da Paralel Programlama

Preparing Video For Download...