The Bigmemory 專案

R 的可擴展資料處理

Michael Kane

Assistant Professor, Yale University

bigmemory

用於儲存、操作與處理可能大於電腦 RAM 的大型矩陣。

R 的可擴展資料處理

big.matrix

  • 建立
  • 讀取
  • 取子集
  • 摘要
R 的可擴展資料處理

「out-of-core」是什麼意思?

  • R 物件保存在 RAM 中

  • 當 RAM 用盡時

    • 會移到磁碟
    • 程式會變慢或當機

最好只在需要處理時,才把資料載入 RAM。

R 的可擴展資料處理

何時使用 big.matrix?

  • 約為 RAM 的 20%
  • 稠密矩陣
R 的可擴展資料處理

bigmemory 概觀

  • bigmemory 實作 big.matrix 資料型別,用於在磁碟上建立、儲存、存取與操作矩陣

  • 資料留在磁碟,並在需要時隱式載入 RAM

R 的可擴展資料處理

bigmemory 概觀

一個 big.matrix 物件:

  • 只需匯入一次
  • 「backing」檔
  • 「descriptor」檔
R 的可擴展資料處理

bigmemory 範例

library(bigmemory)

# 建立新的 big.matrix 物件 x <- big.matrix(nrow = 1, ncol = 3, type = "double", init = 0, backingfile = "hello_big_matrix.bin", descriptorfile = "hello_big_matrix.desc")
R 的可擴展資料處理

backing 與 descriptor 檔

  • backing 檔:矩陣在磁碟上的二進位表示
  • descriptor 檔:包含中繼資料,如列數、欄數、名稱等
R 的可擴展資料處理

bigmemory 範例

# 查看內容
 x[,]
  0    0    0
x
An object of class "big.matrix"
Slot "address":
<pointer: 0x108e2a9a0>
R 的可擴展資料處理

與一般矩陣的相似處

# 修改第 1 列第 1 欄的值
x[1, 1] <- 3
# 驗證是否已更新
x[,]
   3    0    0
R 的可擴展資料處理

一起來練習吧!

R 的可擴展資料處理

Preparing Video For Download...