Pipeline dữ liệu trên Kubernetes
Nhập môn Kubernetes
Frank Heilmann
Platform Architect and Freelance Instructor
Pipeline dữ liệu là gì?
Tập hợp quy trình để di chuyển, chuyển đổi hoặc phân tích dữ liệu
Các bước điển hình:
ETL
:
Extract
dữ liệu từ nhiều nguồn, sau đó
Transform
thành schema có ý nghĩa, cuối cùng
Load
vào đích (ví dụ: data warehouse)
ELT
:
Extract
dữ liệu từ nhiều nguồn, sau đó
Load
vào đích (ví dụ: data lake), cuối cùng
Transform
khi cần để thành schema có ý nghĩa
Pipeline dữ liệu trên Kubernetes
Các bước của một pipeline dữ liệu ánh xạ tốt sang đối tượng Kubernetes:
Bước Extract, Transform, Load: Pod (Deployment hoặc StatefulSet)
Dữ liệu đã Extract và Transform: Persistent Volume
Kubernetes có thể scale Deployment và Storage theo nhu cầu, nhờ đó tăng thông lượng
Công cụ mã nguồn mở cho pipeline dữ liệu
Nhiều phần mềm mã nguồn mở có thể triển khai ngay trên Kubernetes
Ví dụ:
Extract:
Apache NiFi
,
Apache Kafka
với
Kafka Connect
Transform:
Apache Spark
,
Apache Kafka
,
PostgreSQL
Load:
Apache Spark
,
Apache Kafka
với
KSQL
,
PostgreSQL
Lưu trữ trên PV:
Minio
,
Ceph
Danh sách này không đầy đủ
Ayo berlatih!
Nhập môn Kubernetes
Preparing Video For Download...