Pipeline dữ liệu trên Kubernetes

Nhập môn Kubernetes

Frank Heilmann

Platform Architect and Freelance Instructor

Pipeline dữ liệu là gì?

  • Tập hợp quy trình để di chuyển, chuyển đổi hoặc phân tích dữ liệu
  • Các bước điển hình:
    • ETL: Extract dữ liệu từ nhiều nguồn, sau đó Transform thành schema có ý nghĩa, cuối cùng Load vào đích (ví dụ: data warehouse)
    • ELT: Extract dữ liệu từ nhiều nguồn, sau đó Load vào đích (ví dụ: data lake), cuối cùng Transform khi cần để thành schema có ý nghĩa

Pipeline dữ liệu

Nhập môn Kubernetes

Pipeline dữ liệu trên Kubernetes

  • Các bước của một pipeline dữ liệu ánh xạ tốt sang đối tượng Kubernetes:
    • Bước Extract, Transform, Load: Pod (Deployment hoặc StatefulSet)
    • Dữ liệu đã Extract và Transform: Persistent Volume
  • Kubernetes có thể scale Deployment và Storage theo nhu cầu, nhờ đó tăng thông lượng

Pipeline dữ liệu trên Kubernetes

Nhập môn Kubernetes

Công cụ mã nguồn mở cho pipeline dữ liệu

  • Nhiều phần mềm mã nguồn mở có thể triển khai ngay trên Kubernetes
  • Ví dụ:
    • Extract: Apache NiFi, Apache Kafka với Kafka Connect
    • Transform: Apache Spark, Apache Kafka, PostgreSQL
    • Load: Apache Spark, Apache Kafka với KSQL, PostgreSQL
    • Lưu trữ trên PV: Minio, Ceph
  • Danh sách này không đầy đủ
Nhập môn Kubernetes

Ayo berlatih!

Nhập môn Kubernetes

Preparing Video For Download...