Pengantar Data Engineering
Vincent Vankrunkelsven
Data Engineer @ DataCamp
Analitik

Aplikasi

Analitik

Aplikasi

Database Pemrosesan Paralel Masif

Muat dari file ke format penyimpanan kolumnar
# Metode Pandas .to_parquet()
df.to_parquet("./s3://path/to/bucket/customer.parquet")
# Metode PySpark .write.parquet()
df.write.parquet("./s3://path/to/bucket/customer.parquet")
COPY customer
FROM 's3://path/to/bucket/customer.parquet'
FORMAT as parquet
...
pandas.to_sql()
# Transformasi data
recommendations = transform_find_recommendatins(ratings_df)
# Muat ke database PostgreSQL
recommendations.to_sql("recommendations",
db_engine,
schema="store",
if_exists="replace")
Pengantar Data Engineering