Kiểm thử pipeline Polars

Mở rộng và tối ưu hóa pipeline dữ liệu với Polars

Liam Brannigan

Data Scientist & Polars Contributor

Vì sao cần kiểm thử pipeline?

Sơ đồ hiển thị hai DataFrame khác nhau ở một ô.

Mở rộng và tối ưu hóa pipeline dữ liệu với Polars

Vì sao cần kiểm thử pipeline?

Sơ đồ cho thấy một phép biến đổi Polars được kiểm thử với một DataFrame mong đợi.

Mở rộng và tối ưu hóa pipeline dữ liệu với Polars

Yêu cầu đã hoàn tất theo phòng ban

def completed_by_department(requests: pl.LazyFrame) -> pl.LazyFrame:
    return






Mở rộng và tối ưu hóa pipeline dữ liệu với Polars

Yêu cầu đã hoàn tất theo phòng ban

def completed_by_department(requests: pl.LazyFrame) -> pl.LazyFrame:
    return (
        requests




    )
Mở rộng và tối ưu hóa pipeline dữ liệu với Polars

Yêu cầu đã hoàn tất theo phòng ban

def completed_by_department(requests: pl.LazyFrame) -> pl.LazyFrame:
    return (
        requests
        .group_by("DEPARTMENT")
        .len()


    )
Mở rộng và tối ưu hóa pipeline dữ liệu với Polars

Yêu cầu đã hoàn tất theo phòng ban

def completed_by_department(requests: pl.LazyFrame) -> pl.LazyFrame:
    return (
        requests
        .group_by("DEPARTMENT")
        .len()
        .with_columns(pl.col("len").cast(pl.Int32))
        .sort("DEPARTMENT")
    )
Mở rộng và tối ưu hóa pipeline dữ liệu với Polars

Dữ liệu kiểm thử đầu vào

sample = pl.DataFrame({
    "SR_NUMBER": ["SR1", "SR2", "SR3", "SR4"],
    "DEPARTMENT": ["Sanitation", "Water", "Sanitation", "Aviation"],
    "STATUS": ["Completed", "Open", "Completed", "Completed"],
})
shape: (4, 3)
| SR_NUMBER | DEPARTMENT | STATUS    |
| ---       | ---        | ---       |
| str       | str        | str       |
|-----------|------------|-----------|
| SR1       | Sanitation | Completed |
| SR2       | Water      | Open      |
| SR3       | Sanitation | Completed |
| SR4       | Aviation   | Completed |
Mở rộng và tối ưu hóa pipeline dữ liệu với Polars

Kết quả mong đợi

expected = pl.DataFrame(
    {
        "DEPARTMENT": ["Aviation", "Sanitation"],
        "len": [1, 2],
    }
)
shape: (2, 2)
| DEPARTMENT | len |
| ---        | --- |
| str        | i64 |
|------------|-----|
| Aviation   | 1   |
| Sanitation | 2   |
Mở rộng và tối ưu hóa pipeline dữ liệu với Polars

Lấy kết quả thực tế

actual = completed_by_department(


Mở rộng và tối ưu hóa pipeline dữ liệu với Polars

Lấy kết quả thực tế

actual = completed_by_department(
    sample.lazy()
).collect()
Mở rộng và tối ưu hóa pipeline dữ liệu với Polars

Lấy kết quả thực tế

actual = completed_by_department(
    sample.lazy()
).collect()
shape: (3, 2)
| DEPARTMENT | len |
| ---        | --- |
| str        | i32 |
|------------|-----|
| Aviation   | 1   |
| Sanitation | 2   |
| Water      | 1   |
Mở rộng và tối ưu hóa pipeline dữ liệu với Polars

So sánh bằng equals

actual.equals(expected)
False
Mở rộng và tối ưu hóa pipeline dữ liệu với Polars

Import công cụ kiểm thử của Polars

from polars.testing import (
    assert_frame_equal,
    assert_schema_equal,
)
Mở rộng và tối ưu hóa pipeline dữ liệu với Polars

Kiểm thử schema

assert_schema_equal(


)
Mở rộng và tối ưu hóa pipeline dữ liệu với Polars

Kiểm thử schema

assert_schema_equal(
    actual.schema,
    expected.schema,
)
AssertionError: Schemas are different (dtypes do not match)
[left]: [String, Int32]
[right]: [String, Int64]
Mở rộng và tối ưu hóa pipeline dữ liệu với Polars

Sửa schema mong đợi

expected = pl.DataFrame(
    {
        "DEPARTMENT": ["Aviation", "Sanitation"], "len": [1, 2],
    },
    schema={"DEPARTMENT": pl.String, "len": pl.Int32},
)
shape: (2, 2)
| DEPARTMENT | len |
| ---        | --- |
| str        | i32 |
|------------|-----|
| Aviation   | 1   |
| Sanitation | 2   |
Mở rộng và tối ưu hóa pipeline dữ liệu với Polars

Kiểm thử schema

assert_schema_equal(
    actual.schema,
    expected.schema,
)
print("Schema test passed!")
Schema test passed!
Mở rộng và tối ưu hóa pipeline dữ liệu với Polars

Kiểm thử DataFrame

assert_frame_equal(


)
Mở rộng và tối ưu hóa pipeline dữ liệu với Polars

Kiểm thử DataFrame

assert_frame_equal(
    actual,
    expected,
)
AssertionError: DataFrames are different height (row count) mismatch
[left]: 3
[right]: 2
Mở rộng và tối ưu hóa pipeline dữ liệu với Polars

So sánh kết quả thực tế và mong đợi

shape: (3, 2)
| DEPARTMENT | len |
| ---        | --- |
| str        | i32 |
|------------|-----|
| Aviation   | 1   |
| Sanitation | 2   |
| Water      | 1   |
shape: (2, 2)
| DEPARTMENT | len |
| ---        | --- |
| str        | i32 |
|------------|-----|
| Aviation   | 1   |
| Sanitation | 2   |
Mở rộng và tối ưu hóa pipeline dữ liệu với Polars

Sửa truy vấn

def completed_by_department(requests: pl.LazyFrame) -> pl.LazyFrame:
    return (
        requests
        .filter(pl.col("STATUS") == "Completed")
        .group_by("DEPARTMENT")
        .len()
        .with_columns(pl.col("len").cast(pl.Int32))
        .sort("DEPARTMENT")
    )
actual = completed_by_department(
    sample.lazy()
).collect()
Mở rộng và tối ưu hóa pipeline dữ liệu với Polars

Khẳng định cuối cùng

assert_schema_equal(actual.schema, expected.schema)
assert_frame_equal(actual, expected)
print("All tests passed!")
All tests passed!
Mở rộng và tối ưu hóa pipeline dữ liệu với Polars

Ayo berlatih!

Mở rộng và tối ưu hóa pipeline dữ liệu với Polars

Preparing Video For Download...