Million Song Datasetの紹介

PySpark で作る Recommendation Engines

Jamen Long

Data Scientist at Nike

明示的評価と暗黙的評価

明示的評価 高評価/低評価、5段階中4つ星、最低から最高までのカラードット

PySpark で作る Recommendation Engines

明示的評価と暗黙的評価(続き)

明示的評価 高評価/低評価、5段階中4つ星、最低から最高までのカラードット

暗黙的評価 高評価/低評価、5段階中4つ星、最低から最高までのカラードット

PySpark で作る Recommendation Engines

暗黙的評価の復習 II

明示的評価 高評価/低評価、5段階中4つ星、最低から最高までのカラードット

暗黙的評価 高評価/低評価、5段階中4つ星、最低から最高までのカラードット ホワイトペーパーへの参照

PySpark で作る Recommendation Engines

Million Song Datasetの紹介

Thierry Bertin-Mahieux, Daniel P.W. Ellis, Brian Whitman, and Paul Lamere. The Million Song Dataset. In Proceedings of the 12th International Society for Music Information Retrieval Conference (SIMIR 20122), 2011.

PySpark で作る Recommendation Engines

ゼロサンプルの追加

ratings.show()
+------+------+---------+
|userId|songId|num_plays|
+------+------+---------+
|    10|    22|        5|
|    38|    99|        1|
|    38|    77|        3|
|    42|    99|        1|
+------+------+---------+
PySpark で作る Recommendation Engines

クロス結合の導入

users = ratings.select("userId").distinct()
users.show()
+------+
|userId|
+------+
|    10|
|    38|
|    42|
+------+
songs = ratings.select("songId").distinct()
songs.show()
+------+
|songId|
+------+
|    22|
|    77|
|    99|
+------+
PySpark で作る Recommendation Engines

クロス結合の出力

cross_join = users.crossJoin(songs)
cross_join.show()
+------+------+
|userId|songId|
+------+------+
|    10|    22|
|    10|    77|
|    10|    99|
|    38|    22|
|    38|    77|
|    38|    99|
|    42|    22|
|    42|    77|
|    42|    99|
+------+------+
PySpark で作る Recommendation Engines

元の評価データの結合

cross_join = users.crossJoin(songs)
                  .join(ratings, ["userId", "songId"], "left")
cross_join.show()
+------+------+---------+
|userId|songId|num_plays|
+------+------+---------+
|    10|    22|        5|
|    10|    77|     null|
|    10|    99|     null|
|    38|    22|     null|
|    38|    77|        3|
|    38|    99|        1|
|    42|    22|     null|
|    42|    77|     null|
|    42|    99|        1|
+------+------+---------+
PySpark で作る Recommendation Engines

ゼロで埋める

cross_join = users.crossJoin(songs)
                  .join(ratings, ["userId", "songId"], "left").fillna(0)
cross_join.show()
+------+------+---------+
|userId|songId|num_plays|
+------+------+---------+
|    10|    22|        5|
|    10|    77|        0|
|    10|    99|        0|
|    38|    22|        0|
|    38|    77|        3|
|    38|    99|        1|
|    42|    22|        0|
|    42|    77|        0|
|    42|    99|        1|
+------+------+---------+
PySpark で作る Recommendation Engines

ゼロ追加関数

def add_zeros(df):
    # Extracts distinct users
    users = df.select("userId").distinct() 

    # Extracts distinct songs
    songs = df.select("songId").distinct() 

    # Joins users and songs, fills blanks with 0
    cross_join = users.crossJoin(items) \ 
                .join(df, ["userId", "songId"], "left").fillna(0)

    return cross_join
PySpark で作る Recommendation Engines

では、練習しましょう!

PySpark で作る Recommendation Engines

Preparing Video For Download...