자주 쓰이는 단어 시퀀스

Python에서 Spark SQL 입문

Mark Plutowski

Data Scientist

훈련

Python에서 Spark SQL 입문

예측

Python에서 Spark SQL 입문

끝 단어 예측

Python에서 Spark SQL 입문

시퀀스

Python에서 Spark SQL 입문

마지막 시퀀스

Python에서 Spark SQL 입문

The quick brown fox

Python에서 Spark SQL 입문

문장 괄호

Python에서 Spark SQL 입문

다른 종류의 집계

Python에서 Spark SQL 입문

동영상

Python에서 Spark SQL 입문

범주형 데이터

Python에서 Spark SQL 입문

범주형 vs 순서형

  • 범주형: he, hi, she, that, they
  • 순서형: 1, 2, 3, 4, 5
Python에서 Spark SQL 입문

시퀀스 분석

Python에서 Spark SQL 입문

이전/다음 단어

Python에서 Spark SQL 입문

3-튜플

query3 = """
   SELECT 
   id,
   word AS w1,
   LEAD(word,1) OVER(PARTITION BY part ORDER BY id ) AS w2,
   LEAD(word,2) OVER(PARTITION BY part ORDER BY id ) AS w3
   FROM df
""" 
Python에서 Spark SQL 입문

서브쿼리로서의 윈도 함수 SQL

query3agg = """
SELECT w1, w2, w3, COUNT(*) as count FROM (
   SELECT 
   word AS w1,
   LEAD(word,1) OVER(PARTITION BY part ORDER BY id ) AS w2,
   LEAD(word,2) OVER(PARTITION BY part ORDER BY id ) AS w3
   FROM df
)
GROUP BY w1, w2, w3 
ORDER BY count DESC
""" 

spark.sql(query3agg).show()
Python에서 Spark SQL 입문

서브쿼리로서의 윈도 함수 SQL – 출력

+-----+-----+-----+-----+
|   w1|   w2|   w3|count|
+-----+-----+-----+-----+
|  one|   of|  the|   49|
|    i|think| that|   46|
|   it|   is|    a|   46|
|   it|  was|    a|   45|
| that|   it|  was|   38|
|  out|   of|  the|   35|
|.....|.....|.....|.....|
Python에서 Spark SQL 입문

가장 빈번한 3-튜플

+-----+-----+-----+-----+
|   w1|   w2|   w3|count|
+-----+-----+-----+-----+
|  one|   of|  the|   49|
|    i|think| that|   46|
|   it|   is|    a|   46|
|   it|  was|    a|   45|
| that|   it|  was|   38|
|  out|   of|  the|   35|
| that|    i| have|   35|
|there|  was|    a|   34|
|    i|   do|  not|   34|
| that|   it|   is|   33|
| that|   he|  was|   30|
| that|   he|  had|   30|
| that|    i|  was|   28|
+-----+-----+-----+-----+
Python에서 Spark SQL 입문

다른 종류의 집계

query3agg = """
SELECT w1, w2, w3, length(w1)+length(w2)+length(w3) as length FROM (
   SELECT 
   word AS w1,
   LEAD(word,1) OVER(PARTITION BY part ORDER BY id ) AS w2,
   LEAD(word,2) OVER(PARTITION BY part ORDER BY id ) AS w3
   FROM df
   WHERE part <> 0 and part <> 13
)
GROUP BY w1, w2, w3 
ORDER BY length DESC
""" 

spark.sql(query3agg).show(truncate=False)
Python에서 Spark SQL 입문

다른 종류의 집계

+-------------------+-------------------+---------------+------+
|                 w1|                 w2|             w3|length|
+-------------------+-------------------+---------------+------+
|comfortable-looking|           building|    two-storied|    38|
|         widespread|comfortable-looking|       building|    37|
|      extraordinary|      circumstances|      connected|    35|
|      simple-minded|      nonconformist|      clergyman|    35|
|       particularly|          malignant|  boot-slitting|    34|
|       unsystematic|        sensational|     literature|    33|
|       oppressively|        respectable|     frock-coat|    33|
|         relentless|        keen-witted|   ready-handed|    33|
|   travelling-cloak|                and|  close-fitting|    32|
|        ruddy-faced|      white-aproned|       landlord|    32|
|  fellow-countryman|            colonel|       lysander|    32|
+-------------------+-------------------+---------------+------+
Python에서 Spark SQL 입문

Ayo berlatih!

Python에서 Spark SQL 입문

Preparing Video For Download...